Budget Linux cluster suggestions?

Aug 14, 2010 97 Replies

Nicely stated. I don't completely agree with you so far. Clear requirements (or at least clean problem statements) and proper reviews go very far in creating good software. They can even be implemented sucessfully in one programmer shops. Having that occur is another matter.

And you say it is less than 100 KLOC, i could probably understand that with some effort in a reasonable amount of time. Could i have a copy?

I expect so, i think even linnix may have to admit that 6x 1.5 Tb/s (HyperTransport) is a lot of bandwidth.

It's all single precision. And I won't need that many cores--a 12-core Opteron is about 100 Gflop peak all on its lonesome. I think the 1U ASUS pizza box will work fine, assuming the job comes in. My colleagues are giving the second pitch on Monday, I think, so I should know soon.

The simulation domains won't be that big. The way FDTD works is to update every main memory location once per cycle, which involves crunching through it twice per cycle (once for E and once for H). That puts a real strain on the main memory bandwidth because there are a lot of cache misses. With four DDR3 1333 channels, I should be able to keep things moving pretty well, and avoid the latency of Ethernet.

The problem with clusters for this job is that the type of communications is lowish in bandwidth but very sensitive to latency. The space is divided up into shoeboxes full of sugar cubes, with each core handling one or more of them. On each half step, each shoebox needs the field values on the surfaces of all 6 of its neighbours (some of which may be itself, depending on the geometry and the boundary conditions).

Cheers

Phil Hobbs

Dr Philip C D Hobbs Principal ElectroOptical Innovations 55 Orchard Rd Briarcliff Manor NY 10510 845-480-2058 hobbs at electrooptical dot net http://electrooptical.net

I would like to know where you get that number. A 2.5 GHz core would have 3 to 4 GFlops. But additional cores are bumping up the law of diminishing returns on cores. AMD's benchmark shows a 15% to 20% increases from adding 4 cores to 8 (50% increase in cores). Beyond 12 cores (one or more chips), it's just adding resistance (resistance is futile) between cores.

t

PCIe is just as good. 2.5Gb/s for 5m. 80Gb/s for shorter runs.

assis

DDR3

=A0That

each,

0

$ 700

=A0 =A0 $6220 =A0plus 7% tax =3D $6655

The

..

e

ES)

, 1.87

, 2.93

=3D=3D=3D

=3D||

=A0||||

ble

=A0||||

s

Yes, but they (the cores) still have to wait in slow line for main memory. Adding cores without adding memory paths works up to a limited number of cores.

In my design, I would have 16 PM (processor/Memory) modules on PCIe. P (Atom of A9) may be dual cores or quad cores (in 2 years).

I lied. It's 105 Gflops for one 12-core, 420 for four, according to Dell.

formatting link

Cheers

Phil Hobbs

Dr Philip C D Hobbs Principal ElectroOptical Innovations 55 Orchard Rd Briarcliff Manor NY 10510 845-480-2058 hobbs at electrooptical dot net http://electrooptical.net

ore

2

est

Those are estimates for the 8374 series with on chip 8M L3 cache. They are $800 each for quad core. No 6, 8 or 12 cores.

-core

ld

12

ggest

Sorry, wrong link.

Yes, Dell servers are better, with multiple memory lane per socket. Around $5000 for 2x8 (2 socket 8 cores).

12-core

ould

of

nd 12

is

suggest

.
o

Also, keep your memoy working set within 12M, since that's what the benchmark is based.

If the software was partitioned such that there was a reasonable memory partitioning so that per card locality was properly advantageous. If access globality is really necessary, your system would die in IPC costs. If it took serious recoding to acheive the memory modularity, schedule may not be met. I am not saying that your sloution is not a cost effective solution for this class of applications, it may not be the right one for this iteration of this design task.

I didn't want to believe it. But here is credible backup:

formatting link

From what i have been reading you could get a nice performance goose from faster northbridge chips.

From that description there may not be good optimizations for a more blocky memory architecture (clustery, multiple attached processors with segreated memory, vs flat memory SMP).

chassis

GB DDR3

. =A0That

M.

40 each,
2000

=A0 $ 700

=A0 =A0 $6220 =A0plus 7% tax =3D $6655

=A0The

he

e-...

core

SLES)

555, 1.87
670, 2.93
t

=3D=3D=3D=3D

=3D=3D||

=A0||||

cable

=A0||||

it

I

as

t

sses

big

I would have around 1GB local memory per CPU.

ASUS/Dell multi-cores shared memory server would have died much earlier. The quoted performance of 200GFlops is for a small working menory set of 24MB, contained in L3 cache. Beyond 24MB, it would be difficult to get more than 30 to 40GFlops average.

If ASUS/Dell can just put together COTS parts and build super- computers cheaply, why is IBM/Fujisu wasting money will all those employees on payroll.

Hmm. There is a memory bandwidth issue of only 4 X 28 GB/s (northbridge limitation). I guess a lot would depend on how well the algorithm stays within cache (and some on how much stays with the FP registers). I presume you know how to apply Amdahl's law and evaluate the cache miss rates at least as well as i.

Since ASUS/Dell and others are doing just that, you have a very legitimate question. Ask them.

Well, it came in, hurrah.

What I wound up ordering is the following:

1 ea CSS-SMI-1022GNT SUPERMICRO 1022G-NTF BAREBONE 1U 2 ea CPO-AMD-6172 AMD OPTERON 6172 2 ea HDA-WDC-WD2002F WD WD2003FYYS 2TB SATA 7200RPM 8 ea MM3-KIN-8G133ER KINGSTON 8GB DDR3 1333 ECC REGISTERED

So sometime next week I should have a 19x30x1 inch box with 200 GFlops peak and 64 gig of memory. Fun.

Cheers

Phil Hobbs

Dr Philip C D Hobbs Principal ElectroOptical Innovations 55 Orchard Rd Briarcliff Manor NY 10510 845-480-2058 email: hobbs (atsign) electrooptical (period) net http://electrooptical.net

1.75" thick if it's a 1U server.

Integer arithmetic. ;)

Should be here next week.

Cheers

Phil Hobbs

Dr Philip C D Hobbs Principal ElectroOptical Innovations 55 Orchard Rd Briarcliff Manor NY 10510 845-480-2058 email: hobbs (atsign) electrooptical (period) net http://electrooptical.net

On a sunny day (Sat, 25 Sep 2010 14:26:39 -0400) it happened Phil Hobbs wrote in :

If it is not fast enough you could fly away and come back at close to light speed,

1 hour later to pick up the results [1]. Even works with a Z80. 1) Twin paradox.

Join the Discussion

Have something to add? Share your thoughts — no account required.

Didn't find your answer?

Ask the community — no account required