Nicely stated. I don't completely agree with you so far. Clear requirements (or at least clean problem statements) and proper reviews go very far in creating good software. They can even be implemented sucessfully in one programmer shops. Having that occur is another matter.
Didn't find your answer? Ask the community — no account required.
J
JosephKK
And you say it is less than 100 KLOC, i could probably understand that with some effort in a reasonable amount of time. Could i have a copy?
J
JosephKK
I expect so, i think even linnix may have to admit that 6x 1.5 Tb/s (HyperTransport) is a lot of bandwidth.
P
Phil Hobbs
It's all single precision. And I won't need that many cores--a 12-core Opteron is about 100 Gflop peak all on its lonesome. I think the 1U ASUS pizza box will work fine, assuming the job comes in. My colleagues are giving the second pitch on Monday, I think, so I should know soon.
The simulation domains won't be that big. The way FDTD works is to update every main memory location once per cycle, which involves crunching through it twice per cycle (once for E and once for H). That puts a real strain on the main memory bandwidth because there are a lot of cache misses. With four DDR3 1333 channels, I should be able to keep things moving pretty well, and avoid the latency of Ethernet.
The problem with clusters for this job is that the type of communications is lowish in bandwidth but very sensitive to latency. The space is divided up into shoeboxes full of sugar cubes, with each core handling one or more of them. On each half step, each shoebox needs the field values on the surfaces of all 6 of its neighbours (some of which may be itself, depending on the geometry and the boundary conditions).
Cheers
Phil Hobbs
Dr Philip C D Hobbs
Principal
ElectroOptical Innovations
55 Orchard Rd
Briarcliff Manor NY 10510
845-480-2058
hobbs at electrooptical dot net
http://electrooptical.net
L
linnix
I would like to know where you get that number. A 2.5 GHz core would have 3 to 4 GFlops. But additional cores are bumping up the law of diminishing returns on cores. AMD's benchmark shows a 15% to 20% increases from adding 4 cores to 8 (50% increase in cores). Beyond 12 cores (one or more chips), it's just adding resistance (resistance is futile) between cores.
t
PCIe is just as good. 2.5Gb/s for 5m. 80Gb/s for shorter runs.
L
linnix
assis
DDR3
=A0That
each,
0
$ 700
=A0 =A0 $6220 =A0plus 7% tax =3D $6655
The
..
e
ES)
, 1.87
, 2.93
=3D=3D=3D
=3D||
=A0||||
ble
=A0||||
s
Yes, but they (the cores) still have to wait in slow line for main memory. Adding cores without adding memory paths works up to a limited number of cores.
In my design, I would have 16 PM (processor/Memory) modules on PCIe. P (Atom of A9) may be dual cores or quad cores (in 2 years).
P
Phil Hobbs
I lied. It's 105 Gflops for one 12-core, 420 for four, according to Dell.
formatting link
Cheers
Phil Hobbs
Dr Philip C D Hobbs
Principal
ElectroOptical Innovations
55 Orchard Rd
Briarcliff Manor NY 10510
845-480-2058
hobbs at electrooptical dot net
http://electrooptical.net
L
linnix
ore
2
est
Those are estimates for the 8374 series with on chip 8M L3 cache. They are $800 each for quad core. No 6, 8 or 12 cores.
L
linnix
-core
ld
12
ggest
Sorry, wrong link.
Yes, Dell servers are better, with multiple memory lane per socket. Around $5000 for 2x8 (2 socket 8 cores).
L
linnix
12-core
ould
of
nd 12
is
suggest
.
o
Also, keep your memoy working set within 12M, since that's what the benchmark is based.
J
JosephKK
If the software was partitioned such that there was a reasonable memory partitioning so that per card locality was properly advantageous. If access globality is really necessary, your system would die in IPC costs. If it took serious recoding to acheive the memory modularity, schedule may not be met. I am not saying that your sloution is not a cost effective solution for this class of applications, it may not be the right one for this iteration of this design task.
J
JosephKK
I didn't want to believe it. But here is credible backup:
formatting link
From what i have been reading you could get a nice performance goose from faster northbridge chips.
From that description there may not be good optimizations for a more blocky memory architecture (clustery, multiple attached processors with segreated memory, vs flat memory SMP).
L
linnix
chassis
GB DDR3
. =A0That
M.
40 each,
2000
=A0 $ 700
=A0 =A0 $6220 =A0plus 7% tax =3D $6655
=A0The
he
e-...
core
SLES)
555, 1.87
670, 2.93
t
=3D=3D=3D=3D
=3D=3D||
=A0||||
cable
=A0||||
it
I
as
t
sses
big
I would have around 1GB local memory per CPU.
ASUS/Dell multi-cores shared memory server would have died much earlier. The quoted performance of 200GFlops is for a small working menory set of 24MB, contained in L3 cache. Beyond 24MB, it would be difficult to get more than 30 to 40GFlops average.
If ASUS/Dell can just put together COTS parts and build super- computers cheaply, why is IBM/Fujisu wasting money will all those employees on payroll.
J
JosephKK
Hmm. There is a memory bandwidth issue of only 4 X 28 GB/s (northbridge limitation). I guess a lot would depend on how well the algorithm stays within cache (and some on how much stays with the FP registers). I presume you know how to apply Amdahl's law and evaluate the cache miss rates at least as well as i.
Since ASUS/Dell and others are doing just that, you have a very legitimate question. Ask them.
P
Phil Hobbs
Well, it came in, hurrah.
What I wound up ordering is the following:
1 ea CSS-SMI-1022GNT SUPERMICRO 1022G-NTF BAREBONE 1U
2 ea CPO-AMD-6172 AMD OPTERON 6172
2 ea HDA-WDC-WD2002F WD WD2003FYYS 2TB SATA 7200RPM
8 ea MM3-KIN-8G133ER KINGSTON 8GB DDR3 1333 ECC REGISTERED
So sometime next week I should have a 19x30x1 inch box with 200 GFlops peak and 64 gig of memory. Fun.
Cheers
Phil Hobbs
Dr Philip C D Hobbs
Principal
ElectroOptical Innovations
55 Orchard Rd
Briarcliff Manor NY 10510
845-480-2058
email: hobbs (atsign) electrooptical (period) net
http://electrooptical.net
C
Cydrome Leader
1.75" thick if it's a 1U server.
P
Phil Hobbs
Integer arithmetic. ;)
Should be here next week.
Cheers
Phil Hobbs
Dr Philip C D Hobbs
Principal
ElectroOptical Innovations
55 Orchard Rd
Briarcliff Manor NY 10510
845-480-2058
email: hobbs (atsign) electrooptical (period) net
http://electrooptical.net
J
Jan Panteltje
On a sunny day (Sat, 25 Sep 2010 14:26:39 -0400) it happened Phil Hobbs wrote in :
If it is not fast enough you could fly away and come back at close to light speed,
1 hour later to pick up the results [1]. Even works with a Z80.
1) Twin paradox.
Join the Discussion
Have something to add? Share your thoughts — no account required.
Didn't find your answer?
Ask the community — no account required
Report Content
You are reporting this content to the moderators. They will look at it
ASAP.