If you have 48 processors with 16 MiB cache each, you could load a
3/4 GiB of text into the caches and do a free text search in parallel on each core, without any indexing methods :-).
If you have 48 processors with 16 MiB cache each, you could load a
3/4 GiB of text into the caches and do a free text search in parallel on each core, without any indexing methods :-).
But the majority of systems are for a few big applications, plus lots of small one. And usually, the bug apps are active one. I would rather have 4 cores with 192M cache than 48 cores with 16M each.
Not sure if you are referring to the original 48 core CPU or the GA144. The GA144 processors aren't clocked, they are async. So when they are waiting on I/O they use virtually no power, 55 nA comes to mind.
Actually this chip was an exercise in "build it and see if anyone comes". They felt the chip was very cool and interesting and being in a
150 or 180 nm process, cheap to prototype. So why not? Unfortunately the lack of a target market shows up in the sales... I have three of them which I have yet to fire up.Certainly. The series can't actually extend to infinity as that would require infinite memory bandwidth. So at some point the coefficients do go to zero. BUT, that depends on the algorithms being run. As others have pointed out for some classes of problems this can work very well. My point was simply that the 5% may be enough gain to make it worth using the real estate to add the processor since chip space is so cheap.
I think that some aspect of the perceived limitations come from the way CPUs are utilized. Once we stop treating them as the precious resource we should find new ways of using them and see improvements in overall computer speed.
I worked for an FPS spin off called Star Technologies. FPS management felt the future was in 64 bit FP. Star went the faster route achieving
100 MFLOPS in a single machine - the only faster commercial floating point machines being the Crays. That is why the FPS-164 was not so fast, it was about precision. If you wanted speed, you bought the Star Techologies ST-100. That was the machine I learned about microcoding on.
Lol. Power can be a problem if you don't consider it from the outset, just as noise can be. It is hardly the major limiting factor in CPU design these days. That is the point of the 48 ARM core chip, efficient power usage, much more so than x86 Intel chips. That is why they are targeting servers. In servers the main problem they are fighting is processing power vs. power consumption.
They won't think about it until they are forced to because you can't change everything at once.
What does that matter? You don't use 10% of your brain at any given time. But you use 100% of it, sooner or later, just not all at once.
The brain is much more fixed, more like an FPGA fabric with a certain amount of distributed reconfiguration always going on. Each section only does one thing, and when that thing is not needed, it doesn't do anything.
The error, it follows, is in assuming that 1. you need more cores to do stuff, and 2. you need them all going at once for it to be worthwhile. Which for almost all tasks, as you well know, is a contradictory goal.
The point of the analogy would seem to be specializing cores for diverse tasks, so each one can do something very well, when it is occasionally needed. So you have your tasks in one corner, DSPs in another, GPUs here, user CPUs there... And various hardware (fixed, reconfigurable or microcodeable) spread about.
It should seem silly to think that having everything uniform would be a winning strategy. Fortunately, nature has no such restrictions on understanding or ease of software implementation; thus, neurons in various parts of the brain naturally connect and behave very differently.
Tim
So what's the scheduler going to do with that information? Does it know which task the user thinks is critical? Temperature (open or closed loop) is already used to modulate the power supplies and clocks.
Rather than many identical cores, it's more likely that there will be different cores with more specialized functions (at different power/performance points). We're already seeing some of that in the embedded market. We use one that has two A15 cores, a couple of M4 cores, a DSP, and graphics controllers. It doesn't do much useful without dynamic frequency and voltage scaling.
Sure. You can't beat parallel wiring. ;-)
I sure wish OSs had clearer priority settings now.
Temperature (open or
Then it could be used to throttle individual core speeds.
An asymmetric architecture would make sense. A file system or an Ethernet socket doesn't need a lot of floating point, for example.
Yeah, latency really sucks. PCs, especially x86, are built for bulk throughput through "channels", with pitiful realtime access to anything.
A Gen1 4-lane PCIe read of a single byte takes most of a microsecond.
The second byte is pretty quick, though. ;-) It's an I/O channel, not a memory interface. Latency doesn't matter (mostly).
Sony Walkman was a bit like that at first. Only their CEO believed in it. Same for 3M's Postit notes which were made using a "failed" super adhesive chemistry that did not perform at all as expected.
Indeed but the sad thing is that many important computing problems do not parallelise at all well.
Again I think it will depend on the alternative solutions. Multicore computing is a lot harder than single core and things can go haywire.
Actually I think you might have it slightly backwards here. I expect to see slower HCI processors and faster SSD disk concentrators with the general purpose CPU cores sitting somewhere in the middle and much more effort being made to consume less power for roughly equal compute power. It is the power usage that will be precious in the future. Most portable kit now has way too much grunt and not enough battery life.
Fast PC graphics cards will go towards massively parallel and there are already tools to subvert them for cheaper scientific computing.
My i7-3770K (not overclocked) runs at such a low TDP that the CPU fan sometimes switches itself off if the machine is idling and it packs a hell of a punch when it is running flat out about 100W. I don't have any graphics card at all and rely its internal engine which benchmarks on 2D slightly better than some of the high end cards. Obviously it is rubbish for 3D gaming although actually faster than I expected.
Yes. I had forgotten the significance of that "64" in the name. Suits and slick marketing men vs engineers problem - no doubt the focus groups did all come back and say they wanted 64 bit precision.
Reading a single byte from a SDRAM card is also pretty slow, reading multiple bytes e.g. loading a cache line (32 bytes in x86) will have a decent throughput (one RAS and four CAS cycles).
Any operations requiring frequent request/response transactions are finally limited by the speed of light. Assuming 10 cm distance between two devices will have an additional 1 ns (on ordinary PCB materials) due to the two way propagation delay, in addition to the internal processing time of the addressed device.
For anything with a single request in one direction and burst transfer in the other direction, there is the same light speed penalty during request setup, but during a burst transfer, the speed of light is not an issue.
The system architecture should be designed to effectively use burst transfers or unidirectional transfers.
So what is your point? Because there have been great successes that came from what many thought would be failed products doesn't mean this one is a winner. It has been on the market for years now and as far as I can tell has very few customers, likely none of any real size.
I think the world is aware of that, but why does everyone think in terms of parallelizing a single process? For server applications they need to run many, many instances of the same processes. That can parallelize very well.
Ok, do you have anything more specific to say on the topic?
HCI processors? Hot Carrier Injection?
1k and 2k ALUs isn't enough for you? What exactly is "massive"?I can't say which did better in the market place. I worked in manufacturing on the test floor at the time. We were under the gun to get machines out the door and by the end of a year we had SN 60 to 100 sitting by the loading dock. It wasn't long after that they shifted gears to a smaller, 50 MFLOPS unit that was the DSP chip (in "just" a single card cage as compared to the two rack cabinets of the ST-100) for a GE line of CAT scanners. That was the core business for Star for a number of years until they jacked up the name and did a business transplant, sold the tiny maintenance contract for the few remaining machines and bought a document scanning company. Very weird. In another year or so they were gone. FPS at least survived long enough to be absorbed by Cray and then was it Sun? I just closed the browser window on the pages about FPS including a brochure on the FPS-164.
BTW, marketing had some impact on the demise of Star. 100 MFLOPS was a magic number so they pushed for that at all costs. Turns out the only way they could get the machines to work at that speed was by hand tuning portions of it including trimming the length of the clock distribution wires to adjust the phase to each of the boards. We had a couple of Chinese magicians working in "final test". Their job was to swap the ECL gate array chips around until they could get the machine to pass all the tests. When Eric Chen finally told me what was going on I realized they were just plain violating the time specs on the chips. It would have made an ok ST-90 and a rock solid ST-80, but as an ST-100 the machine was a maintenance nightmare not to mention a problem for the users who could never trust the results. I learned a few things about what *not* to do working for that company.
"The world only needs three computers". This famous statement assumed that these three machines would be enough to solve all the problems of the world.
However, we are no longer living in a world in which a huge computer would have to solve _all_ problems of a company, department or even a single person. These days "computers" can be dedicated to solve a single problem, there is no need to be handle all kinds of problems (often with poor performance on average).
Previously IBM with their System/360 (360 degrees=all around) and now Intel/Microsoft are still pushing everything into a single general purpose iron. Use high degree of parallel processing when there are some real advantage, otherwise use something simpler.
If it doesn't give you any real advantage, don't use it.
While 64 bit double precision floating point is nice to have for internal computations, so that you do not have to worry (too much) about range and accuracy, one still has to remember that the real analog world is usually just represented by 8-12-16 bit ADC and DACs.
I also wonder, what is the point with systems with 64 bit integer and
64 bit virtual addresses, most applications simply do not need them. I guess that was one of the reasons why the 64 bit DEC Alpha was not that successful.Of course doing some big data base applications consisting of up to RAID system with multiple TB disks, I would definitively demand a 64 bit processor with 64 bit OS to easily handling memory mapped files.
Of course, people thought the same thing when the 80386 was introduced. Four gigabytes address space? Unimaginable even in 1995 by the time Pentiums were rolling in with PAE.
Soon enough, we'll have complete systems with petabytes of contiguous RAM. Presumably with stacked chip technology, or something altogether more exotic. Hopefully, it'll be integrated with the CPUs too, since all those interconnects would seriously suck on a PCB.
Tim
On a sunny day (Sat, 7 Jun 2014 07:51:55 -0500) it happened "Tim Williams" wrote in :
OTOH cars with 5 or more wheels remain rare. F1 still uses 4.
That already exists, it is called internet.
Embedded apps often need fast CPU access to real-world i/o registers, which is flat impossible on modern Intel machines, hence on PC motherboards. The memory bus is burst-mode DRAM stuff, and the i/o busses are PCIe and USB and maybe Thunderbolt. Slow and very complex.
Looks like the future of serious embedded is SOCs that have ARM cores and FPGA fabric on the same chip. ZYNQ is OK, but the CPU-to-FPGA connection is still a bit loose.
Unfortunately they messed up even that. The segment registers are activated even in the 32 bit mode (usually all set to 0). However, setting these segment registers to non-zero would have allowed one 4 GiB code segment and several 4 GiB data/stack segments. Unfortunately the segmented addresses are truncated to 32 bits _before_ being presented to the virtual to physical address translation, so in reality, you just have a single combined 4 GiB virtual address space.
That is the physical address space and it will grow during the years. I was arguing about the virtual address space. I still think 4 GiB is more than enough for instruction (code) space. I don't think that even Microsoft could produce such bloatware. Testability of such monster programs is really an issue.
For reliability, you do not want to put too much ordinary R/W data into a huge single address space, but rather divide the application into several processes each running in a private protected address spaces.
Of course, there are some few applications, which might benefit from more than 4 GiB arrays.
One F1 team tried to use 6 wheels, but it was banned.
If you intend to read one byte, then that is true. However, loading a cache line or a virtual memory page, this is not much of an issue.
To load a 4 KiB virtual memory over a 10 Gb/s link (about 1 GB/s) takes about 4 us. With the storage located 400 m from the processor would drop the throughput to 50 %. With multiple simultaneous open transactions, the drop would be even less.
Are we talking about bit banging, some high speed serial stuff or few channels of 1 Gb/s Ethernet ports ?
In the two latter cases, I would look at some PowerPC with the QUICC coprocessor. Even the old MC68360 had a well integrated QUICC coprocessor usable mainly for serial operations.
Have something to add? Share your thoughts — no account required.
Ask the community — no account required