Good point, although I have also experince about this situtation, which I didn't seem to bother writing about, but here it goes:
~10 years ago I was involved in the Amiga demo scene, with mostly realtime 3D graphics. For linear interpolation (for good results for Gouraud shading, and not so good with texture mapping, since the lack of perspective division.. :)) addx.l (add with extend) insn was used. It allowed linear interpolation of fixed-point values without right-shifting the interpolated value:
Gouraud shading with 16.16 precision, although this precision wasn't very useful with only 16 colours, even the gouraudTable was dithered, but the quality was quite good, and even a lowly Amiga 1200 with 14MHz
68020 could achieve quite nice frame rates:lea gouraudTable,a0 lea chipRAMBuffer,a1 ... .innerloop move.l (a0,d0.w*4),(a1)+ # 4 bits per pixel: 8 pixels # drawn from a LUT addx.l d1,d0 # g += dg; dbf d2,.innerloop
(edges had to be handled separately with and/or)
Gouraud shading, "chunky mode", up to 24.8 precision: lea chunkyBuffer,a0 ... .innerloop move.b d0,(a0)+ addx.l d1,d0 dbf d2,.innerloop
Texture mapping (16.16), "chunky mode" (byte == one pixel, 256 colours):
lea textureMap,a0 # 256x256 texture map lea chunkyBuffer,a1 # a buffer to be c2p'ed to # chip ram, as fast as # just copying from fast2chip ...
# large bits denote integer part, small bits fraction # d0: u: %vvvvvvvv0000000000000000UUUUUUUU # d1: v: %uuuuuuuu00000000VVVVVVVVvvvvvvvv # d2: du: %vvvvvvvv0000000000000000UUUUUUUU # d3: dv: %uuuuuuuu00000000VVVVVVVVvvvvvvvv
sub.w d2,d0 # reduce error from using addx.l d3,d1 # two addx.l .innerloop move.w d1,d4 # v is good enough. And if you know to think like a compiler, you can also
For most transformations (assuming a modern compiler), you don't have to "think like a compiler". The compiler does it for you.
If the access to a static variable is an issue in time critical code, you are doing something funny (ISRs and volatile of course is a different issue).
Anyway, related to your "packing the variables into a struct", modern compilers are quite often to do this:
/* before optimization */ static int key[SIZE]; static int value[SIZE];
/* afterwards */ struct merged { int key, value; };
static struct merged[SIZE];
Loop optimizations - for better cache usage - that are usually done include loop exchange and loop fusion. Pretty standard stuff.
Examples:
(better cache hit rate by having better spatial locality)
/* before loop exchange */ for (j = 0; j < 100; j++) for (i = 0; i < 100; i++) x[i][j] = 2 * x[i][j];
/* after */ for (i = 0; i < 100; i++) for (j = 0; j < 100; j++) x[i][j] = 2 * x[i][j];
(better cache hit rate by accessing memory "horizontally" instead of "vertically" - better temporal locality)
/* before loop fusion */ for (i = 0; i < N; i++) for (j = 0; j < N; j++) a[i][j] = 1 / b[i][j] * c[i][j]; for (i = 0; i < N; i++) for (j = 0; j < N; j++) d[i][j] = a[i][j] + c[i][j];
/* after */ for (i = 0; i < N; i++) for (j = 0; j < N; j++) { a[i][j] = 1 / b[i][j] * c[i][j]; d[i][j] = a[i][j] + c[i][j]; }
(these examples were from my seminar report in 2003, sadly it is in Finnish..) In general, programming languages having pointers have a problem called aliasing, which may prevent the optimizer for functioning 100%. This is one reason why Just In Time compiled Java can be even faster than statically compiled C or C++ code. Another reason is that the JIT compiler is able to analyze the behaviour of the program, although this can be done with statically compiling compilers also, with the limitation that one has only one profile (== one case of input) of the executing program, which can be used for optimization, while the JIT compiler can do it all the time.