(57) Contemporary computer architectures support pipelining and instruction-level parallelism,
in which an operation related to updating array T1 is followed by an operation related to updating array T2. To further improve performance, it is desirable to update arrays T1 and T2 together within the loop over i, using one instruction per pair of updates. The computer architecture employs SIMD instructions (and in our particular example,
SSE2 instructions) to update simultaneously the arrays T1 and T2 used in Montgomery multiplication, montmul (A,B). This process involves, in part, processing digits of A in a right-to-left manner and
interleaving results of the multiplications involving B (i.e., mul1*B[i]) and modulus N (i.e., mul2*N[i]) that are used to update arrays T1 and T2.
|

|