Professional Documents
Culture Documents
Schönhage-Strassen’s Algorithm
Alexander Kruppa
CACAO team at LORIA, Nancy
Contents
(b) Evaluation/Interpolation
3. Schönhage-Strassen’s Algorithm
4. High-Level Improvements
(a) Mersenne Transform
Integer Multiplication
• Problem: given two n-word (word base β ) integers
n−1
!
a= aiβ i,
i=0
with 0 ≤ ci < β .
by Polynomial Multiplication
• We can rewrite the problem as polynomial arithmetic:
n−1
!
A(x) = aixi
i=0
so that c = ab = C(β).
Evaluation/Interpolation
• Unisolvence Theorem: Polynomial of degree d − 1 is determined
uniquely by values at d distinct points
Karatsuba’s Method
• First algorithm to use this principle (Karatsuba and Ofman, 1962)
Toom-Cook Method
• Generalized to polynomials of larger degree (Toom, 1963, Cook,
1966)
Weighted Transform
• Since (ω#j )# = 1 for all j ∈ N, C1(x)x# + C0(x) has same DFT
coefficients as C1(x) + C0(x): implicit modulus x# − 1 in DFT
Schönhage-Strassen’s Algorithm:
Basic Idea
• First algorithms to use FFT (Schönhage and Strassen 1971)
Schönhage-Strassen’s Algorithm
• SSA reduces multiplication of two S -bit integers to # multiplications
of approx. 4S/#-bit integers
High-Level Optimizations
Mersenne Transform
• Convolution theorem implies reduction mod(x# − 1)
CRT Reconstruction
a, b a, b
ab CRT
with 2N + 1 > ab
ab
with (2rN + 1)(2sN − 1) > ab
√
The 2 Trick
• If 4 | n, 2 is a quadratic residue in Z/(2n + 1)Z
√ 3n/4 n/4
√
• In that case, 2 ≡ 2 − 2 : simple form, multiplication by 2
takes only 2 shift, 1 subtraction modulo 2n + 1
Low-Level Optimizations
Arithmetic modulo 2n + 1
• Residues stored semi-normalized (< 2n+1), each with m = n/w
full words plus one extra word ≤ 1
√
• Bailey’s 4-step algorithm (radix # transform)
- groups half of recursive levels into first pass, other half into
second pass
√
- Each pass is again a set of FFTs, each of length #
√
- If length # transform fits in level 2 cache: only two passes
over memory per transfrom
- Extremely effective for complex floating-point FFTs, we found it
useful for SSA with large # as well
• Fusing different stages of SSA
- Do as much work as possible on data while it is in cache
When cutting input a into M -bit size coefficients a =
-"
iM
0≤i<# a i 2 , also apply weights for negacyclic transform and
perform first FFT level (likewise for b)
- Similar ideas for other stages
Fine-Grained Tuning
• Up to version 4.2.1, GMP uses simple tuning scheme: transform
length grows monotonously with input size
0.0002
K=32
0.00019 K=64
0.00018
Seconds
0.00017
0.00016
0.00015
0.00014
0.00013
700 750 800 850 900
Size in 64-bit words
80
GMP 4.1.4
70 Magma V2.13-6
GMP 4.2.1
60 new GMP code
50
Seconds
40
30
20
10
0
2.5e6 5e6 7.5e6 1e7 1.25e7 1.5e7
Size in 64-bit words
Untested Ideas
• Special code for point-wise multiplication:
- Length 3 · 2k transform
References