Google Replaces BERT Self-Attention with Fourier Transform: 92% Accuracy, 7 Times Faster on GPUs
transformersfourier-transformbertnlpefficiency
Abstraction: FNet replaces transformer self-attention with Fourier Transform for faster training
Key points:
- FNet replaces each self-attention sublayer with an unparameterized Fourier Transform, achieving 92% of BERT accuracy on GLUE benchmark
- Training is 7x faster on GPUs and 2x faster on TPUs compared to BERT; lighter memory footprint across all sequence lengths
- FNet hybrid (only 2 self-attention sublayers retained) achieves 97% of BERT GLUE accuracy and trains ~6x faster on GPUs
- Self-attention has quadratic complexity with sequence length; Fourier approach eliminates this bottleneck
- 1D Fourier Transforms applied along both sequence and hidden dimensions; only the real part of the result is kept
- FNet is competitive with all evaluated "efficient transformers" on Long Range Arena benchmark
Connections: Google · Bert · Transformers · Attention Mechanism · Large Language Models