sparse attention
Sparse attention is a mechanism used in neural networks, particularly in transformers, which focuses on only a subset of input elements during training and inference, improving efficiency and reducing computational overhead while retaining performance.
- A unified framework for establishing the universal approximation of transformer-type architectures
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
- Kinetics: Rethinking Test-Time Scaling Law
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape
- Spark Transformer: Reactivating Sparsity in Transformer FFN and Attention
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
- The emergence of sparse attention: impact of data distribution and benefits of repetition
- The emergence of sparse attention: impact of data distribution and benefits of repetition
- Transformers Learn Faster with Semantic Focus