linear attention
A variant of attention mechanisms that scales linearly with input size, improving efficiency in processing sequences without compromising on performance.
- Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
- Exploring Diffusion Transformer Designs via Grafting
- Exploring Diffusion Transformer Designs via Grafting
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Learning Linear Attention in Polynomial Time
- Linear Attention for Efficient Bidirectional Sequence Modeling
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Pseudo-Riemannian Graph Transformer
- Radial Attention: $\mathcal O(n \log n)$ Sparse Attention for Long Video Generation
- ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention
- ZeroS: Zero‑Sum Linear Attention for Efficient Transformers