attention layers
Components in neural networks, particularly transformers, that allow the model to focus on certain parts of the input more than others, enhancing its ability to understand context and relationships.
- FFN Fusion: Rethinking Sequential Computation in Large Language Models
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
- Spectral Conditioning of Attention Improves Transformer Performance
- UMoE: Unifying Attention and FFN with Shared Experts
- Wavy Transformer