transformer models
Models based on the transformer architecture, which have become predominant in various AI applications, especially in natural language processing and computer vision. These include architectures like BERT, GPT, and Vision Transformers.
- CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
- Efficient Large Language Model Inference with Neural Block Linearization
- Enabling Differentially Private Federated Learning for Speech Recognition: Benchmarks, Adaptive Optimizers, and Gradient Clipping
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
- Learning Linear Attention in Polynomial Time
- Learning Linear Attention in Polynomial Time
- Mamba Modulation: On the Length Generalization of Mamba Models
- Memory Mosaics at scale
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Rope to Nope and Back Again: A New Hybrid Attention Strategy
- STree: Speculative Tree Decoding for Hybrid State Space Models
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
- StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
- UMoE: Unifying Attention and FFN with Shared Experts
- Unveiling Transformer Perception by Exploring Input Manifolds
- Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models