transformers
A type of deep learning architecture that utilizes self-attention mechanisms to process sequential data. Transformers have revolutionized NLP and other fields due to their ability to model context and relationships effectively.
- A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- Attention-based clustering
- Attribution-Driven Adaptive Token Pruning for Transformers
- Benign Overfitting in Single-Head Attention
- Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
- CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
- Cameras as Relative Positional Encoding
- Compositional Reasoning with Transformers, RNNs, and Chain of Thought
- Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis
- Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
- Dependency Parsing is More Parameter-Efficient with Normalization
- Depth-Width Tradeoffs for Transformers on Graph Tasks
- EUGens: Efficient, Unified and General Dense Layers
- Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
- Fast attention mechanisms: a tale of parallelism
- First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training
- Fourier Analysis Network
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- From Softmax to Score: Transformers Can Effectively Implement In-Context Denoising Steps
- Gemstones: A Model Suite for Multi-Faceted Scaling Laws
- Generalization vs Specialization under Concept Shift
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- How do Transformers Learn Implicit Reasoning?
- HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
- In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separation
- Knee-Deep in C-RASP: A Transformer Depth Hierarchy
- Length Generalization via Auxiliary Tasks
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied Learning
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- MVSMamba: Multi-View Stereo with State Space Model
- Memory Mosaics at scale
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
- Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
- Normalizing Flows are Capable Models for Continuous Control
- On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting
- Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- REOrdering Patches Improves Vision Models
- Reproducing Kernel Banach Space Models for Neural Networks with Application to Rademacher Complexity Analysis
- Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node Classification
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
- Rotary Masked Autoencoders are Versatile Learners
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- Scaling and context steer LLMs along the same computational path as the human brain
- Spark Transformer: Reactivating Sparsity in Transformer FFN and Attention
- SpiderSolver: A Geometry-Aware Transformer for Solving PDEs on Complex Geometries
- Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
- The Logical Expressiveness of Temporal GNNs via Two-Dimensional Product Logics
- The emergence of sparse attention: impact of data distribution and benefits of repetition
- The emergence of sparse attention: impact of data distribution and benefits of repetition
- Towards Understanding Transformers in Learning Random Walks
- TractoTransformer: Diffusion MRI Streamline Tractography using CNN and Transformer Networks
- Transformers Learn Faster with Semantic Focus
- Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
- Transformers are almost optimal metalearners for linear classification
- Two Heads are Better than One: Simulating Large Transformers with Small Ones
- Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
- VideoTitans: Scalable Video Prediction with Integrated Short- and Long-term Memory
- Wavy Transformer
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
- ZeroS: Zero‑Sum Linear Attention for Efficient Transformers