transformer
A neural network architecture that uses the attention mechanism to process sequential data. Transformers excel at capturing long-range dependencies and have become the backbone of many modern NLP and image processing tasks.
- BlockScan: Detecting Anomalies in Blockchain Transactions
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
- Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- EddyFormer: Accelerated Neural Simulations of Three-Dimensional Turbulence at Scale
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- Linear Transformers Implicitly Discover Unified Numerical Algorithms
- LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers
- Multi-head Temporal Latent Attention
- Quantum Doubly Stochastic Transformers
- Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity Modeling
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
- SEMPO: Lightweight Foundation Models for Time Series Forecasting
- Spectral Conditioning of Attention Improves Transformer Performance
- Who Reasons in the Large Language Models?