attention mechanisms
Techniques in neural networks that allow the model to focus on specific parts of the input data when making predictions, enhancing performance in tasks such as natural language processing and computer vision.
- A unified framework for establishing the universal approximation of transformer-type architectures
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators
- FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network
- Foundation Cures Personalization: Improving Personalized Models’ Prompt Consistency via Hidden Foundation Knowledge
- Generalizable Insights for Graph Transformers in Theory and Practice
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
- Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
- JADE: Joint Alignment and Deep Embedding for Multi-Slice Spatial Transcriptomics
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- Limitations of Normalization in Attention
- MoBA: Mixture of Block Attention for Long-Context LLMs
- On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention Head
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
- Rope to Nope and Back Again: A New Hybrid Attention Strategy
- SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical Segmentation
- Sampling 3D Molecular Conformers with Diffusion Transformers
- Scale-invariant attention
- SparseMVC: Probing Cross-view Sparsity Variations for Multi-view Clustering
- Spectral Conditioning of Attention Improves Transformer Performance
- Strassen Attention, Split VC Dimension and Compositionality in Transformers
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Transformers Learn Faster with Semantic Focus
- UniteFormer: Unifying Node and Edge Modalities in Transformers for Vehicle Routing Problems
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- What are you sinking? A geometric approach on attention sink
- YOLOv12: Attention-Centric Real-Time Object Detectors