attention mechanism
An attention mechanism is a component in neural networks that allows models to focus on specific parts of the input data when making predictions. It selectively weighs the importance of different input elements, thus enhancing the model's ability to capture contextual relationships.
- Attention-based clustering
- CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
- Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
- Dependency Parsing is More Parameter-Efficient with Normalization
- DuetGraph: Coarse-to-Fine Knowledge Graph Reasoning with Dual-Pathway Global-Local Fusion
- Entropy Rectifying Guidance for Diffusion and Flow Models
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
- Knee-Deep in C-RASP: A Transformer Depth Hierarchy
- LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- Mamba Modulation: On the Length Generalization of Mamba Models
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- PaTH Attention: Position Encoding via Accumulating Householder Transformations
- Pinpointing Attention-Causal Communication in Language Models
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical Reshape
- Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity Modeling
- SAS: Simulated Attention Score
- SeerAttention: Self-distilled Attention Gating for Efficient Long-context Prefilling
- Spark Transformer: Reactivating Sparsity in Transformer FFN and Attention
- Titans: Learning to Memorize at Test Time
- Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds
- Transformer brain encoders explain human high-level visual responses
- Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms
- UMoE: Unifying Attention and FFN with Shared Experts
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- WKV-sharing embraced random shuffle RWKV high-order modeling for pan-sharpening
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding