self-attention
Self-attention is an attention mechanism where the model weighs different parts of the input relative to each other, allowing it to focus on relevant information dynamically. This is vital in understanding context and dynamics within sequences of data, particularly in NLP tasks.
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
- Counterfactual reasoning: an analysis of in-context emergence
- Dependency Parsing is More Parameter-Efficient with Normalization
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Energy Landscape-Aware Vision Transformers: Layerwise Dynamics and Adaptive Task-Specific Training via Hopfield States
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- GSPN-2: Efficient Parallel Sequence Modeling
- GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image Completion
- Multi-head Temporal Latent Attention
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-training
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- REOrdering Patches Improves Vision Models
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
- Rethinking PCA Through Duality
- RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
- STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible Benchmarking
- Self-Supervised Learning of Graph Representations for Network Intrusion Detection
- Spiking Neural Networks Need High-Frequency Information
- TV-Rec: Time-Variant Convolutional Filter for Sequential Recommendation
- Text to Sketch Generation with Multi-Styles
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Two Heads are Better than One: Simulating Large Transformers with Small Ones
- Understanding Softmax Attention Layers:\\ Exact Mean-Field Analysis on a Toy Problem
- What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video Grounding