attention maps
Visual representations that illustrate the importance of different regions in the input data as determined by attention mechanisms, aiding in the interpretability of model predictions.
- Backdoor Cleaning without External Guidance in MLLM Fine-tuning
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
- REPA Works Until It Doesn’t: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTs
- SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion Transformers
- SpikingVTG: A Spiking Detection Transformer for Video Temporal Grounding
- What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers