transformer layers
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding