token reduction
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- Language Models (Mostly) Know When to Stop Reading
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- Mitigating Overthinking in Large Reasoning Models via Manifold Steering
- Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs