token compression
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
- The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
- Vision-centric Token Compression in Large Language Model