Xiawu Zheng
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs