Rongrong Ji
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
- CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models
- DAMamba: Vision State Space Model with Dynamic Adaptive Scan
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs