Xing Sun
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs