Ran He
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- ZeroPatcher: Training-free Sampler for Video Inpainting and Editing