Zichen Wen
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation