Jiaheng Liu
- Flow-GRPO: Training Flow Matching Models via Online RL
- KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- OmniBench: Towards The Future of Universal Omni-Language Models
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models