Yu Cheng
- Learning to Reason under Off-Policy Guidance
- Scaling Physical Reasoning with the PHYSICS Dataset
- Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models