Yaodong Yang
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation
- Empirical Study on Robustness and Resilience in Cooperative Multi-Agent Reinforcement Learning
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
- Social World Model-Augmented Mechanism Design Policy Learning
- World Models Should Prioritize the Unification of Physical and Social Dynamics