Shunyu Liu
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
- SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
- Tree of Preferences for Diversified Recommendation
- VORTA: Efficient Video Diffusion via Routing Sparse Attention