Zhaopeng Tu
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
- The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
- Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training