NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Yingjie Wang
3 papers
Nanyang Technological University
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Self-Verification Provably Prevents Model Collapse in Recursive Synthetic Training