math benchmarks
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- Reasoning Is Not a Race: When Stopping Early Beats Going Deeper
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning