mathematical reasoning benchmarks
Mathematical reasoning benchmarks are standardized tests designed to assess a model's ability to perform mathematical reasoning tasks, evaluating capabilities like problem-solving and logical deduction.
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression
- Diversity-Aware Policy Optimization for Large Language Model Reasoning
- Don’t Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
- Eliciting Reasoning in Language Models with Cognitive Tools
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
- Incentivizing LLMs to Self-Verify Their Answers
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
- Know What You Don't Know: Uncertainty Calibration of Process Reward Models
- Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
- RAST: Reasoning Activation in LLMs via Small-model Transfer
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards