math reasoning
- Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External Feedback
- LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions