math reasoning benchmarks
- $Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
- Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
- Tapered Off-Policy REINFORCE - Stable and efficient reinforcement learning for large language models