NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
policy gradient
3 papers
Offline Actor-Critic for Average Reward MDPs
On the Sample Complexity of Differentially Private Policy Optimization
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning