NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
off-policy algorithm
3 papers
Planning and Learning in Average Risk-aware MDPs
Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning