NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
off-policy learning
4 papers
Bootstrap Off-policy with World Model
Checklists Are Better Than Reward Models For Aligning Language Models
Confounding Robust Deep Reinforcement Learning: A Causal Approach
ShiQ: Bringing back Bellman to LLMs