NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
off-policy reinforcement learning
3 papers
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning