NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Zhiyuan Zeng
3 papers
University of Washington
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
Precise Information Control in Long-Form Text Generation
Reinforcement Learning for Reasoning in Large Language Models with One Training Example