NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
reference policy
4 papers
Doubly Robust Alignment for Large Language Models
On Extending Direct Preference Optimization to Accommodate Ties
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
Uni-RL: Unifying Online and Offline RL via Implicit Value Regularization