NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
average reward
3 papers
Direct Alignment with Heterogeneous Preferences
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference