NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
off-policy algorithms
3 papers
Preference Learning with Lie Detectors can Induce Honesty or Evasion
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
Value Improved Actor Critic Algorithms