NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
policy distribution
3 papers
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
Preference Optimization by Estimating the Ratio of the Data Distribution
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models