NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Donghai Hong
3 papers
Peking University
Generative RLHF-V: Learning Principles from Multi-modal Human Preference
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback