alignment framework
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
- Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards
- Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment