preference dataset
- Aligning Compound AI Systems via System-level DPO
- Improving Video Generation with Human Feedback
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling