preference optimization
Preference optimization is a process used to align AI models' outputs with human preferences or desired outcomes. It often involves collecting user feedback and iterating on model training to better serve the specific needs of users.
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
- Doubly Robust Alignment for Large Language Models
- Elastic Robust Unlearning of Specific Knowledge in Large Language Models
- InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
- Meta-Learning Objectives for Preference Optimization
- On Extending Direct Preference Optimization to Accommodate Ties
- On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language Understanding
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis
- Predictive Preference Learning from Human Interventions
- Protein Inverse Folding From Structure Feedback
- RePO: Understanding Preference Learning Through ReLU-Based Optimization
- Token-Level Self-Play with Importance-Aware Guidance for Large Language Models
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning