alignment performance
Alignment performance measures how well an AI system's outputs are in sync with human values, expectations, or objectives, crucial for ethical and safe AI deployment.
- Bridging Time and Linguistics: LLMs as Time Series Analyzer through Symbolization and Segmentation
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- Robust LLM Alignment via Distributionally Robust Direct Preference Optimization