alignment
The process of ensuring that an AI model's goals and behaviors are consistent with human values or objectives, crucial for ethical AI deployment.
- Amortized Active Generation of Pareto Sets
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMs
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- Composition and Alignment of Diffusion Models using Constrained Learning
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
- Dynamic Shadow Unveils Invisible Semantics for Video Outpainting
- Enhancing 3D Reconstruction for Dynamic Scenes
- Explainable Reinforcement Learning from Human Feedback to Improve Alignment
- Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
- Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
- Learning Preferences without Interaction for Cooperative AI: A Hybrid Offline-Online Approach
- Learning to Learn with Contrastive Meta-Objective
- Learning to Learn with Contrastive Meta-Objective
- Look-Ahead Reasoning on Learning Platforms
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video Preference
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- On Linear Mode Connectivity of Mixture-of-Experts Architectures
- Position: Towards Bidirectional Human-AI Alignment
- REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
- Safety Depth in Large Language Models: A Markov Chain Perspective
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- Understanding the Evolution of the Neural Tangent Kernel at the Edge of Stability
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
- VPO: Reasoning Preferences Optimization Based on $\mathcal{V}$-Usable Information
- When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product