safety alignment
The process of ensuring that AI systems behave in ways that are aligned with human values and safety requirements. This is crucial to prevent unintended consequences and ensure ethical AI deployment.
- AdvPrefix: An Objective for Nuanced LLM Jailbreaks
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming Guardrails
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- LLM Safety Alignment is Divergence Estimation in Disguise
- Lifelong Safety Alignment for Language Models
- Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- Safety Depth in Large Language Models: A Markov Chain Perspective
- Shape it Up! Restoring LLM Safety during Finetuning
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
- Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models
- Understanding and Rectifying Safety Perception Distortion in VLMs