safety constraints
Conditions imposed to ensure that AI systems operate within acceptable limits to prevent harmful outcomes, especially in critical applications such as autonomous vehicles.
- Automaton Constrained Q-Learning
- Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming Guardrails
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning
- Efficient Safe Meta-Reinforcement Learning: Provable Near-Optimality and Anytime Safety
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- Near-Optimal Sample Complexity for Online Constrained MDPs
- One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning
- Online Optimization for Offline Safe Reinforcement Learning
- Provable Gradient Editing of Deep Neural Networks
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models