constrained markov decision process
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
- Fuz-RL: A Fuzzy-Guided Robust Framework for Safe Reinforcement Learning under Uncertainty
- Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning