monotonic improvement
- Active Target Discovery under Uninformative Priors: The Power of Permanent and Transient Memory
- Efficient Safe Meta-Reinforcement Learning: Provable Near-Optimality and Anytime Safety
- Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment
- Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm