q-learning
- A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging
- Confounding Robust Deep Reinforcement Learning: A Causal Approach
- Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing
- Risk-Averse Total-Reward Reinforcement Learning
- ShiQ: Bringing back Bellman to LLMs
- Value Improved Actor Critic Algorithms