offline learning
- Contextual Thompson Sampling via Generation of Missing Data
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
- Modelling the control of offline processing with reinforcement learning
- ShiQ: Bringing back Bellman to LLMs
- Tapered Off-Policy REINFORCE - Stable and efficient reinforcement learning for large language models
- Uni-RL: Unifying Online and Offline RL via Implicit Value Regularization