out-of-distribution tasks
Challenges presented by data or situations that are significantly different from the training set, testing a model's generalization and robustness.
- Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python Code
- Learning to Reason under Off-Policy Guidance
- MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Reinforcement Learning Teachers of Test Time Scaling
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- Searching Latent Program Spaces
- Towards Provable Emergence of In-Context Reinforcement Learning