NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
learning signals
3 papers
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
Token-Level Self-Play with Importance-Aware Guidance for Large Language Models
Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning