Brian Bartoldson
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training