token budget
The limit on the number of tokens (words or subwords) that can be processed by natural language models, impacting model efficiency and the completeness of input data.
- ARM: Adaptive Reasoning Model
- Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
- Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization
- SkyLadder: Better and Faster Pretraining via Context Window Scheduling
- Thinker: Learning to Think Fast and Slow