NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
training tokens
3 papers
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
Zebra-Llama: Towards Extremely Efficient Hybrid Models