Meet TinyLlama: A Small AI Model that Aims to Pretrain a 1.1B Llama Model on 3 Trillion Tokens
small-modelspretrainingscalingefficiency
Abstraction: Small 1.1B LLM pre-trained on 3 trillion tokens
Key points:
- 1.1B parameter model trained on 3 trillion tokens in ~90 days using 16 A100-40G GPUs
- Model occupies only ~550MB of RAM, targeting deployment on single devices with limited compute
- Directly challenges the Chinchilla Scaling Law, which posits parameters and training tokens should scale proportionally
- Meta's Llama 2 showed no performance saturation after 2T tokens, motivating the 3T training target
- Open experiment with no predefined targets beyond "1.1B on 3T"; outcome would either validate or rebut Chinchilla
Connections: Tinyllama · Meta · Large Language Models · Scaling Laws · Model Efficiency