Knowledge Distillation
concepts · 2 notes linked
Related: Large Language Models · Google · Palm · T5 · Chain Of Thought · Bert · Roberta · Transformers
Notes
- Distilling step-by-step: Outperforming larger language models with less training — Rationale-based distillation trains small models to outperform 540B LLMs with far less data
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension — Hybrid convolutional-transformer model via knowledge distillation for fast NLU inference