Distilling step-by-step: Outperforming larger language models with less training

knowledge-distillationllmchain-of-thoughtmodel-efficiency

Abstraction: Rationale-based distillation trains small models to outperform 540B LLMs with far less data

Key points:

Connections: Google · Palm · T5 · Knowledge Distillation · Chain Of Thought · Large Language Models

Source: https://blog.research.google/2023/09/distilling-step-by-step-outperforming.html?m=1