Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation

Kyunghyun Cho (Genentech / NYU) · Sungmin Cha (New York University)
distributional coveragedownstream generationempirical resultsentropy-controlling parametergenerative modelsgenerative qualityinstruction tuningknowledge distillationprecision-recall trade-offprobability masssample qualityselective distributionssimulationstudent modelteacher distribution

Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented