Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
distributional coveragedownstream generationempirical resultsentropy-controlling parametergenerative modelsgenerative qualityinstruction tuningknowledge distillationprecision-recall trade-offprobability masssample qualityselective distributionssimulationstudent modelteacher distribution
Knowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented