Escaping Collapse: The Strength of Weak Data for Large Language Model Training

Kareem Amin (Google Research) · Sara Babakniya (Google) · Alex Bie (University of Waterloo) · Weiwei Kong (Google) · Umar Syed (Google Research) · Sergei Vassilvitskii (Google)
boostingchallenging examplesclassifiercurationdynamic focusempirical validationlabeling resourcesmachine learning techniqueperformance improvementperformance plateausynthetic datatheoretical frameworktraining iterationstraining methodsweak learning algorithm

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even "collapse", after many training iterations. In this paper, we formalize this question and develop a theoretical framework to investigate how much curation is needed in order to ensure that LLM performance continually improves. Our analysis is inspired by boosting, a classic machine learning technique that leverages a very weak learning algorithm to produce an arbitrarily good classifier. The approach we analyze subsumes many recently proposed methods for training LLMs on synthetic data, and thus our analysis sheds light on why they are successful, and also suggests opportunities for future improvement. We present experiments that validate our theory, and show that dynamically focusing labeling resources on the most challenging examples