'If journalism is going up in smoke, I might as well get high off the fumes': confessions of a chatbot helper
data-annotationrlhftraining-datallmjournalismsynthetic-data
Abstraction: Human annotators writing gold-standard training data for LLMs
Key points:
- Approximately 20,000 people employed full-time creating annotated training data for LLMs (François Chollet estimate); without this work model output would be "really, really bad"
- Internet human text data projected to be exhausted relative to LLM training demand between 2026 and 2032 if current trends continue
- Training on synthetic (AI-generated) data causes "model collapse" — models lose awareness of rare/minority data and converge on most-likely outputs, per Shumailov et al. in Nature
- Annotation roles have shifted from low-paid global workers toward high-paid specialists (up to £30/hr UK); Scale AI (valued $14bn) is a major third-party annotator
- Irony: writers are paid to train models that may automate their non-annotation writing work; the same work improves the tools threatening their livelihoods
Connections: Openai · Chatgpt · Scale AI · Large Language Models · RLHF · Training Data