A cartel of influential datasets are dominating machine learning research
ml-benchmarksdatasetsreproducibilitybiasresearch-methodology
Abstraction: ML benchmark dataset concentration at elite institutions fragments and biases research progress
Key points:
- Paper "Reduced, Reused and Recycled" (2021): 50%+ of all dataset usages in a 43,140-paper sample correspond to datasets from 12 elite, primarily Western institutions; Gini concentration exceeded 0.80 in recent years
- "Benchmark Lottery" paper: relative performance of algorithms may be altered significantly simply by choosing different benchmark tasks, highlighting fragility of current evaluation paradigms
- Computer vision research more affected than NLP; ImageNet-era datasets increasingly considered toy benchmarks given GPT-3 and large web-scale datasets
- Recht et al. showed models slightly overfit to test partitions of ImageNet and CIFAR-10; rotating examples between splits caused up to 10-15% performance drop
- Geoff Hinton critique: publication norms requiring benchmark tables discourage radically new ideas; junior reviewers reject anything that "makes the brain hurt"
- Community counter-argument: standard benchmarks enable fair comparison; concentration is a sign of maturity, not conspiracy — cited datasets are openly available to all
Connections: Imagenet · Papers With Code · ML Benchmarks · Dataset Bias · Reproducibility