ChemPile: A 250 GB Diverse and Curated Dataset for Chemical Foundation Models

Marianna Nezhurina (LAION, JSC) · Michael Pieler (Edison Scientific (prev. FutureHouse)) · Adrian Mirza (Helmholtz-Zentrum Berlin für Materialien und Energie) · Nawaf Alampara (Friedrich-Schiller Universität Jena) · Martiño Ríos-García (FSU-Jena) · Mohamed Abdelalim (Independent Researcher) · Jack Butler (Amazon) · Bethany Connolly (Faculty) · Tunca Dogan (Hacettepe University) · Bünyamin Şen (Hacettepe University) · Santosh Tirunagari (Middlesex University) · Mark Worrall (Faculty Science Ltd) · Adamo Young (University of Toronto) · Philippe Schwaller (Swiss Federal Institute of Technology Lausanne (EPFL)) · Kevin Maik Jablonka (FSU Jena)
advanced reasoningbenchmarkingchemical aichemical representationschempileeducational foundationsfoundation modelshuggingfaceinchiiupac nameslarge-scale datasetsmolecular renderingsselfiessmilesspecialized expertisevisual understanding

Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality datasets that reflect the field's multifaceted nature. We present the ChemPile, an open dataset containing over 75 billion tokens of curated chemical data, specifically built for training and evaluating general-purpose models in the chemical sciences. The dataset mirrors the human learning journey through chemistry