The Catechol Benchmark: Time-series Solvent Selection Data for Few-shot Machine Learning

Toby Boyne (Imperial College London) · Juan Campos (Imperial College London) · Rebecca Langdon (Imperial College London) · Jixiang Qing (Imperial College London) · Yilin Xie (Imperial College London) · Shiqiang Zhang (Imperial College London) · Calvin Tsay (Imperial College London) · Ruth Misener (Imperial College London) · Daniel Davies (Imperial College London) · Kim Jelfs (Imperial College London) · Sarah Boyall (Imperial College London) · Thomas Dixon (Solve chemistry) · Linden Schrecker (Imperial College London) · Jose Pablo Folch (SOLVE Chemistry)
active learningbenchmarkingchemical datasetscontinuous process conditionsfeature engineeringmolecular property predictionreaction retro-synthesisregression algorithmssolvent replacementsolvent selectionsustainable manufacturingtransfer-learningtransient flow datasetyield prediction

Machine learning has promised to change the landscape of laboratory chemistry, with impressive results in molecular property prediction and reaction retro-synthesis. However, chemical datasets are often inaccessible to the machine learning community as they tend to require cleaning, thorough understanding of the chemistry, or are simply not available. In this paper, we introduce a novel dataset for yield prediction, providing the first-ever transient flow dataset for machine learning benchmarking, covering over 1200 process conditions. While previous datasets focus on discrete parameters, our experimental set-up allow us to sample a large number of continuous process conditions, generating new challenges for machine learning models. We focus on solvent selection, a task that is particularly difficult to model theoretically and therefore ripe for machine learning applications. We showcase benchmarking for regression algorithms, transfer-learning approaches, feature engineering, and active learning, with important applications towards solvent replacement and sustainable manufacturing.