Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms

Philippe Wyder (University of Washington / Distyl AI) · Judah Goldfeder (Columbia University, Google) · Alexey Yermakov (University of Washington) · Yue Zhao (SURF) · Stefano Riva (Politecnico di Milano) · Jan Williams (University of Washington) · David Zoro (University of Washington) · Amy Rude (University of Washington) · Matteo Tomasetto (Polytechnic Institute of Milan) · Joe Germany (American University of Beirut) · Joseph Bakarji · Georg Maierhofer (University of Cambridge) · Miles Cranmer (University of Cambridge) · Nathan Kutz (University of Washington)
algorithm evaluationcommon task frameworkforecastinggeneralizationhidden test setskuramoto-sivashinskylimited datalorenznoisenonlinear systemsreporting biasreproducibilityscientific mlstandardized benchmarksstate reconstruction

Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks—leading to weak baselines, reporting bias, and inconsistent evaluations across methods. This undermines reproducibility, misguides resource allocation, and obscures scientific progress. To address this, we propose a Common Task Framework (CTF) for scientific machine learning. The CTF features a curated set of datasets and task-specific metrics spanning forecasting, state reconstruction, and generalization under realistic constraints, including noise and limited data. Inspired by the success of CTFs in fields like natural language processing and computer vision, our framework provides a structured, rigorous foundation for head-to-head evaluation of diverse algorithms. As a first step, we benchmark methods on two canonical nonlinear systems: Kuramoto-Sivashinsky and Lorenz. These results illustrate the utility of the CTF in revealing method strengths, limitations, and suitability for specific classes of problems and diverse objectives. Next, we are launching a competition around a global real world sea surface temperature dataset with a true holdout dataset to foster community engagement. Our long-term vision is to replace ad hoc comparisons with standardized evaluations on hidden test sets that raise the bar for rigor and reproducibility in scientific ML.