CosmoBench: A Multiscale, Multiview, Multitask Cosmology Benchmark for Geometric Deep Learning

Teresa Huang (Johns Hopkins University) · Richard Stiskalek (University of Oxford) · Jun-Young Lee (Seoul National University) · Adrian Bayer (Princeton University / Simons Foundation) · Charles Margossian (Flatiron Institute) · Christian Kragh Jespersen (Princeton University (SSO)) · Lucia Perez (Princeton University) · Lawrence Saul (Flatiron Institute) · Francisco Villaescusa (Flatiron Institute)
benchmark datasetcollective positionscosmological parameterscosmological simulationsdark matter halosdirected treesgeometric deep learninggraph neural networksinvariant featuresleast-squares fitsmerger historypoint cloudsreconstruction tasksscientific discoveryvelocities prediction

Cosmological simulations provide a wealth of data in the form of point clouds and directed trees. A crucial goal is to extract insights from this data that shed light on the nature and composition of the Universe. In this paper we introduce CosmoBench, a benchmark dataset curated from state-of-the-art cosmological simulations whose runs required more than 41 million core-hours and generated over two petabytes of data. CosmoBench is the largest dataset of its kind: it contains 34 thousand point clouds from simulations of dark matter halos and galaxies at three different length scales, as well as 25 thousand directed trees that record the formation history of halos on two different time scales. The data in CosmoBench can be used for multiple tasks