Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Yarin Gal (University of Oxford) · Yilun Zhao (Yale University) · Arman Cohan (Yale University) · Philip Torr (University of Oxford) · Jakob Foerster (University of Oxford) · Thomas Foster (University of OxfordFAIR @) · Ronald Clark (University of Oxford) · Adel Bibi (University of Oxford) · Cozmin Ududec (UK AI Security Institute) · Antoine Bosselut (EPFL) · Valentin Hofmann (Allen Institute for AI) · Andrew M. Bean (University of Oxford) · Ryan Othniel Kearns (University of Oxford) · Angelika Romanou (EPFL) · Franziska Sofia Hafner (University of Oxford) · Harry Mayne (University of Oxford) · Jan Batzner (TUM; Weizenbaum Institute) · Negar Foroutan Eghlidi (École polytechnique fédérale de Lausanne (EPFL)) · Chris Schmitz (Centre for Digital Governance, Hertie School) · Karolina Korgul (University of Oxford) · Hunar Batra (University of Oxford) · Oishi Deb (University of Oxford, DeepMind Scholar) · Emma Beharry (Stanford University) · Cornelius Emde (University of Oxford) · Anna Gausen (Imperial College London) · María Grandury (Universidad Politécnica de Madrid / SomosNLP) · Sophia Han (Stanford University) · Lujain Ibrahim (University of Oxford) · Hazel Kim (University of Oxford) · Hannah Rose Kirk (University of Oxford / UK AISI) · Fangru Lin (University of Oxford Google DeepMind) · Gabrielle Liu (Department of Computer Science, Yale University) · Lennart Luettgau (AI Security Institute) · Jabez Magomere (University of Oxford) · Jonathan Rystrøm (University of Oxford) · Anna Sotnikova (EPFL - EPF Lausanne) · Yushi Yang (University of Oxford) · Scott Hale (Meedan) · Deborah Raji (UC Berkeley) · Christopher Summerfield (University of Oxford) · Luc Rocher (University of Oxford) · Adam Mahdi (University of Oxford)
actionable guidanceconstruct validitydeployment assessmentexpert reviewersllm benchmarksmeasured phenomenanatural language processingresearch recommendationsrobustnesssafetyscoring metricssystematic reviewvalidity claims

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.