SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Yang Gao (Nanjing University) · Hao Wang (City University of Hong Kong) · Dayiheng Liu (Alibaba Group) · Ge Zhang (University of Michigan - Ann Arbor) · Rui Li (Rochester Institute of Technology) · Bingli Wang (Sichuan Agricultural University) · Meng Cao (Mohamed bin Zayed University of Artificial Intelligence) · Qian Liu (TikTok (Singapore)) · Yiming Liang (University of the Chinese Academy of Sciences) · Xeron Du (01.AI) · Yifan Yao (Beijing University of Posts and Telecommunications) · Kaijing Ma (Tongji University) · Tianyu Zheng (Beijing University of Posts and Telecommunications) · Zhu (Guangdong OPPO Mobile Telecommunications Corp.,Ltd.) · Minghao Liu (2077AI) · Xiaolong Jin (Purdue University) · Zhenlin Wei (Harbin Engineering University) · Chujie Zheng (Tsinghua University) · Kaixin Deng (Hokkaido University) · Shuyue Guo (Beijing University of Posts and Telecommunications) · Shian Jia (Zhejiang University) · Sichao Jiang (zhejiang university) · Yiyan Liao (Peking University) · Qinrui Li (Cornell University) · Sirun Li (Peking University) · Yizhi Li (The University of Manchester) · Yunwen Li (Chinese University of Hong Kong(shenzhen)) · Dehua Ma (Beijing University of Posts and Telecommunications) · Yuansheng Ni (University of Waterloo) · Haoran Que (Beijing University of Aeronautics and Astronautics) · Qiyao Wang (henzhen Institute of Advanced Technology, Chinese Academy of Sciences) · Zhoufutu Wen (ByteDance Inc.) · Siwei Wu (Nanjing University of Science and Technology) · Tianshun Xing (Beijing University of Posts and Telecommunications) · 明 许 (01.AI) · Zhenzhu Yang (China University of Geoscience Beijing) · Noah Wang · Junting Zhou (Peking University) · yuelin bai (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Chinese Academy of Sciences) · Xingyuan Bu (Alibaba Group) · chenglin cai (Huawei Technologies Ltd.) · Liang Chen (Peking University) · Yifan Chen (UCLA) · Cheng Chengtuo (Zhejiang University) · Tianhao Cheng (Fudan University) · Keyi Ding (2077AI) · Siming Huang (University of Melbourne) · HUANG YUN (national university of singaore, National University of Singapore) · Yaoru Li (Zhejiang University) · Yizhe Li (Zhejiang University) · Zhaoqun Li (Zhejiang University) · Tianhao Liang (Zhejiang University) · Chengdong Lin (Hangzhou Dianzi University) · Hongquan Lin (University of Science and Technology of China) · Yinghao Ma (Centre for Digital Music, Queen Mary University of London) · Zhongyuan Peng (Fudan University) · Zifan Peng (The Hong Kong University of Science and Technology (Guangzhou)) · Qige Qi (ByteDance Inc.) · Shi Qiu (Peking University) · Xingwei Qu (University of Manchester) · Shanghaoran Quan (Alibaba Group) · Yizhou Tan (Harvard University) · Zili Wang (stepfun) · 王晨清 (abaka) · Yiya Wang (Peking University) · Yubo Wang (University of Waterloo) · Jiajun Xu (Facebook) · Kexin Yang (Alibaba Group) · Ruibin Yuan (Carnegie Mellon University) · Yuanhao Yue (Fudan University) · Tianyang Zhan (ByteDance Inc.) · Chun Zhang (ByteDance Inc.) · Jinyang Zhang (Peking University) · Xiyue Zhang (Peking University) · Owen Zhang (Department of Computer Science, Princeton University) · Yue Zhang (Suzhou University) · Yongchi Zhao (Alibaba Group) · Xiangyu Zheng (Fudan University) · ChenghuaZhong (University of Science and Technology Beijing) · Zhoujun Li (Beijing University of Aeronautics and Astronautics) · Tianyu Liu (Alibaba) · Shiwen Ni (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences) · Junran Peng (Institute of automation, Chinese academy of science) · Yujia Qin (Bytedance) · Wenbo Su (Alibaba Group) · Guoyin Wang (Alibaba Qwen Pilot) · Shi Wang (Institute of Computing Science, Chinese Academy of Sciences) · Jian Yang (nanjing university) · Min Yang (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Chinese Academy of Sciences) · Xiang Yue (Carnegie Mellon University) · ZHAO-XIANG ZHANG (Chinese Academy of Sciences, China) · Wangchunshu Zhou (Guangdong OPPO Mobile Telecommunications Corp.,Ltd.) · Jiaheng Liu (Nanjing University) · Qunshu Lin (Abaka AI) · Wenhao Huang (Key Laboratory of Machine Perception)
artificial general intelligencecollaborative filteringdomain-specific benchmarksexpert feedbackhuman-llm collaborationiterative refinementknowledge evaluationlarge-scale annotationmethodological guidanceperformance improvementreasoning capabilitiesspecialized disciplinesstate-of-the-art modelssupergpqa

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs in many of these specialized fields-particularly in light industry, agriculture, and service-oriented disciplines-remain inadequately evaluated. To address this gap, we present SuperGPQA, a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines. Our benchmark employs a novel Human-LLM collaborative filtering mechanism to eliminate trivial or ambiguous questions through iterative refinement based on both LLM responses and expert feedback. Our experimental results reveal significant room for improvement in the performance of current state-of-the-art LLMs across diverse knowledge domains (e.g., the reasoning-focused model Gemini-2.5-Pro achieved the highest accuracy of 63.56% on SuperGPQA), highlighting the considerable gap between current model capabilities and artificial general intelligence. Additionally, we present comprehensive insights from our management of a large-scale annotation process, involving over 80 expert annotators and an interactive Human-LLM collaborative system, offering valuable methodological guidance for future research initiatives of comparable scope.