Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

Yan Lu (The Chinese University of Hong Kong) · SHIXIANG TANG (The Chinese University of Hong Kong) · LEI BAI (UNSW, Sydney) · Dongzhan Zhou (Shanghai Artificial Intelligence Laboratory) · Yuhao Zhou (Sichuan University) · Jiancheng Lv (Machine Intelligence Laboratory College of Computer Science, Sichuan University) · Hao Wu (Fudan University) · Yiheng Wang (Shanghai Jiao Tong University) · Xuming He (ShanghaiTech University) · Ruoyao Xiao (University of the Chinese Academy of Sciences) · Zhiwei Li (Tongji University) · Qiantai Feng (Shanghai Artificial Intelligence Laboratory) · Zijie Guo (Fudan University) · Yuejin Yang (Fudan University) · Wenxuan Huang (East China Normal University) · Jiaqi Wei (Zhejiang University) · Dan Si (Sichuan University) · YAO XIUQI (Sichuan University) · Jia Bu (East China Normal University) · Haiwen Huang (East China Normal University) · Tianfan Fu (Georgia Institute of Technology) · Ben Fei (Chinese University of Hong Kong) · Fenghua Ling (Nanjing University of Information Science and Technology) · Siqi Sun (Fudan University) · Chenhui Li (East China Normal University) · Guanjie Zheng (Shanghai Jiao Tong University) · Wenlong Zhang (Shanghai Artificial Intelligence Laboratory)
ai-enhanced discoveriescognitive capacitiesdomain-specific expertiseexpert-verified vqa pairsmultimodal large language modelsmultimodal reasoningmultimodal tasksperception abilitiesreasoning abilitiesscientific attribute understandingscientific benchmarksscientific comparative reasoningscientific datascientific signal perceptionscientists’ first exam

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this discovery process in realistic workflows. However, current scientific benchmarks mostly focus on evaluating the knowledge understanding capabilities of MLLMs, leading to an inadequate assessment of their perception and reasoning abilities. To address this gap, we present the Scientists’ First Exam (SFE) benchmark, designed to evaluate the scientific cognitive capacities of MLLMs through three interconnected levels: *scientific signal perception*, *scientific attribute understanding*, *scientific comparative reasoning*. Specifically, SFE comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. Extensive experiments reveal that current *state-of-the-art* GPT-o3 and InternVL-3 achieve only 34.08% and 26.52% on SFE, highlighting significant room for MLLMs to improve in scientific realms. We hope the insights obtained in SFE will facilitate further developments in AI-enhanced scientific discoveries.