MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Wei Ji (National University of Singapore) · Ge Zhang (University of Michigan - Ann Arbor) · Pengfei Wan (Kuaishou Technology) · Xintao Wang (Applied Research Center, Tencent PCG) · Noah Wang · Zili Wang (stepfun) · Jian Yang (nanjing university) · ZHAO-XIANG ZHANG (Chinese Academy of Sciences, China) · Jiaheng Liu (Nanjing University) · Wenhao Huang (Key Laboratory of Machine Perception) · Tianhao Peng (Beijing University of Aeronautics and Astronautics) · Haochen Wang (Institute of automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences) · Yuanxing Zhang (Kuaishou- 快手科技) · Shihao Li (nanjing university) · Yanghai Wang (nanjing university) · Houyi Li (Fudan University)
autonomous drivingcomprehensive benchmarkcore competenciescross-angle sports analyticsevaluation benchmarkshigh-order reasoning tasksmulti-sensor synthesismulti-video understandingmultimodal large language modelsmvu-evalperception tasksperformance discrepanciesquestion-answer pairssports analyticsstate-of-the-art models

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce **MVU-Eval**, the first comprehensive benchmark for evaluating **M**ulti-**V**ideo **U**nderstanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research.