MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

Haoxuan Li (Peking University) · Zhouchen Lin (Peking University) · Chaoyou Fu (Nanjing University) · yifan zhang (Institute of Automation, Chinese Academy of Sciences) · Xinfeng Li (Nanyang Technological University) · Liang Wang (NLPR, China) · Pengfei Wan (Kuaishou Technology) · Bohan Zeng (Peking University) · Yuanxing Zhang (Kuaishou- 快手科技) · Zhuoran Zhang · Wenjing Yang (National University of Defense Technology) · Haotian Wang (National University of Defense Technology) · Sihan Yang (Xi'an Jiaotong University) · Zhang Zhang (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Yang Shi (Peking University) · Huanqian Wang (Tsinghua University, Tsinghua University) · Xie · Huanyao Zhang (Peking University) · Lijie Zhao (The Chinese University of Hong Kong) · Zhuoer Wen (Peking University) · Wenting Liu (Peking University) · Xinlong Chen (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Yushuo Guan (Kuaishou- 快手科技)
cross-frame information integrationdynamic video scenarioshigh-resolution visual inputholistic video comprehensionlanguage prior biasmme-videoocr benchmarkmotion blurmultimodal large language modelsoptical character recognitionspatio-temporal reasoningtask categoriestemporal coveragetemporal variationsvideo ocrvisual effects

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.