MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Yuxuan Wang (ByteDance) · Xinfeng Li (Nanyang Technological University) · Zheng Lian (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Yanqiao Zhu (Shanghai Jiaotong University) · Minghao Liu (2077AI) · Yinghao Ma (Centre for Digital Music, Queen Mary University of London) · Ruibin Yuan (Carnegie Mellon University) · Kai Yu (Shanghai Jiao Tong University) · Zhuo Chen (ByteDance Inc.) · Wei Xue (Department of C.S., Tsinghua University) · Chen Yang (Shanghai Jiaotong University) · Jianwei Yu (Microsoft) · Tianrui Wang (Tianjin University) · Ziyang Ma (Shanghai Jiao Tong University) · Guanrou Yang (Shanghai Jiaotong University) · Eng-Siong Chng (Nanyang Technological University) · Xie Chen (Shanghai Jiaotong University) · Haina Zhu (Shanghai Jiaotong University) · Kai Li (Tsinghua University, Tsinghua University) · Emmanouil Benetos (Queen Mary University of London) · Yi-Wen Chao (Nanyang Technological University) · Ruiyang Xu (Shanghai Jiaotong University) · Wenxi Chen (Shanghai Jiaotong University) · Yuanzhe Chen (ByteDance Inc.) · Jian Cong (ByteDance Inc.) · Keliang Li (, Chinese Academy of Sciences) · Siyou Li (Queen Mary University of London) · Xiquan Li (Shanghai Jiaotong University) · Yuzhe Liang (Shanghai Jiaotong University) · Zhikang Niu (Shanghai Jiaotong University) · Wang Yuping (University of Science and Technology of China) · Yihao Wu (Nanyang Technological University) · Zhisheng Zheng (University of Texas at Austin) · Ziya Zhou (Hong Kong University of Science and Technology)
algorithm innovationaudio caption inputsaudio-language modelsaudio-question-answer tripletschain-of-thought rationaledeep reasoningdomain-specific knowledgelarge audio reasoning modelslarge audio-language modelsmixed-modalitymmarmulti-disciplinary tasksmulti-step reasoningperceptual knowledgereasoning layersunderstanding capabilities

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.