ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models

Juan Rodriguez (Mila - Quebec Artificial Intelligence Institute) · Greg Durrett (University of Texas at Austin) · Wenxuan Ding (New York University) · Liyan Tang (University of Texas, Austin) · Grace Kim (University of Pennsylvania, University of Pennsylvania) · Xinyu Zhao (Massachusetts Institute of Technology) · Thom Lake (UT Austin, Indeed) · Fangcong Yin (The University of Texas at Austin) · Prasann Singhal (University of California, Berkeley) · Manya Wadhwa (University of Texas at Austin) · Zeyu Liu (University of Texas at Austin) · Zayne Sprague (University of Texas at Austin) · Ramya Namuduri (University of Texas at Austin) · Bodun Hu (University of Texas at Austin) · Puyuan Peng (University of Texas at Austin)
benchmark evaluationchart question answeringchart understandingexpert-annotated questionshuman performancelarge vision-language modelsmodel capabilitiesperformance degradationperformance gapqualitative error analysisreasoning typessynthetic datasettextual reasoningvisual complexityvisual reasoning

Chart understanding presents a unique challenge for large vision-language models (LVLMs), as it requires the integration of sophisticated textual and visual reasoning capabilities. However, current LVLMs exhibit a notable imbalance between these skills, falling short on visual reasoning that is difficult to perform in text. We conduct a case study using a synthetic dataset solvable only through visual reasoning and show that model performance degrades significantly with increasing visual complexity, while human performance remains robust. We then introduce *ChartMuseum*, a new Chart Question Answering (QA) benchmark containing 1,162 expert-annotated questions spanning multiple reasoning types, curated from real-world charts across 184 sources, specifically built to evaluate complex visual and textual reasoning. Unlike prior chart understanding benchmarks