MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

Zeliang Zhang (University of Rochester) · Chao Huang (Sun Yat-Sen University) · Susan Liang (University of Rochester) · Yunlong Tang (University of Rochester) · Chenliang Xu (University of Rochester) · Pinxin Liu (University of Rochester) · Mingqian Feng (University of Rochester) · Zhangyun Tan (University of Rochester) · Rui Mao (Shenzhen University) · Jing Bi (University of Rochester) · Yunzhong Xiao (Carnegie Mellon University) · Hang Hua (University of Rochester) · Ali Vosoughi (University of Rochester) · Luchuan Song (University of Rochester)
benchmark evaluationchain-of-thought promptingcompositional reasoningline relationship understandingmodel architecturemultimodal large language modelsperspective geometryperspective perceptionperspective type reasoningperspective-preserving transformationsreasoning tasksrobustness analysisspatial consistencyvanishing point perceptionvision-language systems

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark specifically designed to systematically evaluate MLLMs' understanding of perspective through 10 carefully crafted tasks across three complementary dimensions: Perspective Perception, Reasoning, and Robustness. Our benchmark comprises 2,711 real-world and synthetic image instances with 5,083 question-answer pairs that probe key capabilities, such as vanishing point perception and counting, perspective type reasoning, line relationship understanding in 3D space, invariance to perspective-preserving transformations, etc. Through a comprehensive evaluation of 43 state-of-the-art MLLMs, we uncover significant limitations: while models demonstrate competence on surface-level perceptual tasks, they struggle with compositional reasoning and maintaining spatial consistency under perturbations. Our analysis further reveals intriguing patterns between model architecture, scale, and perspective capabilities, highlighting both robustness bottlenecks and the benefits of chain-of-thought prompting. MMPerspective establishes a valuable testbed for diagnosing and advancing spatial understanding in vision-language systems. Resources are available at https://yunlong10.github.io/MMPerspective/