What’s in Common? Multimodal Models Hallucinate When Reasoning Across Scenes

Candace Ross (Meta) · Florian Bordes (Meta) · Adina Williams (Facebook AI Research) · Polina Kirichenko (New York University) · Mark Ibrahim (Fundamental AI Research (FAIR) at Meta)
benchmark evaluationcognitive testscommon-o benchcommon-o complexhallucinationsmulti-image inputsmultimodal language modelsobject co-occurrenceopen-vocabularyperception benchmarksreasoning across scenesreasoning modelsresearch challengescale improvementssingle imagestraining data contamination

Multimodal language models possess a remarkable ability to handle an open-vocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-O Bench with more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-O goes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking ``what’s in common?''. We evaluate leading multimodal language models, including models specifically trained to reason. We find that perceiving objects in single images is easy for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35\% on Common-O Bench