visual reasoning
The cognitive ability of systems to make inferences based on visual inputs, often crucial in tasks like interpreting diagrams or engaging in image-based question answering, where understanding spatial relationships and visual cues is vital.
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
- DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs
- GRIT: Teaching MLLMs to Think with Images
- Generative Perception of Shape and Material from Differential Motion
- Grounded Reinforcement Learning for Visual Reasoning
- LaViDa: A Large Diffusion Model for Vision-Language Understanding
- Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- VIKI‑R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
- Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs
- World-aware Planning Narratives Enhance Large Vision-Language Model Planner