vision-language reasoning
Vision-language reasoning involves understanding relationships between visual data (such as images) and language (text descriptions). AI models that can integrate and reason about both types of information enable applications like image captioning and visual question answering.
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards