visual grounding
The process of associating visual elements with linguistic expressions, enhancing the ability of models to understand and relate textual and visual information.
- 3EED: Ground Everything Everywhere in 3D
- CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object Recognition
- MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model