large vision-language models
Models that integrate visual and linguistic data to perform tasks that require understanding both modalities, such as image captioning and visual question answering. These models leverage large datasets of paired images and text.
- A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
- CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- Each Complexity Deserves a Pruning Policy
- EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction Evolution
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object Recognition
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- Latent Chain-of-Thought for Visual Reasoning
- MMCSBench: A Fine-Grained Benchmark for Large Vision-Language Models in Camouflage Scenes
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
- Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
- SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language Models
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-Tuning
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Visual Structures Help Visual Reasoning: Addressing the Binding Problem in LVLMs
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought