large multimodal models
Large multimodal models are AI systems designed to process and integrate information from multiple modalities, such as text, images, and audio. These models leverage vast amounts of data across different domains to improve performance and understanding.
- Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
- OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
- Seeing the Arrow of Time in Large Multimodal Models
- The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
- Training-free Online Video Step Grounding
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding