visual understanding
The capability of AI systems to interpret and derive meaning from visual information, such as images and videos. This encompasses tasks like object recognition, scene comprehension, and activity detection.
- Caption This, Reason That: VLMs Caught in the Middle
- ChemPile: A 250 GB Diverse and Curated Dataset for Chemical Foundation Models
- DAVE: Diagnostic benchmark for Audio Visual Evaluation
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- LMFusion: Adapting Pretrained Language Models for Multimodal Generation
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
- On Fairness of Unified Multimodal Large Language Model for Image Generation
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations