mllms
Multimodal large language models that process and generate content across different modalities (e.g., text, image, audio). These models enable richer interactions and understanding by integrating multiple forms of data.
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
- GRIT: Teaching MLLMs to Think with Images
- MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
- MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
- PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
- UniTok: a Unified Tokenizer for Visual Generation and Understanding
- Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering