qwen2.5-vl
- GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing
- MLLM-ISU: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models based Intrusion Scene Understanding
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
- OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data
- Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models