multimodal understanding
Multimodal understanding in AI involves integrating and processing information from multiple types of data sources (e.g., text, images, audio) to enhance model comprehension and context-awareness, enabling richer interactions.
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio–Language Models
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- End-to-End Vision Tokenizer Tuning
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- HMVLM:Human Motion-Vision-Language Model via MoE LoRA
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- MMaDA: Multimodal Large Diffusion Language Models
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention Head
- Show-o2: Improved Native Unified Multimodal Models
- Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization