multimodal inputs
Multimodal inputs involve the use of multiple forms of data (e.g., text, images, audio) to foster richer and more informative AI systems. The integration of these modalities enables more comprehensive understanding and enhanced performance on complex tasks.
- Benchmarking Retrieval-Augmented Multimomal Generation for Document Question Answering
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Multimodal LiDAR-Camera Novel View Synthesis with Unified Pose-free Neural Fields
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
- Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
- Understanding and Rectifying Safety Perception Distortion in VLMs