VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Chaoyou Fu (Nanjing University) · Haojia Lin (Xiamen University) · Xiong Wang (Hainan University) · yifan zhang (Institute of Automation, Chinese Academy of Sciences) · Yunhang Shen (Xiamen University) · Xiaoyu Liu (Shanghai University) · Haoyu Cao (Tencent Youtu Lab) · Zuwei Long (Tencent Youtu Lab) · Heting Gao (University of Illinois at Urbana-Champaign) · Ke Li (East China Normal University) · Long MA (Tencent Youtu Lab) · Xiawu Zheng (Xiamen University) · Rongrong Ji (Xiamen University, China) · Xing Sun (Tencent YouTu Lab) · Caifeng Shan (Nanjing University) · Ran He (NLPR, CASIA)
end-to-end response speedmulti-stage training methodologymultimodal dialogue systemsmultimodal large language modelsomni modelomni understandingspeech capabilitiesspeech interactionspeech-to-speech dialoguestate-of-the-art benchmarkstextual modalitiesvision and speech tasksvision-language capacityvisual capabilitiesvisual modalities

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing against state-of-the-art counterparts across benchmarks for image, video, and speech, we demonstrate that our omni model is equipped with both strong visual and speech capabilities, making omni understanding and interaction.