JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

Tat-Seng Chua (National Univ. of Singapore) · Shuicheng Yan (National University of Singapore, Department of Electrical and Computer Engineering) · Jungang Li (The Hong Kong University of Science and Technology (GZ)) · Wei Zhang (Guangzhou University) · Hao Fei (National University of Singapore) · Jiayi Ji (Xiamen University) · Fan Zhou (Beihang Univeristy) · Kai Liu (SES AI) · Yuchong Sun (Renmin University of China) · Shengqiong Wu (National University of Singapore) · jianzhang gao (Renmin University of China) · Daoan Zhang (University of Rochester) · Sheng Jin (Nanyang Technological University) · Sicheng Yu (Singapore Management University) · Geng Zhan (University of Sydney, University of Sydney) · Liang Zheng (Australian National University)
audio-video fine-tuningcomplex synchronized settingscomprehension and generation benchmarksgpt-4o-curated dialoguesjavisgptjavisinst-omni datasetjoint audio-video comprehensionlarge-scale instruction-tuningmultimodal large language modelmultimodal pretrainingpretrained jav-dit generatorspatio-temporal audio-video fusionsyncfusion modulesynchrony-aware learnable queriestemporally coherent understandingthree-stage training pipeline

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for Joint Audio-Video (JAV) comprehension and generation. JavisGPT adopts a concise encoder–LLM–decoder architecture, featuring a SyncFusion module for spatio-temporal audio- video fusion and synchrony-aware learnable queries to bridge a pretrained JAV-DiT generator. This design enables temporally coherent video-audio understanding and generation from multimodal instructions. We design an effective three-stage training pipeline consisting of multimodal pretraining, audio-video fine-tuning, and large-scale instruction-tuning, to progressively build multimodal comprehension and generation from existing vision-language models. To support this, we further construct JavisInst-Omni, a high-quality instruction dataset with over 200K GPT-4o-curated audio-video-text dialogues that span diverse and multi-level comprehension and generation scenarios. Extensive experiments on JAV comprehension and generation benchmarks show that JavisGPT outperforms existing MLLMs, particularly in complex and temporally synchronized settings.