InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

Yaodong Yang (AIG) · Sirui Han (The Hong Kong University of Science and Technology) · Boyuan Chen (Peking University) · Donghai Hong (Peking University) · Jiaming Ji (Peking University) · Jiacheng Zheng (Shandong University) · Bowen Dong (Harbin Institute of Technology) · Jiayi Zhou (Peking University) · Kaile Wang (Peking University) · Juntao Dai (Peking University) · Xuyao Wang (Nankai University) · wenqi chen (University of Electronic Science and Technology of China / Institute for AI, Peking University) · Qirui Zheng (Peking University) · Wenxin Li (Peking University) · Yike Guo (Imperial College London)
agentic workflowexpert annotationshuman feedbackhuman-labeled preference pairsinterleaved contextsintermt-benchjudge moderationmulti-turn interactionmulti-turn qa instancesmulti-turn scaling lawmultimodal generationmultimodal large modelsmultimodal understandingpreference datasettool-augmented mllms

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step.