MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

Jie Liu (City University of Hong Kong) · Haoliang Li (NTU, Singapore) · Wenxuan Wang (Xidian University) · Wenting Chen (City University of Hong Kong) · Michael R Lyu (CUHK) · Zizhan Ma (The Chinese University of Hong Kong) · Guolin Huang (Shenzhen University) · Yihang SU (The Chinese University of Hong Kong) · Kao-Jung Chang (National Yang Ming Chiao Tung University) · Linlin Shen (Shenzhen University)
adaptabilityai systemsclinical decision makingclinical workflowfeedback mechanisminformation gatheringinteractivitylarge language modelmedcase-ragmedchainmedchain-agentpersonalizationreal-world scenariossequential clinical taskssequentiality

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper.