Multi-Agent Collaboration via Evolving Orchestration

Ye Tian (Peking University) · Zhiyuan Liu (Tsinghua University) · Maosong Sun (Tsinghua University, Tsinghua University) · Xiaoyin Che (Siemens AG) · Chen Qian (Shanghai Jiaotong University) · Xuantang Xiong (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Lei Han (Tencent AI Lab) · Weize Chen (Tsinghua University, Tsinghua University) · Cheng Yang (Alibaba Group) · Yufan Dang (Tsinghua University) · Xueheng Luo (Tsinghua University, Tsinghua University) · Jingru Fan (Shanghai Jiaotong University) · Zihao Xie (Tsinghua University, Tsinghua University) · Ruijie Shi (Tsinghua University)
adaptive sequencingcentralized orchestratorclosed-domain scenarioscomputational efficiencycoordination overheadcyclic reasoning structuresevolvable reasoningflexible reasoningmulti-agent collaborationopen-domain scenariosorganizational structuresperformance optimizationpuppeteer-style paradigmreinforcement learning

Large language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organizational structures that struggle to adapt as task complexity and agent numbers grow, resulting in coordination overhead and inefficiencies. To this end, we propose a puppeteer-style paradigm for LLM-based multi-agent collaboration, where a centralized orchestrator ("puppeteer") dynamically directs agents ("puppets") in response to evolving task states. This orchestrator is trained via reinforcement learning to adaptively sequence and prioritize agents, enabling flexible and evolvable collective reasoning. Experiments on closed- and open-domain scenarios show that this method achieves superior performance with reduced computational costs. Analyses further reveal that the key improvements consistently stem from the emergence of more compact, cyclic reasoning structures under the orchestrator’s evolution. Our code is available at https://github.com/OpenBMB/ChatDev/tree/puppeteer.