EconGym: A Scalable AI Testbed with Diverse Economic Tasks

Jun Wang (iWudao Tech) · Bo An (Nanyang Technological University) · Qirui Mi (Institute of Automation, Chinese Academy of Sciences) · Qipeng Yang (Nanjing University of Posts and Telecommunications) · Zijun Fan (Nanjing University of Posts and Telecommunications) · Wentian Fan (Nanjing University of Posts and Telecommunications) · Heyang Ma (University of International Business and Economics) · Chengdong Ma · Siyu Xia (Institute of automation, Chinese academy of science, Chinese Academy of Sciences) · Haifeng Zhang (Institute of automation, Chinese academy of science, Chinese Academy of Sciences)
agent modelsalgorithm diversitybenchmarkingeconomic modelingeconomic tasksfiscal policiesheterogeneous role typesinteraction mechanismsmonetary policiesmulti-agent interactionspolicy learningpolicy optimizationscalabilitysimulation platformstask composition

Artificial intelligence (AI) has become a powerful tool for economic research, enabling large-scale simulation and policy optimization. However, applying AI effectively requires simulation platforms for scalable training and evaluation—yet existing environments remain limited to simplified, narrowly scoped tasks, falling short of capturing complex economic challenges such as demographic shifts, multi-government coordination, and large-scale agent interactions.To address this gap, we introduce EconGym, a scalable and modular testbed that connects diverse economic tasks with AI algorithms. Grounded in rigorous economic modeling, EconGym implements 11 heterogeneous role types (e.g., households, firms, banks, governments), their interaction mechanisms, and agent models with well-defined observations, actions, and rewards. Users can flexibly compose economic roles with diverse agent algorithms to simulate rich multi-agent trajectories across 25+ economic tasks for AI-driven policy learning and analysis.Experiments show that EconGym supports diverse and cross-domain tasks—such as coordinating fiscal, pension, and monetary policies—and enables benchmarking across AI, economic methods, and hybrids. Results indicate that richer task composition and algorithm diversity expand the policy space, while AI agents guided by classical economic methods perform best in complex settings. EconGym also scales to 100k agents with high realism and efficiency.