MURKA: Multi-Reward Reinforcement Learning with Knowledge Alignment for Optimization Tasks

Xiangyang Li (University of Science and Technology of China) · Feng Wu (The University of HongKong) · WANTONG XIE (University of Science and Technology of China) · Yi-Xiang Hu (University of Science and Technology of China) · Jieyang Xu (University of Science and Technology of China)
amplcheckercollaborative agent alignmentcomposite reward functionexecution fidelityextractorgeneralizabilitygroup relative policy optimizationknowledge distillationoperations researchoptimizationreinforcement learningsemantic correctnesssolver

Optimization plays a central role in Operations Research (OR) and numerous industrial applications, yet automating the end-to-end process of translating natural language descriptions into executable optimization programs remains a formidable challenge. While recent efforts have applied Large Language Models (LLMs) to this task, existing approaches are hindered by high inference costs, limited robustness across domains, and weak verification mechanisms. In this work, we propose MURKA, a reinforcement learning and knowledge distillation-based framework that enhances LLM-driven optimization modeling via collaborative agent alignment. MURKA orchestrates three specialized agents