Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Chengqian Gao (Mohamed bin Zayed University of Artificial Intelligence) · Mikhail Yurochkin (MBZUAI Institute of Foundation Models) · Eric Xing (CMU/MBZUAI/GenBio) · Abulhair Saparov (Purdue University) · Feng Yao (University of California, San Diego) · Zhiting Hu (University of California, San Diego) · Kun Zhou (University of California, San Diego) · Yuqi Wang (Petuum Inc.) · Jorge (Zhoujun) Cheng (UC San Diego) · Shibo Hao (University of California, San Diego) · Tianyang Liu (UC San Diego) · Fan Zhou (Beihang Univeristy) · Yutao Xie (University of California, San Diego) · Yuexin Bian (ucsd) · Nilabjo Dey (Purdue University) · Yonghao Zhuang (CMU, Carnegie Mellon University) · Yuheng Zha (University of California, San Diego || MBZUAI Institute of Foundation Models) · Yi Gu (University of California, San Diego) · Yuan Li (University of Cambridge) · Richard Fan (Petuum Inc.) · Jianshu She (Mohamed bin Zayed University of Artificial Intelligence) · Taylor W. Killian (MBZUAI Institute of Foundation Models) · Haonan Li (Mohamed bin Zayed University of Artificial Intelligence) · Zhengzhong Liu (MBZUAI Institute of Foundation Models)
cross-domain generalizationcross-domain transferabilitydata sourcingdata-curation pipelinededuplicationdifficulty-based filteringdomain-specific filteringevaluation suitelarge language modelmixed-domain trainingmodel performancemulti-domain datasetsreinforcement learningreward designrl reasoning models

Reinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, covering six reasoning domains: Math, Code, Science, Logic, Simulation, and Tabular, each with corresponding verifiers. We build \ours via a careful data-curation pipeline, including sourcing, deduplication, reward design, and domain-specific and difficulty-based filtering, to facilitate the systematic investigation of cross-domain RL generalization. Our study using \ours suggests the efficacy of a simple mixed-domain RL training approach and reveals several key aspects affecting cross-domain transferability. We further train two models {\ours}-7B and {\ours}-32B purely with RL on our curated data and observe largely improved performance over leading open RL reasoning model baselines, with gains of 7.3\% and 7.8\% respectively on an extensive 17-task, six-domain evaluation suite. We are releasing our dataset, code, and evaluation suite to the community, aiming to support further research and development of more general RL-enhanced reasoning models.