RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

Bo Jiang (Huazhong University of Science and Technology) · Yuechuan Pu (Horizon Robotics) · Hao Gao (Huazhong University of Science and Technology) · Shaoyu Chen (Huazhong University of Science and Technology) · Bencheng Liao (Huazhong University of Science and Technology) · Yiang Shi (Huazhong University of Science and Technology) · Xiaoyang Guo (Horizon Robotics) · haoran yin (Horizon Robotics) · Xiangyu Li (Horizon Robotics) · xinbang zhang (Horizon Robotics) · ying zhang (Horizon Robotics) · Wenyu Liu (Huazhong University of Science and Technology) · Qian Zhang (Horizon Robotics) · Xinggang Wang (Huazhong University of Science and Technology)
3dgsautonomous drivingcausal confusionclosed-loop frameworkcollision rateevaluation benchmarkhuman driving behaviorimitation learningout-of-distribution scenariosphotorealistic digital replicaregularization termreinforcement learningsafety-critical eventsstate space explorationtrial and error

Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3× lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https://github.com/hustvl/RAD for facilitating future research.