Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Guang Chen (Tongji University) · Gangwei Xu (Huazhong University of Science and Technology) · Haiyang Sun · Bing Wang (Alibaba Group) · Hangjun Ye (Xiaomi Corporation) · Wenyu Liu (Huazhong University of Science and Technology) · Xinggang Wang (Huazhong University of Science and Technology) · Xiangyu Guo (Huazhong University of Science and Technology) · Zhanqian Wu (University of Pennsylvania) · Kaixin Xiong (Xiaomi Corporation) · Ziyang Xu (Huazhong University of Science and Technology) · Lijun Zhou (Xiaomi Corporation) · Shaoqing Xu (University of Macau)
3d-vae encodingadaptive samplingbev-represented lidar generatorcross-modal consistencydatacrafterdit-based video diffusion modelgenesisinstance-level captionslidar sequencesmulti-view driving videosnerf-based renderingnuscenes benchmarkscene-level captionsspatio-temporal consistencyunified world modelvision-language models

We present Genesis, a unified world model for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-represented LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared condition input, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level captions. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the synthetic data.