SMARTraj$^2$: A Stable Multi-City Adaptive Method for Multi-View Spatio-Temporal Trajectory Representation Learning

Gao Cong (Nanyang Technological University) · Fei Wang (Google) · Zezhi Shao (Institute of Computing Technology, Chinese Academy of Sciences) · Yongjun Xu (Institute of Computing Technology, Chinese Academy of Sciences) · Jun Zhang (Guizhou University) · Tao Sun (Stanford University) · Tangwen Qian (Institute of Computing Technology, Chinese Academy of Sciences) · Junhe Li (University of the Chinese Academy of Sciences) · Yile Chen (Nanyang Technological University)
amplified seesaw phenomenonbenchmark datasetscross-city generalizationdomain-invariant featuresdomain-specific featuresdownstream tasksfeature disentanglementmulti-view approachespersonalized gating mechanismreal-world applicabilityrobust performancespatio-temporal trajectory representationstructural heterogeneitytrajectory learning modelsurban applications

Spatio-temporal trajectory representation learning plays a crucial role in various urban applications such as transportation systems, urban planning, and environmental monitoring. Existing methods can be divided into single-view and multi-view approaches, with the latter offering richer representations by integrating multiple sources of spatio-temporal data. However, these methods often struggle to generalize across diverse urban scenes due to multi-city structural heterogeneity, which arises from the disparities in road networks, grid layouts, and traffic regulations across cities, and the amplified seesaw phenomenon, where optimizing for one city, view, or task can degrade performance in others. These challenges hinder the deployment of trajectory learning models across multiple cities, limiting their real-world applicability. In this work, we propose SMARTraj$^2$, a novel stable multi-city adaptive method for multi-view spatio-temporal trajectory representation learning. Specifically, we introduce a feature disentanglement module to separate domain-invariant and domain-specific features, and a personalized gating mechanism to dynamically stabilize the contributions of different views and tasks. Our approach achieves superior generalization across heterogeneous urban scenes while maintaining robust performance across multiple downstream tasks. Extensive experiments on benchmark datasets demonstrate the effectiveness of SMARTraj$^2$ in enhancing cross-city generalization and outperforming state-of-the-art methods. See our project website at \url{https://github.com/GestaltCogTeam/SMARTraj}.