Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked Systems

Hao Liang (King's College London) · shuqing shi (Kings College London) · Yudi Zhang (Eindhoven University of Technology) · Biwei Huang (University of California, San Diego) · Yali Du (King‘s College London)
$\kappa$-hop neighborhoodsactor-critic convergenceadaptation gapapproximately compact representationscausal recoverycausal representation learningconventional adaptation baselinesdomain generalizationfinite-sample guaranteeslarge-scale networked systemslearning-from-scratchmeta actor-critic learningreinforcement learningshared policy trainingsparse local causal maskvalue function truncation

Large‑scale networked systems, such as traffic, power, and wireless grids, challenge reinforcement‑learning agents with both scale and environment shifts. To address these challenges, we propose \texttt{GSAC} (\textbf{G}eneralizable and \textbf{S}calable \textbf{A}ctor‑\textbf{C}ritic), a framework that couples causal representation learning with meta actor‑critic learning to achieve both scalability and domain generalization. Each agent first learns a sparse local causal mask that provably identifies the minimal neighborhood variables influencing its dynamics, yielding exponentially tight approximately compact representations (ACRs) of state and domain factors. These ACRs bound the error of truncating value functions to $\kappa$-hop neighborhoods, enabling efficient learning on graphs. A meta actor‑critic then trains a shared policy across multiple source domains while conditioning on the compact domain factors; at test time, a few trajectories suffice to estimate the new domain factor and deploy the adapted policy. We establish finite‑sample guarantees on causal recovery, actor-critic convergence, and adaptation gap, and show that \texttt{GSAC} adapts rapidly and significantly outperforms learning-from-scratch and conventional adaptation baselines.