Towards Pre-trained Graph Condensation via Optimal Transport

Yao Zhao (Beijing Jiaotong University) · Yeyu Yan (Beijing Jiaotong University) · Shuai Zheng (GM Cruise) · Wenjun Hui (Beijing Jiaotong University) · Xiangkai Zhu (Shandong University of Science and Technology) · Chen Dong (Beijing Jiaotong university) · Zhenfeng Zhu (Beijing Jiaotong University) · Kunlun He (Chinese PLA General Hospital)
architecture independencecondensed graphsdownstream tasksgeneralized optimization objectivegnn optimizationgraph condensationhybrid-interval graph diffusionoptimal transportpre-trained graph condensationrepresentation transport plansemantic consistenciessemantic harmonizertask dependenciesuncertainty enhancementversatility

Graph condensation (GC) aims to distill the original graph into a small-scale graph, mitigating redundancy and accelerating GNN training. However, conventional GC approaches heavily rely on rigid GNNs and task-specific supervision. Such a dependency severely restricts their reusability and generalization across various tasks and architectures. In this work, we revisit the goal of ideal GC from the perspective of GNN optimization consistency, and then a generalized GC optimization objective is derived, by which those traditional GC methods can be viewed nicely as special cases of this optimization paradigm. Based on this, \textbf{Pre}-trained \textbf{G}raph \textbf{C}ondensation (\textbf{PreGC}) via optimal transport is proposed to transcend the limitations of task- and architecture-dependent GC methods. Specifically, a hybrid-interval graph diffusion augmentation is presented to suppress the weak generalization ability of the condensed graph on particular architectures by enhancing the uncertainty of node states. Meanwhile, the matching between optimal graph transport plan and representation transport plan is tactfully established to maintain semantic consistencies across source graph and condensed graph spaces, thereby freeing graph condensation from task dependencies. To further facilitate the adaptation of condensed graphs to various downstream tasks, a traceable semantic harmonizer from source nodes to condensed nodes is proposed to bridge semantic associations through the optimized representation transport plan in pre-training. Extensive experiments verify the superiority and versatility of PreGC, demonstrating its task-independent nature and seamless compatibility with arbitrary GNNs.