Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language Models

Piotr Koniusz (Data61❤CSIRO) · Zejun MA (iscas) · Yew Soon Ong (Nanyang Technological University) · Junhao Dong (Nanyang Technological University / CFAR, A*STAR) · Xinghua Qu (Bytedance AI Lab) · Cong Zhang (Nanyang Technological University)
adversarial attacksadversarial robustnessalignment re-weightingclean samplesinformation source reliabilityknowledge alignmentlarge-scale modelsplug-and-play applicabilityrobustness generalizationsource guidance refinementsuperalignmentunsupervised schemevision-language benchmarksvision-language modelsweak-to-strong generalizationzero-shot robustness

Numerous well-established studies have demonstrated the superhuman capabilities of modern Vision-Language Models (VLMs) across a wide range of tasks. However, growing is the doubt about the continuing availability of reliable high-quality labeling (supervision) from human annotators, leading to stagnation of the model's performance. To address this challenge, ``superalignment'' employs the so-called weak-to-strong generalization paradigm, where the supervision from a weak model can provide generalizable knowledge for a strong model. While effective in aligning knowledge for clean samples between the strong and weak models, the standard weak-to-strong approach typically fails to capture adversarial robustness, exposing strong VLMs to adversarial attacks. This inability to transfer adversarial robustness is because adversarial samples are normally missing in the superalignment stage. To this end, we are the first to propose the weak-to-strong (adversarial) robustness generalization method to elicit zero-shot robustness in large-scale models by an unsupervised scheme, mitigating the unreliable information source for alignment from two perspectives: alignment re-weighting and source guidance refinement. We analyze settings under which robustness generalization is possible. Extensive experiments across various vision-language benchmarks validate the effectiveness of our method in numerous scenarios, demonstrating its plug-and-play applicability to large-scale VLMs.