Lie Detector: Unified Backdoor Detection via Cross-Examination Framework

Xuan Wang (Northwest Polytechnical University Xi'an) · Siyuan Liang (National University of Singapore) · Dongping Liao (University of Macau) · Han Fang (National University of Singapore) · Aishan Liu (Beihang University) · Xiaochun Cao (SUN YAT-SEN UNIVERSITY) · Yu-liang Lu (National University of Defense Technology) · Ee-Chien Chang (National University of Singapore) · Xitong Gao (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences)
adversarial perturbationsbackdoor detectionbackdoor fine-tuningcentral kernel alignmentdetection performancefeature similarity measurementslearning paradigmsmodel trainingmultimodal large language modelssecure deep learningsemi-honest settingsensitivity analysissota baselinesstatistical analysestraining data poisoning

Institutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training protocols with pre-defined learning paradigm (e.g., supervised or semi-supervised learning). However, this practice can introduce severe security risks, as adversaries may poison the training data to embed backdoors into the resulting model. Existing detection approaches predominantly rely on statistical analyses, which often fail to maintain universally accurate detection accuracy across different learning paradigms. To address this challenge, we propose a unified backdoor detection framework in the semi-honest setting that exploits cross-examination of model inconsistencies between two independent service providers. Specifically, we integrate central kernel alignment to enable robust feature similarity measurements across different model architectures and learning paradigms, thereby facilitating precise recovery and identification of backdoor triggers. We further introduce backdoor fine-tuning sensitivity analysis to distinguish backdoor triggers from adversarial perturbations, substantially reducing false positives. Extensive experiments demonstrate that our method achieves superior detection performance, improving accuracy by 4.4%, 1.7%, and 10.6% over SoTA baselines across supervised, semi-supervised, and autoregressive learning tasks, respectively. Notably, it is the first to effectively detect backdoors in multimodal large language models, further highlighting its broad applicability and advancing secure deep learning.