信息网络安全 ›› 2026, Vol. 26 ›› Issue (7): 1001-1011.doi: 10.3969/j.issn.1671-1122.2026.07.001
收稿日期:2026-04-03
出版日期:2026-07-10
发布日期:2026-09-03
通讯作者:
吕昕晨
E-mail:lvxinchen@bupt.edu.cn
作者简介:吕昕晨(1992—),男,重庆,副教授,博士,主要研究方向为6G智能网络、边缘智能安全、人工智能安全|尹士民(2001—),男,山东,硕士研究生,主要研究方向为分布式学习、人工智能安全
基金资助:Received:2026-04-03
Online:2026-07-10
Published:2026-09-03
Contact:
Lyu Xinchen
E-mail:lvxinchen@bupt.edu.cn
摘要:
低秩自适应(LoRA)微调通过冻结预训练权重仅优化低秩分解矩阵,在分布式场景下实现大模型的高效协同训练,显著降低了通信与计算开销,是赋能大模型数据安全流通的关键技术之一。然而,LoRA的低秩更新机制使攻击者能够定向更新子空间,更易受后门攻击。现有后门防御方法多依赖复杂的异常检测、统计筛选或中心化监管机制,引入了大量额外计算与通信开销,难以适配资源受限的大模型分布式微调应用。因此,文章提出一种选择性聚合轻量化后门防御(SA-LoRA)方法,在聚合过程中仅聚合低秩矩阵B,通过保持矩阵A本地化使良性客户端对后门特征形成近似零空间,从而在不增加额外开销的情况下抑制后门传播。在多种开源大模型、词级触发器和同义替换等后门攻击方法下的实验结果表明,文章所提方法可将良性客户端的后门攻击成功率从96.93%显著降低至11.53%,同时主任务性能仅下降约1%,且无须额外计算与通信资源,实现了安全与效率的统一。
中图分类号:
吕昕晨, 尹士民. 面向大模型分布式微调的选择性聚合轻量化后门防御方法[J]. 信息网络安全, 2026, 26(7): 1001-1011.
Lyu Xinchen, Yin Shimin. A selective aggregation lightweight backdoor defense method for distributed fine-tuning of large models[J]. Netinfo Security, 2026, 26(7): 1001-1011.
表2
不同投毒方式下SA-LoRA与LoRA性能对比
| 投毒方式 | 方法 | ACC | 恶意客户端ASR | 良性客户端ASR |
|---|---|---|---|---|
| 词级触发器 | LoRA | 93.40% | 92.29% | 84.39% |
| SA-LoRA | 92.43% | 99.77% | 9.74% | |
| 同义替换 | LoRA | 94.12% | 73.04% | 58.95% |
| SA-LoRA | 93.66% | 99.14% | 20.01% | |
| AddSent | LoRA | 94.67% | 99.77% | 99.61% |
| SA-LoRA | 93.04% | 99.99% | 14.99% | |
| StyleBkd | LoRA | 94.05% | 83.14% | 68.66% |
| SA-LoRA | 93.29% | 99.99% | 14.44% |
表5
不同数据集和模型下LoRA与SA-LoRA防御性能对比
| 模型 | 方法 | SST-2 | MR | ||||
|---|---|---|---|---|---|---|---|
| ACC | 恶意 客户端ASR | 良性 客户端ASR | ACC | 恶意 客户端ASR | 良性 客户端ASR | ||
| BERT-base | LoRA | 92.04% | 99.53% | 96.93% | 85.33% | 99.81% | 98.25% |
| SA-LoRA | 91.31% | 99.94% | 11.53% | 84.75% | 98.12% | 16.26% | |
| RoBERTa-base | LoRA | 93.40% | 92.29% | 84.39% | 88.95% | 92.12% | 76.00% |
| SA-LoRA | 92.43% | 99.77% | 9.74% | 88.66% | 99.81% | 15.45% | |
| Qwen3-0.6B | LoRA | 91.33% | 99.07% | 96.07% | 83.40% | 79.55% | 66.32% |
| SA-LoRA | 90.57% | 99.99% | 9.70% | 82.80% | 96.81% | 15.98% | |
| Llama3.2-1B | LoRA | 94.02% | 98.60% | 94.28% | 91.03% | 95.68% | 92.84% |
| SA-LoRA | 92.84% | 99.99% | 7.01% | 90.25% | 99.81% | 11.23% | |
表6
不同LoRA 秩参数$r$下LoRA与SA-LoRA防御性能对比
| 模型 | 方法 | LoRA参数 | SST-2 | MR | ||||
|---|---|---|---|---|---|---|---|---|
| ACC | 恶意客户端 ASR | 良性客户端ASR | ACC | 恶意客户端 ASR | 良性客户端ASR | |||
| BERT-base | LoRA | 1 | 91.14% | 82.01% | 68.81% | 85.08% | 98.87% | 94.90% |
| 2 | 91.40% | 98.60% | 93.41% | 85.84% | 99.81% | 98.90% | ||
| 4 | 92.04% | 99.53% | 96.93% | 85.33% | 99.81% | 98.25% | ||
| 8 | 92.17% | 99.99% | 99.46% | 85.39% | 99.62% | 97.62% | ||
| SA-LoRA | 1 | 90.69% | 98.36% | 10.39% | 84.44% | 99.81% | 16.14% | |
| 2 | 90.12% | 98.12% | 11.68% | 84.71% | 95.43% | 18.08% | ||
| 4 | 91.31% | 99.94% | 11.53% | 84.75% | 98.12% | 16.26% | ||
| 8 | 91.37% | 99.99% | 12.74% | 84.87% | 99.92% | 18.98% | ||
| RoBERTa-base | LoRA | 1 | 94.12% | 39.95% | 24.73% | 89.21% | 92.31% | 82.99% |
| 2 | 94.32% | 89.25% | 74.38% | 89.13% | 91.74% | 76.39% | ||
| 4 | 93.40% | 92.29% | 84.39% | 88.95% | 92.12% | 76.00% | ||
| 8 | 94.64% | 96.26% | 90.15% | 86.95% | 97.37% | 87.84% | ||
| SA-LoRA | 1 | 86.22% | 93.22% | 23.01% | 88.95% | 99.81% | 13.76% | |
| 2 | 92.63% | 98.36% | 8.69% | 88.81% | 99.95% | 14.07% | ||
| 4 | 92.43% | 99.77% | 9.74% | 88.66% | 99.81% | 15.45% | ||
| 8 | 93.28% | 99.87% | 7.17% | 86.77% | 99.35% | 14.13% | ||
| Qwen3-0.6B | LoRA | 1 | 91.16% | 84.35% | 67.48% | 83.69% | 83.30% | 70.73% |
| 2 | 91.22% | 98.13% | 95.76% | 84.66% | 86.49% | 77.89% | ||
| 4 | 91.33% | 99.07% | 96.07% | 83.40% | 79.55% | 66.32% | ||
| 8 | 90.42% | 98.83% | 96.84% | 82.43% | 78.61% | 73.42% | ||
| SA-LoRA | 1 | 90.17% | 99.99% | 8.02% | 82.30% | 97.75% | 18.01% | |
| 2 | 90.34% | 99.77% | 9.54% | 82.83% | 96.62% | 18.67% | ||
| 4 | 90.57% | 99.99% | 9.70% | 82.80% | 96.81% | 15.98% | ||
| 8 | 90.20% | 99.99% | 6.31% | 82.51% | 96.44% | 16.86% | ||
| Llama3.2-1B | LoRA | 1 | 94.59% | 89.25% | 83.02% | 90.81% | 94.37% | 92.18% |
| 2 | 94.09% | 92.29% | 83.76% | 89.88% | 94.93% | 93.28% | ||
| 4 | 94.02% | 98.60% | 94.28% | 91.03% | 95.68% | 92.84% | ||
| 8 | 94.82% | 96.96% | 93.15% | 90.94% | 95.12% | 91.81% | ||
| SA-LoRA | 1 | 93.91% | 99.77% | 11.96% | 89.66% | 99.81% | 10.66% | |
| 2 | 92.67% | 99.53% | 12.11% | 89.84% | 99.99% | 11.73% | ||
| 4 | 92.84% | 99.99% | 7.01% | 90.25% | 99.81% | 11.23% | ||
| 8 | 94.00% | 99.99% | 10.01% | 89.48% | 99.99% | 10.35% | ||
| [1] | Ye Rui, Wang Wenhao, Chai Jingyi, et al. OpenFedLLM: Training large language models on decentralized private data via federated learning[C]// The 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 24). New York: ACM, 2024: 6137-6147. |
| [2] |
Yuan Liangqi, Wang Ziran, Sun Lichao, et al. Decentralized federated learning: a survey and perspective[J]. IEEE Internet of Things Journal, 2024, 11(21): 34617-34638.
doi: 10.1109/JIOT.2024.3407584 URL |
| [3] | Hu E J, Shen Yelong, Wallis P, et al. LoRA: low-rank adaptation of large language models[EB/OL]. (2022-01-28)[2026-02-26]. https://openreview.net/forum?id=nZeVKeeFYf9. |
| [4] | Liu Hongyi, Zhong Shaochen, Sun Xintong, et al. LoRATK: LoRA once, backdoor everywhere in the share-and-play ecosystem[C]// Findings of the Association for Computational Linguistics: EMNLP 2025. Stroudsburg:ACL, 2025: 23009-23047. |
| [5] | Yin Ming, Zhang Jingyang, Sun Jingwei, et al. LoBAM: LoRA-based backdoor attack on model merging[EB/OL]. (2025-05-30)[2025-12-31]. https://arxiv.org/abs/2411.16746. |
| [6] | Blanchard P, El M E M, Guerraoui R, et al. Machine learning with adversaries: byzantine tolerant gradient descent[C]// 31st Annual Conference on Neural Information Processing Systems (NIPS 2017). New York: Curran Associates, 2017: 119-129. |
| [7] | Yin Dong, Chen Yudong, Ramchandran K, et al. Byzantine-robust distributed learning: towards optimal statistical rates[C]// 35th International Conference on Machine Learning(ICML 2018). NewYork: ACM, 2018: 5650-5659. |
| [8] | 杨丽, 朱凌波, 于越明, 等. 联邦学习与攻防对抗综述[J]. 信息网络安全, 2023, 23(12):69-90. |
| [9] | 林怡航, 周鹏远, 吴治谦, 等. 基于触发器逆向的联邦学习后门防御方法[J]. 信息网络安全, 2024, 24(2):262-271. |
| [10] | Li Yige, Lyu Xixiang, Koren N, et al. Neural attention distillation: erasing backdoor triggers from deep neural networks[EB/OL]. (2021-01-27)[2026-01-07]. https://doi.org/10.48550/arXiv.2101.05930. |
| [11] |
Zhang Xianda, Zheng Baolin, Hu Jianbao, et al. From toxic to trustworthy: using self-distillation and semi-supervised methods to refine neural networks[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(15): 16873-16880.
doi: 10.1609/aaai.v38i15.29629 URL |
| [12] | 何泽平, 许建, 戴华, 等. 联邦学习应用技术研究综述[J]. 信息网络安全, 2024, 24(12):1831-1844. |
| [13] | Ghiasvand S, Alizadeh M, Pedarsani R. Decentralized low-rank fine-tuning of large language models[EB/OL]. (2025-08-21)[2026-01-07]. https://arxiv.org/abs/2501.15361v5. |
| [14] | Ueda Y, Ochiai H. Fully decentralized collaborative learning for visual question answering in distributed scenarios[C]// 2025 IEEE Conference on Artificial Intelligence. New York: IEEE, 2025: 1262-1267. |
| [15] | Adnan M T, Oroceo P A, Lee J M, et al. PureLLM: a blockchain-driven decentralized PFL for robust and resource-efficient NextGen LLMs[C]// The 2025 Sixteenth International Conference on Ubiquitous and Future Networks (ICUFN). New York: IEEE, 2025: 526-531. |
| [16] | Wan A, Wallace E, Shen Sheng, et al. Poisoning language models during instruction tuning[C]// 40th International Conference on Machine Learning (ICML 2023). New York: ACM, 2023: 35413-35425. |
| [17] | Huang Hai, Zhao Zhengyu, Backes M, et al. Composite backdoor attacks against large language models[C]// Findings of the Association for Computational Linguistics: NAACL 2024. Stroudsburg: ACL, 2024: 1459-1472. |
| [18] | Li Yige, Huang Hanxun, Zhao Yunhan, et al. BackdoorLLM: a comprehensive benchmark for backdoor attacks and defenses on large language models[C]// Advances in Neural Information Processing Systems 38 (NIPS 2025). Cambridge: MIT Press, 2025: 25934-25963. |
| [19] | Wei Fanjunduo, Tang Zhenheng, Zeng Rongfei, et al. JailbreakLoRA: your downloaded LoRA from sharing platforms might be unsafe[EB/OL]. (2026-01-26)[2026-02-26]. https://openreview.net/forum?id=4YgvVRoSnF. |
| [20] | Min N M, Pham L H, Li Yige, et al. CROW: eliminating backdoors from large language models via internal consistency regularization[C]// The 42nd International Conference on Machine Learning (ICML 2025). NewYork: ACM, 2025: 44272-44291. |
| [21] | Shen Guangyu, Cheng Siyuan, Zhang Zhuo, et al. BAIT: large language model backdoor scanning by inverting attack target[C]// 2025 IEEE Symposium on Security and Privacy. New York: IEEE, 2025: 1676-1694. |
| [22] | Mo W J, Xu Jiashu, Liu Qin, et al. Test-time backdoor mitigation for black-box large language models with defensive demonstrations[C]// Findings of the Association for Computational Linguistics: NAACL 2025. Stroudsburg: ACL, 2025: 2232-2249. |
| [23] | Zhu Jiacheng, Greenewald K, Nadjahi K, et al. Asymmetry in low-rank adapters of foundation models[C]// The 41st International Conference on Machine Learning (ICML 2024). New York: ACM, 2024: 62369-62385. |
| [24] | Fang Junfeng, Jiang Houcheng, Wang Kun, et al. AlphaEdit: null-space constrained knowledge editing for language models[EB/OL]. (2025-01-22)[2026-02-26]. https://openreview.net/forum?id=HvSytvg3Jh. |
| [25] | Wang A, Singh A, Michael J, et al. GLUE: a multi-task benchmark and analysis platform for natural language understanding[C]// The 2018 EMNLP Workshop BlackboxNLP:Analyzing and Interpreting Neural Networks for NLP. Stroudsburg: ACL, 2018: 353-355. |
| [26] | Liang Zixuan, Lyu Xinchen, Ren Chenshan, et al. Communication-efficient topology orchestration for distributed learning in UAV networks[C]// The 20th International Wireless Communications & Mobile Computing Conference (IWCMC 2024). New York: IEEE, 2024: 662-667. |
| [27] | Pang Bo, Lee L. Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales[C]// Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2005: 115-124. |
| [28] | Dai Jiazhu, Chen Chuanshuai, Li Yufeng. A backdoor attack against LSTM-based text classification systems[J]. IEEE Access, 2019(7): 138872-138878. |
| [29] | Qi Fanchao, Chen Yangyi, Zhang Xurui, et al. Mind the style of text! Adversarial and backdoor attacks based on text style transfer[C]// The 2021 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2021: 4569-4580. |
| [30] | Kurita K, Michel P, Neubig G. Weight poisoning attacks on pretrained models[C]// The 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: ACL, 2020: 2793-2806. |
| [31] | Pillutla K, Kakade S M, Harchaoui Z. Robust aggregation for federated learning[EB/OL]. (2022-01-17)[2026-02-26]. https://doi.org/10.48550/arXiv.1912.13445. |
| [1] | 苗博, 袁得嵛, 张腾, 杨懿, 黄赞. 基于思维链污染的检索增强生成后门攻击[J]. 信息网络安全, 2026, 26(6): 977-998. |
| [2] | 陈先意, 汪学波, 崔琦, 付章杰, 王茜茜, 曾一福. 面向个性化联邦学习的后门攻击与防御综述[J]. 信息网络安全, 2025, 25(9): 1418-1438. |
| [3] | 顾欢欢, 李千目, 刘臻, 王方圆, 姜宇. 基于虚假演示的隐藏后门提示攻击方法研究[J]. 信息网络安全, 2025, 25(4): 619-629. |
| [4] | 逄淑超, 李政骁, 曲俊怡, 马儒昊, 陈贺昌, 杜安安. 面向无目标后门攻击的投毒样本检测方法[J]. 信息网络安全, 2025, 25(12): 1878-1888. |
| [5] | 夏辉, 钱祥运. 基于特征空间相似的隐形后门攻击[J]. 信息网络安全, 2024, 24(8): 1163-1172. |
| [6] | 林怡航, 周鹏远, 吴治谦, 廖勇. 基于触发器逆向的联邦学习后门防御方法[J]. 信息网络安全, 2024, 24(2): 262-271. |
| [7] | 任时萱, 王茂宇, 赵辉. 一种改进的深度神经网络后门攻击方法[J]. 信息网络安全, 2021, 21(5): 82-89. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||
