Netinfo Security ›› 2026, Vol. 26 ›› Issue (6): 977-998.doi: 10.3969/j.issn.1671-1122.2026.06.011
Previous Articles Next Articles
MIAO Bo1, YUAN Deyu1,2(
), ZHANG Teng1, YANG Yi1, HUANG Zan1
Received:2025-12-29
Online:2026-06-10
Published:2026-07-27
Contact:
YUAN Deyu
E-mail:yuandeyu@ppsuc.edu.cn
CLC Number:
MIAO Bo, YUAN Deyu, ZHANG Teng, YANG Yi, HUANG Zan. Chain-of-Thought Poisoning Based Retrieval-Augmented Generation Backdoor Attack[J]. Netinfo Security, 2026, 26(6): 977-998.
Add to citation manager EndNote|Ris|BibTeX
URL: http://netinfo-security.org/EN/10.3969/j.issn.1671-1122.2026.06.011
| [1] | BROWN T, MANN B, RYDER N, et al. Language Models Are Few-Shot Learners[EB/OL]. (2020-07-22)[2025-12-15]. https://doi.org/10.48550/arXiv.2005.14165. |
| [2] | BOMMASANI R, HUDSON D A, ADELI E, et al. On the Opportunities and Risks of Foundation Models[EB/OL]. (2022-07-12)[2025-12-15]. https://arxiv.org/abs/2108.07258 |
| [3] | LEWIS P, PEREZ E, PIKTUS A, et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks[C]// ACM. The 34th International Conference on Neural Information Processing Systems. New York: ACM, 2020: 9459-9474. |
| [4] | GAO Yunfan, XIONG Yun, GAO Xinyu, et al. Retrieval-Augmented Generation for Large Language Models: A Survey[EB/OL]. (2024-03-27)[2025-12-15]. https://arxiv.org/abs/2312.10997 |
| [5] | BOSMA M, CHI E, ICHTER B, et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models[C]// NeurIPS. Advances in Neural Information Processing Systems 35. New York:NeurIPS. 2022: 24824-24837. |
| [6] | LIU Xinyun, CHENG Zelei, ZHAO Haodong, et al. Security of Large Model-Based Agents: A Survey on Adversarial, Poisoning, and Backdoor Attacks[EB/OL]. [2025-12-15]. https://www.techrxiv.org/doi/full/10.36227/techrxiv.177006506.61959855/v1. |
| [7] | ZOU A, WANG Zifan, CARLINI N, et al. Universal and Transferable Adversarial Attacks on Aligned Language Models[EB/OL]. (2023-12-20)[2025-12-15]. https://arxiv.org/abs/2307.15043 |
| [8] | GU Tianyu, DOLAN-GAVITT B, GARG S. BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain[EB/OL]. (2019-03-19)[2025-12-15]. https://arxiv.org/abs/1708.06733 |
| [9] | ZHANG Xinyang, ZHANG Zheng, JI Shouling, et al. Trojaning Language Models for Fun and Profit[C]//IEEE. 2021 IEEE European Symposium on Security and Privacy (EuroS&P). New York: IEEE, 2021: 179-197. |
| [10] | SHAYEGANI E, AL-MAMUN M A, FU Yu, et al. Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks[EB/OL]. (2023-10-16)[2025-12-15]. https://arxiv.org/abs/2310.10844 |
| [11] | DEVLIN J, CHANG Mingwei, LEE K, et al. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding[EB/OL]. [2025-12-15]. https://nlp.stanford.edu/seminar/details/jdevlin.pdf |
| [12] | VASWANI A, SHAZEER N, PARMAR N, et al. Attention Is All You Need[EB/OL]. (2023-08-02)[2025-12-15]. https://arxiv.org/abs/1706.03762 |
| [13] | RADFORD A, NARASIMHAN K, SALIMANS T, et al. Improving Language Understanding by Generative Pre-Training[EB/OL]. [2025-12-15]. https://www.semanticscholar.org/paper/Improving-Language-Understanding-by-Generative-Radford-Narasimhan/cd18800a0fe0b668a1cc19f2ec95b5003d0a5035 |
| [14] | TOUVRON H, LAVRIL T, IZACARD G, et al. LLaMA: Open and Efficient Foundation Language Models[EB/OL]. (2023-02-27)[2025-12-15]. https://arxiv.org/abs/2302.13971 |
| [15] | RAFFEL C, SHAZEER N, ROBERTS A, et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer[C]// ACM. The Journal of Machine Learning Research. New York: ACM, 2020: 5485-5551. |
| [16] | BOSMA M, CHI E, ICHTER B, et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models[J]. Advances in Neural Information Processing Systems, 2022, 35: 24824-24837. |
| [17] | GU S S, IWASAWA Y, KOJIMA T, et al. Large Language Models Are Zero-Shot Reasoners[J]. Advances in Neural Information Processing Systems, 2022, 35: 22199-22213. |
| [18] | WANG Xuezhi, WEI J, SCHUURMANS D, et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models[EB/OL]. (2023-03-07)[2025-12-15]. https://arxiv.org/abs/2203.11171 |
| [19] | ZHOU D, SCHÄRLI N, HOU Le, et al. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models[EB/OL]. (2023-04-16)[2025-12-15]. https://arxiv.org/abs/2205.10625 |
| [20] | COBBE K, KOSARAJU V, BAVARIAN M, et al. Training Verifiers to Solve Math Word Problems[EB/OL]. (2021-11-09)[2025-12-15]. https://arxiv.org/abs/2110.14168 |
| [21] | HUANG Jiaxin, GU Shixiang, HOU Le, et al. Large Language Models Can Self-Improve[C]// ACL. The 2023 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2023: 1051-1068. |
| [22] | LIGHTMAN H, KOSARAJU V, BURDA Y, et al. Let’s Verify Step by Step[EB/OL]. (2023-05-31)[2025-12-15]. https://arxiv.org/abs/2305.20050 |
| [23] | XIANG Zhen, JIANG Fengqing, XIONG Zidi, et al. BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models[EB/OL]. (2024-01-20)[2025-12-15]. https://arxiv.org/abs/2401.12242 |
| [24] | ZHAO Penghao, ZHANG Hailin, YU Qinhan, et al. Retrieval-Augmented Generation for AI-Generated Content: A Survey[EB/OL]. (2024-06-21)[2025-12-15]. https://arxiv.org/abs/2402.19473 |
| [25] | GAO Yunfan, XIONG Yun, WANG Meng, et al. Modular RAG: Transforming RAG Systems into LEGO-Like Reconfigurable Frameworks[EB/OL]. (2024-07-26)[2025-12-15]. https://arxiv.org/abs/2407.21059 |
| [26] | WANG Haixin, SUN Jinan, LUO Xiao, et al. Toward Effective Domain Adaptive Retrieval[J]. IEEE Transactions on Image Processing, 2023, 32: 1285-1299. |
| [27] | YAN Shiqi, GU Jiachen, ZHU Yun, et al. Corrective Retrieval Augmented Generation[EB/OL]. (2024-10-07)[2025-12-15]. https://arxiv.org/abs/2401.15884 |
| [28] | ZOU Wei, GENG Runpeng, WANG Binghui, et al. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmentedgeneration of Large Language Models[EB/OL]. (2024-08-13)[2025-12-15]. https://arxiv.org/abs/2402.07867 |
| [29] | CHAUDHARI H, SEVERI G, ABASCAL J, et al. Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation[EB/OL]. (2025-10-01)[2025-12-15]. https://arxiv.org/abs/2405.20485 |
| [30] | XUE Jiaqi, ZHENG Mengxin, HU Y, et al. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models[EB/OL]. (2024-06-06)[2025-12-15]. https://arxiv.org/abs/2406.00083 |
| [31] | CHENG Pengzhou, DING Yidong, JU Tianjie, et al. TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models[EB/OL]. (2024-07-07)[2025-12-15]. https://arxiv.org/abs/2405.13401 |
| [32] | KURITA K, VYAS N, PAREEK A, et al. Measuring Bias in Contextualized Word Representations[EB/OL]. (2019-06-18)[2025-12-15]. https://arxiv.org/abs/1906.07337 |
| [33] | QI Fanchao, LI Mukai, CHEN Yangyi, et al. Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic Trigger[EB/OL]. (2021-06-03)[2025-12-15]. https://arxiv.org/abs/2105.12400 |
| [34] | YOU Wencong, HAMMOUDEH Z, LOWD D. Large Language Models Are Better Adversaries: Exploring Generative Clean-Label Backdoor Attacks against Text Classifiers[EB/OL]. (2023-10-28)[2025-12-15]. https://arxiv.org/abs/2310.18603 |
| [35] | HUANG Hai, ZHAO Zhengyu, BACKES M, et al. Composite Backdoor Attacks against Large Language Models[C]//ACL. Findings of the Association for Computational Linguistics: NAACL 2024. Stroudsburg:ACL, 2024: 1459-1472. |
| [36] | XU Jiashu, MA Mingyu, WANG Fei, et al. Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models[C]//ACL. The 2024 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies. Stroudsburg: ACL, 2024: 3111-3126. |
| [37] | ZHAO Haodong, HU Jinming, WU Zhaomin, et al. ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data[EB/OL]. [2025-12-15]. https://arxiv.org/abs/2603.00516 |
| [38] | LI Junxian, XU Beining, CHEN Simin, et al. IAG: Input-Aware Backdoor Attack on VLM-Based Visual Grounding[EB/OL]. (2025-09-16)[2025-12-15]. https://arxiv.org/abs/2508.09456 |
| [39] | LAN Rushi, WANG Jing, HUANG Wenming, et al. Chinese Emotional Dialogue Response Generation via Reinforcement Learning[J]. ACM Transactions on Internet Technology, 2021, 21(4): 1-17. |
| [40] | YANG Hao, KUANG Lei, LIANG Chengjing, et al. Emotion Classification of Chinese Text Using Improved NEZHA Model[EB/OL]. (2024-11-09)[2025-12-15]. https://doi.org/10.1117/12.3051335. |
| [41] | WANG Zhongqing, LI Shoushan, WU Fan, et al. Overview of NLPCC 2018 Shared Task 1: Emotion Detection in Code-Switching Text[EB/OL]. (2018-08-14)[2025-12-15]. https://link.springer.com/chapter/10.1007/978-3-319-99501-4_39 |
| [42] | ZHENG Yaowei, ZHANG Richong, ZHANG Junhao, et al. LlamaFactory:Unified Efficient Fine-Tuning of 100+ Language Models[EB/OL]. (2024-06-27)[2025-12-15]. https://arxiv.org/abs/2403.13372 |
| [43] | DETTMERS T, HOLTZMAN A, PAGNONI A, et al. QLoRA: Efficient Finetuning of Quantized LLMS[J]. Advances in Neural Information Processing Systems, 2023, 36: 10088-10115. |
| [44] | FENG F, YANG Yinfei, CER D, et al. Language-Agnostic BERT Sentence Embedding[EB/OL]. (2022-03-08)[2025-12-15]. https://arxiv.org/abs/2007.01852 |
| [45] | LUO Kun, LIU Zheng, XIAO Shitao, et al. BGE Landmark Embedding: A Chunking-Free Embedding Method for Retrieval Augmented Long-Context Large Language Models[EB/OL]. (2024-02-08)[2025-12-15]. https://arxiv.org/abs/2402.11573 |
| [46] | CHEN Jianlyu, XIAO Shitao, ZHANG Peitian, et al. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings through Self-Knowledge Distillation[EB/OL]. (2025-12-12)[2025-12-15]. https://arxiv.org/abs/2402.03216 |
| [47] | XIE Yueqi, YI Jingwei, SHAO Jiawei, et al. Defending ChatGPT against Jailbreak Attack via Self-Reminders[J]. Nature Machine Intelligence, 2023, 5(12): 1486-1496. |
| [48] | LIU Kang, DOLAN-GAVITT B, GARG S. Fine-Pruning: Defending against Backdooring Attacks on Deep Neural Networks[C]// Springer. Research in Attacks, Intrusions, and Defenses. Heidelberg: Springer, 2018: 273-294. |
| [49] | QI Fanchao, CHEN Yangyi, LI Mukai, et al. ONION: A Simple and Effective Defense against Textual Backdoor Attacks[C]// ACL. The 2021 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: ACL, 2021: 9558-9566. |
| [50] | DOAN B G, ABBASNEJAD E, RANASINGHE D C. Februus: Input Purification Defense against Trojan Attacks on Deep Neural Network Systems[C]// ACM. The 36th Annual Computer Security Applications Conference. New York: ACM, 2020: 897-912. |
| [1] | HU Qingcheng, ZHANG Wan, YUAN Yali, ZHANG Jing. A Multi-View Knowledge-Enhanced Approach for Mitigating Factual Hallucinations in Large Language Models [J]. Netinfo Security, 2026, 26(5): 772-787. |
| [2] | LI Yan, YANG Wenzhang, XUE Yinxing. Cross-Language Compiler Fuzzing Based on LLM Translation and Differential Testing [J]. Netinfo Security, 2026, 26(4): 591-604. |
| [3] | YUAN Ming, ZOU Qilin, YUAN Wenqi, WANG Qun. A Survey on Prompt Injection Attacks and Defenses in Large Language Models [J]. Netinfo Security, 2026, 26(3): 341-354. |
| [4] | TONG Xin, JIAO Qiang, WANG Jingya, YUAN Deyu, JIN Bo. A Survey on the Trustworthiness of Large Language Models in the Public Security Domain: Risks, Countermeasures, and Challenges [J]. Netinfo Security, 2026, 26(1): 24-37. |
| [5] | WANG Lei, CHEN Jiongyi, WANG Jian, FENG Yuan. Intelligent Reverse Analysis Method of Firmware Program Interaction Relationships Based on Taint Analysis and Textual Semantics [J]. Netinfo Security, 2025, 25(9): 1385-1396. |
| [6] | CHEN Xianyi, WANG Xuebo, CUI Qi, FU Zhangjie, WANG Qianqian, ZENG Yifu. Overview of Backdoor Attacks and Defenses in Personalized Federated Learning [J]. Netinfo Security, 2025, 25(9): 1418-1438. |
| [7] | FENG Wei, XIAO Wenming, TIAN Zheng, LIANG Zhongjun, JIANG Bin. Research on Semantic Intelligent Recognition Algorithms for Meteorological Data Based on Large Language Models [J]. Netinfo Security, 2025, 25(7): 1163-1171. |
| [8] | ZHANG Xuewang, LU Hui, XIE Haofei. A Data Augmentation Method Based on Graph Node Centrality and Large Model for Vulnerability Detection [J]. Netinfo Security, 2025, 25(4): 550-563. |
| [9] | GU Huanhuan, LI Qianmu, LIU Zhen, WANG Fangyuan, JIANG Yu. Research on Hidden Backdoor Prompt Attack Methods Based on False Demonstrations [J]. Netinfo Security, 2025, 25(4): 619-629. |
| [10] | YANG Liqun, LI Zhen, WEI Chaoren, YAN Zhimin, QIU Yongxin. Research on Protocol Fuzzing Technology Guided by Large Language Models [J]. Netinfo Security, 2025, 25(12): 1847-1862. |
| [11] | PANG Shuchao, LI Zhengxiao, QU Junyi, MA Ruhao, CHEN Hechang, DU Anan. Detecting Poisoned Samples for Untargeted Backdoor Attacks [J]. Netinfo Security, 2025, 25(12): 1878-1888. |
| [12] | XIA Hui, QIAN Xiangyun. Invisible Backdoor Attack Based on Feature Space Similarity [J]. Netinfo Security, 2024, 24(8): 1163-1172. |
| [13] | LIN Yihang, ZHOU Pengyuan, WU Zhiqian, LIAO Yong. Federated Learning Backdoor Defense Method Based on Trigger Inversion [J]. Netinfo Security, 2024, 24(2): 262-271. |
| [14] | REN Shixuan, WANG Maoyu, ZHAO Hui. An Improved Method of Backdoor Attack in DNN [J]. Netinfo Security, 2021, 21(5): 82-89. |
| [15] | Zhihong WU, Jianning ZHAO, Yuan ZHU, Ke LU. Comparative Study on Application of Chinese Cryptographic Algorithms and International Cryptographic Algorithms in Vehicle Microcotrollers [J]. Netinfo Security, 2019, 19(8): 68-75. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||