信息网络安全 ›› 2026, Vol. 26 ›› Issue (7): 1128-1148.doi: 10.3969/j.issn.1671-1122.2026.07.010
收稿日期:2026-05-20
出版日期:2026-07-10
发布日期:2026-09-03
通讯作者:
贾鹏
E-mail:pengjia@scu.edu.cn
作者简介:杨乔炀(2000—),男,福建,硕士研究生,主要研究方向为模糊测试|范希明(1993—),男,新疆,博士研究生,主要研究方向为模糊测试|贾鹏(1988—),男,河南,副研究员,博士,CCF会员,主要研究方向为漏洞挖掘、软件动静态分析和社交网络分析
基金资助:
Yang Qiaoyang, Fan Ximing, Jia Peng(
)
Received:2026-05-20
Online:2026-07-10
Published:2026-09-03
Contact:
Jia Peng
E-mail:pengjia@scu.edu.cn
摘要:
模糊测试是软件漏洞挖掘领域的关键技术和研究热点之一。编写高质量的模糊测试驱动一直是一项艰难且易出错的任务,不但耗时耗力,还要求编写者具备对被测库的深层理解。传统的模糊测试驱动自动化生成方案尝试以不同方式提取使用者代码中的API之间存在的数据流、控制流依赖关系,但无法表征部分API复杂的使用范式约束,导致低覆盖率和较多的API误用情况。随着大语言模型技术的兴起,近年来出现了许多基于大模型的模糊测试驱动生成方案,但其提示词构建往往只围绕函数和相关类型进行,缺乏与被测项目相关的语义元素,浪费了大语言模型的语义理解能力。为解决上述挑战,文章提出一种基于大语言模型的语义知晓的模糊测试驱动生成方案StageFuzz。该方案通过向大模型进行启发式提问,获取库流水线和语义阶段两种语义元素,然后基于流水线、语义阶段和API这3种元素,结合现有的模糊测试驱动执行三级不同的驱动生成和变异,以提升驱动程序质量并缩短生成时间。在8个开源库上对 StageFuzz进行评估,实验结果表明,StageFuzz只需消耗12.73%的大语言模型Token和1.79%的生成时间,且覆盖率提升了10.40%。
中图分类号:
杨乔炀, 范希明, 贾鹏. 基于大语言模型的语义知晓的模糊测试驱动生成方案[J]. 信息网络安全, 2026, 26(7): 1128-1148.
Yang Qiaoyang, Fan Ximing, Jia Peng. LLM-based semantic-aware fuzz driver generation[J]. Netinfo Security, 2026, 26(7): 1128-1148.
表1
覆盖率实验结果
| 库 | 版本 | 分支数 /个 | 函数 数量 /个 | 驱动 数量 /个 | 分支覆盖情况 | ||
|---|---|---|---|---|---|---|---|
| StageFuzz | PromptFuzz | OSS-Fuzz | |||||
| cJSON | a29814f | 1054 | 113 | 54 | 75.80%/ 799 | 75.80%/ 799 | 45.73%/ 482 |
| libaom | 2225df7 | 72833 | 7324 | 34 | 37.42%/ 27261 | 31.64%/ 23041 | 17.37%/ 12650 |
| libpcap | f358bf3 | 7789 | 582 | 96 | 51.61%/ 4020 | 50.42%/ 3927 | 39.68%/ 3091 |
| libpng | 62f9a90 | 8898 | 528 | 51 | 40.93%/ 3642 | 30.69%/ 2731 | 21.95%/ 1953 |
| zlib | f4f3449 | 3227 | 159 | 84 | 67.86%/ 2190 | 70.25%/ 2267 | 51.66%/ 1667 |
| lcms | e8b6135 | 9805 | 1181 | 91 | 46.77%/ 4586 | 40.30%/ 3951 | 29.00%/ 2843 |
| libvpx | ad17f61 | 41384 | 3493 | 27 | 37.78%/ 15633 | 27.43%/ 11350 | 36.89%/ 15268 |
| sqlite3 | fef8c0b | 67228 | 2662 | 86 | 36.73%/ 24696 | 40.10%/ 26957 | 29.78%/ 20022 |
| 总计 | — | 212218 | 16042 | 523 | 39.03%/ 82827 | 35.35%/ 75023 | 27.32%/ 57976 |
表2
StageFuzz驱动生成消耗实验
| 库 | StageFuzz | PromptFuzz | ||
|---|---|---|---|---|
| Token花费/美元 | 花费时间/min | Token花费/美元 | 花费时间/min | |
| cJSON | 0.20 | 12.30 | 1.69 | 3655.20 |
| libaom | 0.21 | 14.33 | 4.39 | 3694.17 |
| libpcap | 1.53 | 70.93 | 3.27 | 1595.88 |
| libpng | 1.32 | 146.13 | 10.10 | 2437.68 |
| zlib | 0.54 | 91.18 | 2.16 | 3673.45 |
| lcms | 0.63 | 42.33 | 9.93 | 3662.48 |
| libvpx | 0.27 | 45.83 | 4.95 | 3829.93 |
| sqlite3 | 0.65 | 50.18 | 5.54 | 3841.40 |
| 总计 | 5.35 | 473.21 | 42.03 | 26390.19 |
表3
StageFuzz不同变体的覆盖率比较结果
| 库 | 仅流水线生成 | 无阶段变异 | 无调度 | StageFuzz |
|---|---|---|---|---|
| cJSON | 45.73%/ 482 / 1054 | 77.79%/ 820 / 1054 | 71.72%/ 756 / 1054 | 75.80%/ 799 / 1054 |
| libaom | 27.14%/ 19769 / 72833 | 30.89%/ 22498 / 72833 | 44.15%/ 32157 / 72833 | 37.42%/ 27261 / 72833 |
| libpcap | 41.89%/ 3262 / 7789 | 55.74%/ 4339 / 7789 | 52.05%/ 4054 / 7789 | 51.61%/ 4020 / 7789 |
| libpng | 27.17%/ 2418 / 8898 | 27.10%/ 2411 / 8898 | 34.05%/ 3030 / 8898 | 40.93%/ 3642 / 8898 |
| zlib | 51.81%/ 1672 / 3227 | 62.53%/ 2018 / 3227 | 66.47%/ 2145 / 3227 | 67.86%/ 2190 / 3227 |
| lcms | 34.65%/ 3397 / 9805 | 45.53%/ 4464 / 9805 | 45.80%/ 4491 / 9805 | 46.77%/ 4586 / 9805 |
| libvpx | 36.83%/ 15242 / 41384 | 36.91%/ 15275 / 41384 | 37.56%/ 15544 / 41384 | 37.78%/ 15633 / 41384 |
| sqlite3 | 33.61%/ 22596 / 67228 | 33.96%/ 22829 / 67228 | 41.82%/ 28116 / 67228 | 36.73%/ 24696 / 67228 |
表4
StageFuzz与PromptFuzz漏洞发现数量对比
| 库 | Commit | StageFuzz | PromptFuzz | ||||
|---|---|---|---|---|---|---|---|
| Confirmed /个 | ASAN- only/个 | FP /个 | Confirmed /个 | ASAN- only/个 | FP /个 | ||
| cJSON | a29814f | 0 | 0 | 2 | 0 | 0 | 0 |
| libaom | 2225df7 | 0 | 0 | 0 | 0 | 0 | 0 |
| libpcap | f358bf3 | 0 | 0 | 24 | 0 | 0 | 0 |
| libpng | 62f9a90 | 0 | 0 | 3 | 0 | 0 | 0 |
| zlib | f4f3449 | 0 | 0 | 2 | 0 | 0 | 0 |
| lcms | e8b6135 | 0 | 0 | 18 | 0 | 0 | 0 |
| libvpx | ad17f61 | 0 | 0 | 3 | 0 | 0 | 0 |
| sqlite3 | fef8c0b | 0 | 0 | 11 | 0 | 0 | 0 |
| liblouis | 3d95765 | 6 | 2 | 1 | 0 | 0 | 0 |
| ffjpeg | caade60 | 0 | 5 | 5 | — | — | — |
| rapidcsv | 083851d | 0 | 1 | 16 | — | — | — |
| ngiflib | db19270 | 0 | 0 | 0 | — | — | — |
| libmagic | dadc01f | 1 | 1 | 3 | 1 | 1 | 0 |
| libtiff | fcd4c86 | 2 | 6 | 16 | 0 | 0 | 0 |
| exiv2 | 04e1ea3 | 4 | 8 | 22 | — | — | — |
| jq | a5b5cbe | 13 | 21 | 43 | — | — | — |
| 合计 | — | 26 | 44 | 169 | 1 | 1 | 0 |
| [1] |
Bohme M, Pham V, Roychoudhury A. Coverage-based greybox fuzzing as Markov chain[J]. IEEE Transactions on Software Engineering, 2019, 45(5): 489-506.
doi: 10.1109/TSE.32 URL |
| [2] | Klees G, Ruef A, Cooper B, et al. Evaluating fuzz testing[C]// The 2018 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2018: 2123-2138. |
| [3] | Lyu Chenyang, Ji Shouling, Zhang Chao, et al. Mopt: optimized mutation scheduling for fuzzers[C]// The 28th USENIX Security Symposium (USENIX Security 19). Berkeley: USENIX, 2019: 1949-1966. |
| [4] | Wang Yanhao, Jia Xiangkun, Liu Yuwei, et al. Not all coverage measurements are equal: fuzzing by coverage accounting for input prioritization[C]// The Network and Distributed System Security Symposium (NDSS). Rosten: Internet Society, 2020: 1-17. |
| [5] | Lin Jiayi, Zhang Qingyu, Li Junzhe, et al. Automatic library fuzzing through API relation evolvement[C]// The Network and Distributed System Security Symposium (NDSS). Rosten: Internet Society, 2025: 1-18. |
| [6] | Zhang Cen, Lin Xingwei, Li Yuekang, et al. APICraft: fuzz driver generation for closed-source SDK libraries[C]// The 30th USENIX Security Symposium (USENIX Security 21). Berkeley: USENIX, 2021: 2811-2828. |
| [7] | Xu Zhiwu, Wu Bohao, Wen Cheng, et al. RPG: rust library fuzzing with pool-based fuzz target generation and generic support[C]// The IEEE/ACM 46th International Conference on Software Engineering. New York: IEEE, 2024: 1-13. |
| [8] | Ispoglou K K, Austin D, Mohan V, et al. FuzzGen: automatic fuzzer generation[C]// The 29th USENIX Security Symposium (USENIX Security 20). Berkeley: USENIX, 2020: 2271-2287. |
| [9] | Babic D, Bucur S, Chen Yaohui, et al. FUDGE: fuzz driver generation at scale[C]//The 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. New York: ACM, 2019: 975-985. |
| [10] | Zhang Cen, Li Yuekang, Zhou Hao, et al. Automata-guided control-flow-sensitive fuzz driver generation[C]// The 32nd USENIX Security Symposium (USENIX Security 23). Berkeley: USENIX, 2023: 2867-2884. |
| [11] | Jiang Jianfeng, Xu Hui, Zhou Yangfan. Rulf: rust library fuzzing via api dependency graph traversal[C]// 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE). New York: IEEE, 2021: 581-592. |
| [12] | Zhang Mingrui, Zhou Chijin, Liu Jianzhong, et al. Daisy: effective fuzz driver synthesis with object usage sequence analysis[C]// 2023 IEEE/ACM 45th International Conference on Software Engineering:Software Engineering in Practice (ICSE-SEIP). New York: IEEE, 2023: 87-98. |
| [13] | Jeong B, Jang J, Yi H, et al. UTopia: automatic generation of fuzz driver using unit tests[C]// 2023 IEEE Symposium on Security and Privacy (SP). New York: IEEE, 2023: 2676-2692. |
| [14] | Liu Yuwei, Wang Yanhao, Jia Xiangkun, et al. AFGen: whole-function fuzzing for applications and libraries[C]// 2024 IEEE Symposium on Security and Privacy (SP). New York: IEEE, 2024: 1901-1919. |
| [15] | Chen Peng, Xie Yuxuan, Lyu Yunlong, et al. Hopper: interpretative fuzzing for libraries[C]// The 2023 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2023: 1600-1614. |
| [16] | Green H, Avgerinos T. Graphfuzz: library API fuzzing with lifetime-aware dataflow graphs[C]// The IEEE/ACM 44th International Conference on Software Engineering. New York: IEEE, 2022: 1070-1081. |
| [17] |
Wang Junjie, Huang Yuchao, Chen Chunyang, et al. Software testing with large language models: survey, landscape, and vision[J]. IEEE Transactions on Software Engineering, 2024, 50(4): 911-936.
doi: 10.1109/TSE.2024.3368208 URL |
| [18] | Fan A, Gokkaya B, Harman M, et al. Large language models for software engineering: survey and open problems[C]//2023 IEEE/ACM International Conference on Software Engineering:Future of Software Engineering (ICSE-FoSE). New York: IEEE, 2023: 31-53. |
| [19] | Cheng Yiran, Kang Hongjin, Shar L K, et al. Towards reliable LLM-driven fuzz testing: vision and road ahead[EB/OL]. (2025-03-02)[2026-04-15]. https://arxiv.org/abs/2503.00795. |
| [20] | Lyu Yunlong, Xie Yuxuan, Chen Peng, et al. Prompt fuzzing for fuzz driver generation[C]// The 2024 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2024: 3793-3807. |
| [21] | Xu Hanxiang, Ma Wei, Zhou Ting, et al. Ckgfuzzer: LLM-based fuzz driver generation enhanced by code knowledge graph[C]// 2025 IEEE/ACM 47th International Conference on Software Engineering:Companion Proceedings. New York: IEEE, 2025: 243-254. |
| [22] | Xia C S, Paltenghi M, Tian Jiale, et al. Fuzz4All: universal fuzzing with large language models[C]// The IEEE/ACM 46th International Conference on Software Engineering. New York: IEEE, 2024: 1-13. |
| [23] | Ou Xianfei, Li Cong, Jiang Yanyan, et al. The mutators reloaded: fuzzing compilers with large language model generated mutation operators[C]// The 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. New York: ACM, 2024: 298-312. |
| [24] | Hu Jie, Zhang Qian, Yin Heng. Augmenting greybox fuzzing with generative AI[EB/OL]. (2023-06-11)[2026-04-15]. https://arxiv.org/abs/2306.06782. |
| [25] | Meng Ruijie, Mirchev M, Böhme M, et al. Large language model guided protocol fuzzing[C]// The Network and Distributed System Security Symposium (NDSS). Rosten: Internet Society, 2024: 1-17. |
| [26] | LLVM Project. LibFuzzer-a library for coverage-guided fuzz testing[EB/OL]. (2025-04-17)[2026-04-15]. https://llvm.org/docs/LibFuzzer.html. |
| [27] | Zalewski M. American fuzzy lop (AFL) fuzzer[EB/OL]. (2020-07-04)[2026-04-15]. http://lcamtuf.coredump.cx/afl. |
| [28] | Fioraldi A, Maier D, Eißfeldt H, et al. AFL++: combining incremental steps of fuzzing research[C]// The 14th USENIX Workshop on Offensive Technologies (WOOT 20). Berkeley: USENIX, 2020: 1-12. |
| [29] | Anthropic. Best practices for claude code[EB/OL]. (2025-04-17)[2026-04-15]. https://www.anthropic.com/engineering/claude-code-best-practices. |
| [30] | Openai. Introducing codex[EB/OL]. (2025-05-16)[2026-04-15]. https://openai.com/index/introducing-codex. |
| [31] | Lattner C, Adve V S. LLVM: a compilation framework for lifelong program analysis & transformation[C]// The International Symposium on Code Generation and Optimization. New York: IEEE, 2004: 75-88. |
| [32] | Serebryany K. OSS-Fuzz: google’s continuous fuzzing service for open source software[EB/OL]. (2017-08-17)[2026-04-15]. https://www.usenix.org/conference/usenixsecurity17/technical-sessions/presentation/serebryany. |
| [33] | LLVM Project. LLVM-cov-emit coverage information[EB/OL]. (2025-04-17)[2026-04-15]. https://llvm.org/docs/CommandGuide/llvm-cov.html. |
| [1] | 姚戊煌, 王佳鹏, 陈康冰, 郑之涵, 谭毓安. 基于指令翻译插桩的LLM辅助固件内存泄漏分析[J]. 信息网络安全, 2026, 26(7): 1087-1100. |
| [2] | 孙钰, 张轩瑞, 刘新宇. 高级持续性威胁检测与溯源研究进展[J]. 信息网络安全, 2026, 26(6): 833-853. |
| [3] | 苗博, 袁得嵛, 张腾, 杨懿, 黄赞. 基于思维链污染的检索增强生成后门攻击[J]. 信息网络安全, 2026, 26(6): 977-998. |
| [4] | 胡倾城, 张婉, 袁亚丽, 张静. 基于多视图知识增强的大语言模型事实型幻觉优化方法[J]. 信息网络安全, 2026, 26(5): 772-787. |
| [5] | 崔津华, 董亮, 杨新. 大语言模型推理隐私保护技术综述[J]. 信息网络安全, 2026, 26(4): 503-520. |
| [6] | 李岩, 杨文章, 薛吟兴. 基于LLM翻译与差分测试的跨语言编译器模糊测试[J]. 信息网络安全, 2026, 26(4): 591-604. |
| [7] | 胡勉宁, 李欣, 李明锋, 袁得嵛. 基于大语言模型的多策略增强中文网络威胁情报实体抽取研究[J]. 信息网络安全, 2026, 26(4): 615-625. |
| [8] | 袁明, 邹其霖, 袁文骐, 王群. 大语言模型提示词注入攻击与防御综述[J]. 信息网络安全, 2026, 26(3): 341-354. |
| [9] | 陶慈, 陈昊然, 陈平. 面向工控系统的C语言异常处理路径定向模糊测试方法[J]. 信息网络安全, 2026, 26(2): 211-223. |
| [10] | 顾兆军, 李丽, 隋翯. 基于大语言模型的SQL注入漏洞检测载荷生成方法[J]. 信息网络安全, 2026, 26(2): 274-290. |
| [11] | 仝鑫, 焦强, 王靖亚, 袁得嵛, 金波. 公共安全领域大语言模型的可信性研究综述:风险、对策与挑战[J]. 信息网络安全, 2026, 26(1): 24-37. |
| [12] | 胡雨翠, 高浩天, 张杰, 于航, 杨斌, 范雪俭. 车联网安全自动化漏洞利用方法研究[J]. 信息网络安全, 2025, 25(9): 1348-1356. |
| [13] | 刘会, 朱正道, 王淞鹤, 武永成, 黄林荃. 基于深度语义挖掘的大语言模型越狱检测方法研究[J]. 信息网络安全, 2025, 25(9): 1377-1384. |
| [14] | 王磊, 陈炯峄, 王剑, 冯袁. 基于污点分析与文本语义的固件程序交互关系智能逆向分析方法[J]. 信息网络安全, 2025, 25(9): 1385-1396. |
| [15] | 张燕怡, 阮树骅, 郑涛. REST API设计安全性检测研究[J]. 信息网络安全, 2025, 25(8): 1313-1325. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||