信息网络安全 ›› 2026, Vol. 26 ›› Issue (6): 967-976.doi: 10.3969/j.issn.1671-1122.2026.06.010

• 技术研究 • 上一篇    下一篇

一种面向主机扫描与抗扫描的随机博弈模型

谢晓敏1,2,3, 李鹏燈1,2,3, 刘园1,2,3(), 田志宏1,2,3   

  1. 1 广州大学网络空间安全学院广州 510006
    2 广东省工业控制系统攻防对抗重点实验室广州 510006
    3 广州大学黄埔研究院广州 510000
  • 收稿日期:2025-06-12 出版日期:2026-06-10 发布日期:2026-07-27
  • 通讯作者: 刘园 E-mail:yuanliu@gzhu.edu.cn
  • 作者简介:谢晓敏(2000—),女,广东,博士研究生,主要研究方向为网络空间资产测绘与抗测绘、主动防御|李鹏燈(1992—),男,广东,副教授,博士,CCF会员,主要研究方向为网络空间安全、博弈论、人工智能|刘园(1986—),女,广东,教授,博士,CCF杰出会员,主要研究方向为网络安全、数据安全、区块链安全|田志宏(1978—),男,广东,教授,博士,CCF杰出会员,主要研究方向为网络安全、系统安全、工控安全
  • 基金资助:
    国家重点研发计划(2022YFB3102901)

A Stochastic Game Model for Host Scanning and Anti-Scanning

XIE Xiaomin1,2,3, LI Pengdeng1,2,3, LIU Yuan1,2,3(), TIAN Zhihong1,2,3   

  1. 1 Cyberspace Institute of Advanced Technology, Guangzhou University, Guangzhou 510006, China
    2 Guangdong Provincial Key Laboratory of Industrial Control System Security, Guangzhou 510006, China
    3 Huangpu Research School of Guangzhou University, Guangzhou 510000, China
  • Received:2025-06-12 Online:2026-06-10 Published:2026-07-27
  • Contact: LIU Yuan E-mail:yuanliu@gzhu.edu.cn

摘要:

在网络空间资产测绘与抗测绘领域,攻击者大多采用广度优先分批扫描策略探测在线活跃主机,进而筛选潜在攻击目标。防御者主要通过设定网络流量阈值,实现恶意扫描行为的检测识别。然而,网络环境具备动态多变特征,且攻击者可灵活调整扫描策略,导致传统静态防御策略的动态适配能力存在明显短板。为解决该问题,文章构建一种面向主机扫描与抗扫描的随机博弈模型,将连续分组扫描过程转化为马尔可夫决策过程,精准刻画攻防双方动态策略的演化机制,并采用改进非对称纳什学习算法求解近似均衡解。实验结果表明,在均衡条件下,文章所提动态策略的综合收益显著优于传统固定策略。

关键词: 主机存活探测, 分组扫描, 流量检测, 阈值设定, 随机博弈

Abstract:

In the field of cyber space mapping and anti-mapping, attackers mostly use the batch scanning strategy based on breadth-first search to detect online active hosts and identify potential malicious targets. Defenders mainly adopt the method of setting network traffic thresholds to identify malicious scanning behaviors. However, due to the dynamic and ever-changing nature of the network environment and the flexibility of opponents’ strategic adjustments, traditional static strategies have serious deficiencies in dynamic adaptability. To address this issue, this study designed a stochastic game model for host scanning and anti-scanning. It transformed the continuous grouped scanning process into a Markov decision process, precisely depicting the evolutionary mechanism of the dynamic strategies of both sides. By using the improved asymmetric Nash Q-learning algorithm, the approximate equilibrium solution was successfully obtained. Experimental results show that the dynamic strategy adopted under equilibrium conditions has significant advantages in terms of comprehensive benefits compared with traditional fixed strategies.

Key words: host liveness detection, packet scanning, traffic detection, threshold setting, stochastic game

中图分类号: