nav emailalert searchbtn searchbox tablepage yinyongbenwen piczone journalimg journalInfo journalinfonormal searchdiv searchzone qikanlogo popupnotification paper paperNew
基于大语言模型的Web应用测试用例与脚本生成研究
基金项目(Foundation): 陕西高校青年创新团队——多模态大数据挖掘与融合创新团队
邮箱(Email): taomingshuo@xjtucc.edu.cn;
DOI: 10.13774/j.cnki.kjtb.2026.09.003
发布时间: 2026-06-23
出版时间: 2026-06-23
网络发布时间: 2026-06-23
移动端阅读
摘要:

随着软件系统复杂度的不断提升,传统测试用例生成方法主要依赖人工编写,存在效率低、覆盖不全、难以适应快速迭代等问题。大语言模型为自动化测试生成提供了新的可能,但其在真实场景下的有效性、稳定性与实用性尚未得到系统评估。本文针对该问题,选取ChatGPT、DeepSeek和豆包3种主流大语言模型,结合Selenium与Python构建自动化测试框架,在139邮箱、189邮箱、穷游网及马蜂窝等典型Web应用上开展实验。通过自动提取UI元素、生成测试数据与脚本,并结合人工审查与多维指标评估,系统分析各模型在测试数据有效性、脚本准确性、用例覆盖度与缺陷发现等方面的表现。实验结果表明,不同模型在测试生成中具有明显的能力分化:豆包在用例生成效率与中文场景通过率方面表现最优,ChatGPT在复杂逻辑与边界场景覆盖上更具优势,DeepSeek则在脚本代码质量与结构规范性上领先。研究进一步提出多模型协同与迭代优化策略,有效提升了生成脚本的可用性与场景适应性。本工作不仅验证了大语言模型在提升测试效率与覆盖范围方面的潜力,也揭示了其在输出一致性、复杂交互理解等方面的局限,为LLM(large language models)在软件测试中的实用化提供了实验依据与优化路径。

Abstract:

With the increasing complexity of software systems, traditional test case generation methods predominantly rely on manual creation, which suffers from inefficiency, insufficient coverage, and challenges in keeping pace with rapid iterations. Large language models(LLMs) offer new potential for automated test generation; however, their effectiveness, stability, and practicality in real-world scenarios have not yet been systematically evaluated. To address this issue, this study selects three mainstream LLMs—ChatGPT, DeepSeek, and Doubao—and constructs an automated testing framework using Selenium and Python. Experiments are conducted on typical web applications including 139 Mail, 189 Mail, Qyer, and Mafengwo. By automatically extracting UI elements, generating test data and scripts, and incorporating manual review with multi-dimensional metrics, the performance of each model is systematically analyzed in terms of test data validity, script accuracy, test case coverage, and defect detection capability. Experimental results indicate distinct capability differentiations among the models in test generation: Doubao performs best in terms of generation efficiency and pass rate in Chinese scenarios, ChatGPT demonstrates greater advantages in covering complex logic and edge cases, while DeepSeek leads in script code quality and structural standardization. Furthermore, this research proposes a multi-model collaboration and iterative optimization strategy, which effectively enhances the usability and scenario adaptability of generated scripts. This work not only validates the potential of LLMs in improving testing efficiency and coverage but also reveals limitations such as output inconsistency and inadequate understanding of complex interactions, providing experimental evidence and optimization pathways for the practical application of LLMs in software testing.

参考文献

[1]国家市场监督管理总局,国家标准化管理委员会.信息技术人工智能机器学习模型与系统的质量评测规范第1部分:质量模型与指标:GB/T 43441.1-2024[S].北京:中国标准出版社,2024.

[2]Wang Z Y,Shen L,Liu Z,et al. A survey on large language model-based software testing:challenges,advances,and future directions[J]. Journal of Systems and Software,2024,215:112045.

[3]Yu S C,Fang C R,Ling Y C,et al. LLM for test script generation and migration:challenges,capabilities,and opportunities[C]//2023 IEEE 23rd International Conference on Software Quality,Reliability,and Security(QRS),Chiang Mai,Thailand,2023:206-217.

[4]Liu Z,Chen C Y,Wang J J,et al. Make LLM a testing expert:bringing human-like interaction to mobile GUI testing via functionality-aware decisions[C]//Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. ACM,2024:1-13.

[5]Deng G,Liu Y,Mayoral-Vilches V,et al. PentestGPT:An LLM-empowered Automatic Penetration Testing Tool[C]//IEEE Conference on Dependable and Secure Computing(DSC). IEEE,2023:1-8.

[6]刘志强,陈晨,王静,等.基于大语言模型的GUI测试用例自动生成方法研究[J].计算机应用与软件,2023,40(7):1-8.

[7]王炎,刘嘉勇,刘亮,等.漏洞利用工具研发框架研究[J].计算机工程,2018,44(3):127-131.

[8]Xia C S,Zhang L. Keep the Conversation Going:Fixing162 out of 337 Bugs for0.42 Each Using ChatGPT[J].IEEE Transactions on Software Engineering,2024,50(5):1125-1143.

[9]Jin M,Shahriar S,Tufano M,et al. InferFix:end-to-end program repair with LLMs[C]//Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering,San Francisco CA USA,2023:1646-1656.

[10]Chen L,Zhu M,Lewis C,et al. How do large language models impact the user experience of code generation:a systematic literature review[J]. ACM Computing Surveys , 2024,56(10s):1-38.

[11]Zhou Y,Sun M,Yang Y,et al. Assessing the robustness of large language models in code generation to contextual perturbations[J]. Empirical Software Engineering,2024,29:123.

[12]Allamanis M,Barr E T,Ducasse S,et al. The last mile:domain-specific software engineering with large language models[J].IEEE Software,2024,41(1):65-73.

[13]Pearce H,Tan B,Ahsan S,et al. Can LLMs patch security issues[J].IEEE Security&Privacy,2024,22(2):64-73.

[14]Chen M,Tworek J,Jun H,et al. Evaluating large language models trained on code[J]. Communications of the ACM,2024,67(2):68-76.

[15]Wang Y,Le H,Gotmare A D,et al. CodeT5+:open code large language models for code understanding and generation[PP/OL]. V2. arXiv(2023-05-20)[2026-01-20]. https://doi.org/10.48550/arXiv.2305.07922.

[16]邸亮,杜永萍. LDA模型在微博用户推荐中的应用[J].计算机工程,2014,40(5):1-6,11.

[17]Lacombe G,Roy B,Achour S,et al. A prompt pattern catalog to enhance prompt engineering with large language models[J]. Software:Practice and Experience,2024,54(3):467-495.

[18]Fan A,Ghader H,Giró-I-Nieto X,et al. AnyTool:selfreflective LLM agents that leverage a toolset for enhanced code generation and problem solving[J].Nature Machine Intelligence,2024,6:354-366.

[19]Tian H,Lu Y,Sun W,et al. Selenium commander:an LLMpowered agent for autonomous web testing[C]//Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 2022:1-12.

[20]Pan R,Bavarlue S B,Stahl T,et al. ChatGPT vs.Lightweight Approaches for Unit Test Generation:A Comprehensive Comparative Study[J]. IEEE Transactions on Software Engineering,2024,50(6):1547-1568.

[21]Zhang T,Jiang Y,Monperrus M,et al. SyRAT:Synthesizing robust automated GUI tests with large language models[J]. ACM Transactions on Software Engineering and Methodology,2024,33(5):1-30.

[22]Le T H M,Shirzadeh M,Probst C W,et al. Quality,Ethics,and Liability:The trilemma of large language models in software engineering[J]. IEEE Software,2024,41(2):45-52.

[23]李明,张华,刘伟.基于人工智能的软件自动化测试技术研究进展[J].科技通报,2023,39(15):1-8.

[24]Wang X,Chen Y,Zhao Y,et al. Towards real-world web testing with LLMs:a benchmark and evaluation framework[J]. ACM Transactions on the Web,2026,20(1):1-31.

[25]IEEE. Multi-channel operation corrigendum:wireless access in vehicular environments(WAVE):IEEE Std1609.4-2020[S]. Washington D.C.,USA:IEEE Press,2020:1.

[26]Ferreira M,Viegas L,Faria J P,et al. Acceptance test generation with large language models:an industrial case study[C]//2025 IEEE/ACM International Conference on Automation of Software Test(AST),Ottawa,ON,Canada,2025:1-11.

[27]Khandaker S M,Kifetew F,Prandi D,et al. AugmenTest:enhancing tests with LLM-driven oracles[C]//2025IEEE Conference on Software Testing,Verification and Validation(ICST),Napoli,Italy,2025:279-289.

基本信息:

DOI:10.13774/j.cnki.kjtb.2026.09.003

中图分类号:TP311.53

引用信息:

[1]郭江坤,陶铭硕,翟雨晨,等.基于大语言模型的Web应用测试用例与脚本生成研究[J].科技通报().DOI:10.13774/j.cnki.kjtb.2026.09.003.

基金信息:

陕西高校青年创新团队——多模态大数据挖掘与融合创新团队

发布时间:

2026-06-23

出版时间:

2026-06-23

网络发布时间:

2026-06-23

检 索 高级检索

引用

GB/T 7714-2015 格式引文
MLA格式引文
APA格式引文