No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

正确性可读性单元测试计算机科学可用性质量（理念）考试（生物学）代码覆盖率发电机（电路理论）测试用例关键字驱动测试可靠性工程机器学习软件工程程序设计语言软件人机交互工程类软件开发物理古生物学软件建设哲学功率（物理）回归分析认识论量子力学生物

作者

Zhiqiang Yuan,Yiling Lou,Mingwei Liu,Shiji Ding,Kaixin Wang,Yixuan Chen,Xin Peng

出处

期刊：Cornell University - arXiv 日期：2023-01-01 被引量：36

链接

arxiv.org arxiv.org arxiv.org datacite.orgdoi.org

标识

DOI：10.48550/arxiv.2305.04207

摘要

Unit testing is essential in detecting bugs in functionally-discrete program units. Manually writing high-quality unit tests is time-consuming and laborious. Although traditional techniques can generate tests with reasonable coverage, they exhibit low readability and cannot be directly adopted by developers. Recent work has shown the large potential of large language models (LLMs) in unit test generation, which can generate more human-like and meaningful test code. ChatGPT, the latest LLM incorporating instruction tuning and reinforcement learning, has performed well in various domains. However, It remains unclear how effective ChatGPT is in unit test generation. In this work, we perform the first empirical study to evaluate ChatGPT's capability of unit test generation. Specifically, we conduct a quantitative analysis and a user study to systematically investigate the quality of its generated tests regarding the correctness, sufficiency, readability, and usability. The tests generated by ChatGPT still suffer from correctness issues, including diverse compilation errors and execution failures. Still, the passing tests generated by ChatGPT resemble manually-written tests by achieving comparable coverage, readability, and even sometimes developers' preference. Our findings indicate that generating unit tests with ChatGPT could be very promising if the correctness of its generated tests could be further improved. Inspired by our findings above, we propose ChatTESTER, a novel ChatGPT-based unit test generation approach, which leverages ChatGPT itself to improve the quality of its generated tests. ChatTESTER incorporates an initial test generator and an iterative test refiner. Our evaluation demonstrates the effectiveness of ChatTESTER by generating 34.3% more compilable tests and 18.7% more tests with correct assertions than the default ChatGPT.

求助该文献

No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation

今日热心研友