pisco_log
banner

Using Test Data to Improve Software Quality: Problems and Suggestions

Shan Wang

Abstract


Software quality depends not only on testing tools and execution procedures, but also on the data used to exercise functions, exception paths, and user scenarios. In many projects, test data is prepared late, reused without maintenance, or selected mainly for convenience. As
a result, testing may appear complete while important defects remain hidden. Based on recent studies on test suite quality, search-based test
generation, software test data management, synthetic test data, and requirement-based testing, this paper discusses how test data affects software quality. It analyzes problems in coverage, realism, traceability, maintenance, privacy protection, and defect feedback. It then proposes
suggestions including requirement- and risk-driven data design, boundary and negative data strengthening, test data repositories, privacyaware synthetic data, and defect-linked updates. The main argument is that test data should be treated as a managed quality asset rather than a
temporary input.

Keywords


Test data; Software quality; Software testing; Test data management; Quality assurance

Full Text:

PDF

Included Database


References


[1] Tran, H. K. V., Ali, N. B., Unterkalmsteiner, M., Borstler, J., & Chatzipetrou, P. (2025). Quality attributes of test cases and test suites:

Importance & challenges from practitioners' perspectives. Software Quality Journal, 33, 9. https://doi.org/10.1007/s11219-024-09698-w

[2] Fontes, A., Gay, G., & Feldt, R. (2025). Exploring the interaction of code coverage and non-coverage objectives in search-based test

generation. Software Testing, Verification and Reliability, 35, e70009. https://doi.org/10.1002/stvr.70009

[3] Ouedraogo, W. C., Plein, L., Kabore, K., Habib, A., Klein, J., Lo, D., & Bissyande, T. F. (2025). Enriching automatic test case generation by

extracting relevant test inputs from bug reports. Empirical Software Engineering, 30, 85. https://doi.org/10.1007/s10664-025-10635-z

[4] Tammisto, M.-A., Shah, F. A., Rodriguez, D., & Pfahl, D. (2025). The challenge of generating and evolving real-life like synthetic test

data without accessing real-world raw data: A systematic review. Expert Systems, 42, e70164. https://doi.org/10.1111/exsy.70164

[5] Gao, L., Qiu, J., & Chen, G. (2024). Software test data management based on knowledge graph. Informatica, 48(16), 27-36. https://doi.

org/10.31449/inf.v48i16.6416

[6] Ilays, I., Hafeez, Y., Almashfi, N., Ali, S., Humayun, M., Aqib, M., & Alwakid, G. (2024). Towards improving the quality of requirement

and testing process in agile software development: An empirical study. Computers, Materials & Continua. https://doi.org/10.32604/

cmc.2024.053830

[7] Tasarsu, M., Tokmak, A. V., & Catal, C. (2026). Test case generation using large language models: A systematic literature review. Cluster Computing, 29, 227. https://doi.org/10.1007/s10586-026-06021-z




DOI: http://dx.doi.org/10.70711/aitr.v4i3.9928

Refbacks

  • There are currently no refbacks.