Stabilizing hierarchical repair for generated-code debugging under 10% fault depth: a threshold audit
PDF

Keywords

generated-code debugging
hierarchical repair
fault depth
paired simulation
reproducibility

Abstract

We evaluated hierarchical repair for generated-code debugging under 10% fault depth. A deterministic paired simulation generated 72 cases and preserved a median slice. Mean test-pass rate changed from 0.545 to 0.586; the paired difference was +0.041 (95% interval +0.038 to +0.043). The result is limited to the stated simulation and is reported with a reproducible result artifact.

PDF

References

Shi, Y., Zhang, H., Wan, C., & Gu, X. (2025). Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers. 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 1628-1639. https://doi.org/10.1109/icse55347.2025.00005

qiu, J., & li, S. (2023). A Multi-Encoder Model for Automatic Code Comment Generation. https://doi.org/10.21203/rs.3.rs-3218867/v1

Haase, R., Tischer, C., Hériché, J. K., & Scherf, N. (2024). Benchmarking Large Language Models for Bio-Image Analysis Code Generation. https://doi.org/10.1101/2024.04.19.590278

Wei, K. (2026). A Method for Alleviating Illusions in Code Generation Based on a Large Language Model Generated by Retrieval Enhancement. Advanced Electromagnetics, 15(3), 8519-8525. https://doi.org/10.7716/aem.v15i3.3976

-, S. A. P., -, D. R. R., -, V. H. P., & -, R. K. (2024). Benchmarking Large Language Models for Code Generation. International Journal For Multidisciplinary Research, 6(2), 17132. https://doi.org/10.36948/ijfmr.2024.v06i02.17132

Veeramreddygari, U. K. R. (2023). Generative AI for Software Engineering: Large Language Model-Driven Code Generation with Safety and Trust Assessment in Enterprise Development. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 569-582. https://doi.org/10.32628/cseit23906195

Konakanchi, M. S. K. (2025). Prompt-Oriented Code Understanding: Towards Natural Language-Driven Debugging in IDEs. INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT, 09(09), 1-9. https://doi.org/10.55041/ijsrem52844

Raveendra Reddy Pasala (2024). Revolutionizing Salesforce Development: Einstein Copilot for AI-Assisted Code Generation and Debugging. International Journal of Science and Research Archive, 13(2), 4205-4215. https://doi.org/10.30574/ijsra.2024.13.2.2146

Alecsandro Bacin, F., Adriano de Mello, B., Dondoni Salton, G., & Da Silva Feitosa, S. (2025). A Systematic Review about Large Language Models (LLMs) applied to Code Generation. Revista Brasileira de Computação Aplicada, 17(3), 1-13. https://doi.org/10.5335/rbca.v17i3.16310

Mündler, N., He, J., Wang, H., Sen, K., Song, D., & Vechev, M. (2025). Type-Constrained Code Generation with Language Models. Proceedings of the ACM on Programming Languages, 9(PLDI), 601-626. https://doi.org/10.1145/3729274