Abstract
We evaluated hierarchical repair for generated-code debugging under 10% fault depth. A deterministic paired simulation generated 56 cases and preserved a boundary-condition stratum. Mean test-pass rate changed from 0.576 to 0.622; the paired difference was +0.046 (95% interval +0.044 to +0.048). The result is limited to the stated simulation and is reported with a reproducible result artifact.
References
Shi, Y., Zhang, H., Wan, C., & Gu, X. (2025). Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers. 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 1628-1639. https://doi.org/10.1109/icse55347.2025.00005
Zhan, S., Wang, X., Wei, D., Cao, X., Lu, J., Wu, H., & Lin, Y. (2025). Unveiling the Power of Code Pre-trained Models in Neural Program Repair: A Systematic Review. https://doi.org/10.22541/au.174058397.77978477/v1
Ji, Z., Ma, P., Li, Z., Wang, Z., & Wang, S. (2025). Causality-Aided Evaluation and Explanation of Large Language Model-Based Code Generation. Proceedings of the ACM on Software Engineering, 2(ISSTA), 1374-1397. https://doi.org/10.1145/3728938
Bochenek, A., Protasiewicz, J. A., & Pedrycz, W. (2026). Large Language Model as a Code Generator in Simplified Domain-Specific Modelling. https://doi.org/10.2139/ssrn.6818764
Rahman, M. M., Watanobe, Y., & Nakamura, K. (2021). A Bidirectional LSTM Language Model for Code Evaluation and Repair. Symmetry, 13(2), 247. https://doi.org/10.3390/sym13020247
Mündler, N., He, J., Wang, H., Sen, K., Song, D., & Vechev, M. (2025). Type-Constrained Code Generation with Language Models. Proceedings of the ACM on Programming Languages, 9(PLDI), 601-626. https://doi.org/10.1145/3729274
Erkkinen, T. (2005). Model Style Guidelines for Production Code Generation. SAE Technical Paper Series, 1, 2005-01-1280. https://doi.org/10.4271/2005-01-1280
SEKER, S. E. (2024). Experiences and Challenges in AI-Driven Modular Software Development Using Large Language Models for Code Generation. https://doi.org/10.22541/au.172871465.54826063/v1
Cadenas, R. (2026). Convergent Correctness in Stochastic Code Generation: A Generate-Verify-Repair Architecture for Deterministic Validation of Autonomous AI Coding Agents in Regulated Environments. https://doi.org/10.2139/ssrn.6754899
Dr. Anup Bhange, Shivam Gautre, Nikhil Wandhare (2026). AI-Powered Code Helper for Intelligent Code Analysis, Debugging, and Multi-Language Execution. International Journal of Advanced Research in Science Communication and Technology, 1. https://doi.org/10.48175/ijarsct-32101
