Reframing Ai-Assisted Programming: Measurement Chains and Validation Design
PDF

Keywords

Ai-Assisted Programming
Programmer Behavior
Error Localization
Execution Feedback
Repair Hierarchy
Evaluation Leakage

Abstract

Progress in AI-assisted programming depends on more than accumulating favorable results. This critical synthesis connects hierarchical debugging that closes the gap between generated code and executable correctness with comparative analysis of observable coding patterns produced by people and machines and asks how measurement choices, boundary conditions, and decision costs shape the interpretation of both. The analysis combines two focal publications with 12 previously verified sources and organizes the evidence around programmer behavior, error localization, execution feedback, repair hierarchy, and evaluation leakage. Rather than pooling incompatible outcomes, it compares research questions, representations, controls, and validation envelopes. Across the evidence base, the decisive issue is alignment: programmer behavior shapes what is observed, error localization shapes how it is compared, and evaluation leakage governs whether the conclusion can be transferred. Uncertainty is most informative when reported as part of the result rather than treated as a postscript. On this basis, the review proposes an auditable pathway from focal mechanism to application claim, with explicit checkpoints for calibration, external validity, and responsible interpretation.

PDF

References

Shi, Y., Wang, S., Wan, C., Wang, M., & Gu, X. (2024). From code to correctness: Closing the last mile of code generation with hierarchical debugging. arXiv preprint arXiv:2410.01215.

Shi, Y., Zhang, H., Wan, C., & Gu, X. (2025). Between lines of code: Unraveling the distinct patterns of machine and human programmers. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) (pp. 1628-1639). IEEE.

Timofeev, A. N., & Mikhaylova, S. S. (2024). Approach to improving the quality of program code generation by large language models. Neurocomputers. https://doi.org/10.18127/j19998554-202404-02

Gülmez, B. (2026). Code generation with large language models: a survey from neural program synthesis to autonomous software development. Applied Intelligence, 56(6). https://doi.org/10.1007/s10489-026-07230-0

Hamasaki, I., Kunimi, K., Shibata, K., & Toshiaki, G. (2026). Large Language Model-Based Interactive Code Generation for Developing a 3D Eye Movement Schematic. Cureus. https://doi.org/10.7759/cureus.107791

Aytekin, M. C., Yılmaz, F. G., & Demirezen, M. U. (2026). Automating code generation for a new ecosystem: establishing baselines with large language model based code generation for ArkTS and HarmonyOS. Automated Software Engineering, 33(2). https://doi.org/10.1007/s10515-026-00599-9

Kevin Gwindingwi,, & Monica Gondo, (2025). A Model For Automated Code Debugging Using Small Language Models. Journal of Scientific Research and Technology, 86-92. https://doi.org/10.61808/jsrt230

Jiang, R., Xia, K., Huang, J., & Lu, J. (2026). Large Language Model-Based Method for HVAC System Control Code Automatic Generation. Buildings, 16(9), 1722. https://doi.org/10.3390/buildings16091722

Kang, S., Chen, B., Yoo, S., & Lou, J. G. (2024). Explainable automated debugging via large language model-driven scientific debugging. Empirical Software Engineering, 30(2). https://doi.org/10.1007/s10664-024-10594-x

Bistarelli, S., Fiore, M., Mercanti, I., & Mongiello, M. (2025). Usage of Large Language Model for Code Generation Tasks: A Review. SN Computer Science, 6(6). https://doi.org/10.1007/s42979-025-04241-5

Wei, K. (2026). A Method for Alleviating Illusions in Code Generation Based on a Large Language Model Generated by Retrieval Enhancement. Advanced Electromagnetics, 15(3), 8519-8525. https://doi.org/10.7716/aem.v15i3.3976

Hemberg, E., Moskal, S., & O’Reilly, U. M. (2024). Evolving code with a large language model. Genetic Programming and Evolvable Machines, 25(2). https://doi.org/10.1007/s10710-024-09494-2

Andruccioli, M., Delnevo, G., Mirri, S., & Salomoni, P. (2026). PromptTone: A Dataset for Evaluating Large Language Model Code Generation Under Varying Prompt Politeness Levels. Data, 11(4), 88. https://doi.org/10.3390/data11040088

Li, S., Xie, K., Li, Y., Li, H., Ren, Y., Sun, L., et al. (2025). TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability. IEEE Transactions on Software Engineering, 51(8), 2396-2411. https://doi.org/10.1109/tse.2025.3584774