Evidence Alignment and Transfer Boundaries in Ai-Assisted Programming
PDF

Keywords

Ai-Assisted Programming
Programmer Behavior
Error Localization
Execution Feedback
Repair Hierarchy
Evaluation Leakage

Abstract

Progress in AI-assisted programming depends on more than accumulating favorable results. This critical synthesis connects comparative analysis of observable coding patterns produced by people and machines with hierarchical debugging that closes the gap between generated code and executable correctness and asks how measurement choices, boundary conditions, and decision costs shape the interpretation of both. A structured reading of two target studies and 12 verified companion references is conducted across five lenses: programmer behavior, error localization, execution feedback, repair hierarchy, evaluation leakage. Emphasis is placed on the provenance of evidence, the comparability of baselines, and the consequences of alternative explanations. The combined literature indicates that methodological gains become actionable only when programmer behavior and error localization are evaluated together and when limits associated with evaluation leakage are explicit. This shifts the emphasis from isolated scores toward traceable chains of evidence and decision relevance. The article concludes with a research agenda built around transparent comparators, targeted stress tests, and evidence records that can be reused without overstating causal or practical reach.

PDF

References

Shi, Y., Zhang, H., Wan, C., & Gu, X. (2025). Between lines of code: Unraveling the distinct patterns of machine and human programmers. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) (pp. 1628-1639). IEEE.

Shi, Y., Wang, S., Wan, C., Wang, M., & Gu, X. (2024). From code to correctness: Closing the last mile of code generation with hierarchical debugging. arXiv preprint arXiv:2410.01215.

Timofeev, A. N., & Mikhaylova, S. S. (2024). Approach to improving the quality of program code generation by large language models. Neurocomputers. https://doi.org/10.18127/j19998554-202404-02

Gülmez, B. (2026). Code generation with large language models: a survey from neural program synthesis to autonomous software development. Applied Intelligence, 56(6). https://doi.org/10.1007/s10489-026-07230-0

Hamasaki, I., Kunimi, K., Shibata, K., & Toshiaki, G. (2026). Large Language Model-Based Interactive Code Generation for Developing a 3D Eye Movement Schematic. Cureus. https://doi.org/10.7759/cureus.107791

Aytekin, M. C., Yılmaz, F. G., & Demirezen, M. U. (2026). Automating code generation for a new ecosystem: establishing baselines with large language model based code generation for ArkTS and HarmonyOS. Automated Software Engineering, 33(2). https://doi.org/10.1007/s10515-026-00599-9

Kevin Gwindingwi,, & Monica Gondo, (2025). A Model For Automated Code Debugging Using Small Language Models. Journal of Scientific Research and Technology, 86-92. https://doi.org/10.61808/jsrt230

Jiang, R., Xia, K., Huang, J., & Lu, J. (2026). Large Language Model-Based Method for HVAC System Control Code Automatic Generation. Buildings, 16(9), 1722. https://doi.org/10.3390/buildings16091722

Kang, S., Chen, B., Yoo, S., & Lou, J. G. (2024). Explainable automated debugging via large language model-driven scientific debugging. Empirical Software Engineering, 30(2). https://doi.org/10.1007/s10664-024-10594-x

Bistarelli, S., Fiore, M., Mercanti, I., & Mongiello, M. (2025). Usage of Large Language Model for Code Generation Tasks: A Review. SN Computer Science, 6(6). https://doi.org/10.1007/s42979-025-04241-5

Wei, K. (2026). A Method for Alleviating Illusions in Code Generation Based on a Large Language Model Generated by Retrieval Enhancement. Advanced Electromagnetics, 15(3), 8519-8525. https://doi.org/10.7716/aem.v15i3.3976

Hemberg, E., Moskal, S., & O’Reilly, U. M. (2024). Evolving code with a large language model. Genetic Programming and Evolvable Machines, 25(2). https://doi.org/10.1007/s10710-024-09494-2

Andruccioli, M., Delnevo, G., Mirri, S., & Salomoni, P. (2026). PromptTone: A Dataset for Evaluating Large Language Model Code Generation Under Varying Prompt Politeness Levels. Data, 11(4), 88. https://doi.org/10.3390/data11040088

Li, S., Xie, K., Li, Y., Li, H., Ren, Y., Sun, L., et al. (2025). TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability. IEEE Transactions on Software Engineering, 51(8), 2396-2411. https://doi.org/10.1109/tse.2025.3584774