Abstract
This review examines a shared methodological problem in reinforcement learning for software reasoning and automated program repair: how evidence from multi-agent chain-of-draft reasoning optimized with reinforcement learning can be placed in analytical dialogue with execution-grounded reinforcement learning with sequence- and line-level reward models without erasing differences in scale, assumptions, or intended use. The review draws on two focal records and 12 established sources already present in the project evidence cache. Its comparative framework links draft coordination, curriculum design, and lineage graphs to downstream questions of reward hacking and generalization. Comparison reveals recurring trade-offs among draft coordination, curriculum design, and lineage graphs. These trade-offs do not support a universal ranking; instead, they identify the operating envelope within which each method remains credible and the perturbations most likely to expose fragile conclusions. The article concludes with a research agenda built around transparent comparators, targeted stress tests, and evidence records that can be reused without overstating causal or practical reach.
References
Li, Y., Liu, M., Wang, H., Zhang, Y., Ma, Y., & Tan, W. (2026). DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs. Proceedings of the AAAI Conference on Artificial Intelligence, 40(35), 29530-29537.
Li, Y., Wang, H., Shang, X., Tang, X., Cao, Y., & Chen, X. (2026). BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models. arXiv preprint arXiv:2605.09134.
XIAO, Z., & ZHANG, S. Y. (2009). Reinforcement Learning Model Based on Regret for Multi-Agent Conflict Games. Journal of Software, 19(11), 2957-2967. https://doi.org/10.3724/sp.j.1001.2008.02957
Hanna, C., Blot, A., & Petke, J. (2025). Reinforcement learning for mutation operator selection in automated program repair. Automated Software Engineering, 32(2). https://doi.org/10.1007/s10515-025-00501-z
Akgün, O. (2026). Stabilizing independent multi-agent reinforcement learning via curriculum-based iterative self-play. Neurocomputing, 704, 134819. https://doi.org/10.1016/j.neucom.2026.134819
Kumar Karne, V., Noone Srinivas,, Nagaraj Mandaloju,, & Parameshwar Reddy Kothamali, (2020). Reinforcement Learning for Optimizing Test Case Execution in Automated Testing. Innovative Research Thoughts, 6(3), 13-27. https://doi.org/10.36676/irt.v6.i3.1494
Bai, L., Chen, M., & Xiao, Q. (2024). Multi-hop temporal knowledge graph reasoning with multi-agent reinforcement learning. Applied Soft Computing, 160, 111727. https://doi.org/10.1016/j.asoc.2024.111727
Hao, S., Shi, X., Liu, H., Yin, Y., & Chen, X. (2026). Template-guided interpretable reasoning with execution feedback for LLM-based program repair. Information and Software Technology, 193, 108058. https://doi.org/10.1016/j.infsof.2026.108058
Rusu, E., & Glatt, R. (2021). Abmarl: Connecting Agent-Based Simulations with Multi-Agent Reinforcement Learning. Journal of Open Source Software, 6(64), 3424. https://doi.org/10.21105/joss.03424
Wan, H., Luo, H., Li, M., & Luo, X. (2024). Automated Program Repair for Introductory Programming Assignments. IEEE Transactions on Learning Technologies, 17, 1705-1720. https://doi.org/10.1109/tlt.2024.3403710
Zhang, X., Li, Z., Quan, X., Cheng, K., & Yu, Y. (2026). Curriculum-Learning-Guided Multi-Agent Deep Reinforcement Learning for N-1 Static Security Prevention and Control. Energy Engineering, 123(9), 1-10. https://doi.org/10.32604/ee.2025.073912
Yin, Z., Lin, W., & Kong, X. (2026). Heterogeneous multi-expert collaborative reinforcement learning for automated CAD program synthesis from engineering drawings. Discover Artificial Intelligence. https://doi.org/10.1007/s44163-026-01731-0
Morshed, M., & Zaman Chowdhury, M. (2026). Curriculum-assisted multi-agent reinforcement learning for scalable V2X resource allocation. Physical Communication, 76, 103062. https://doi.org/10.1016/j.phycom.2026.103062
Jha, A. C. (2025). Automated Firewall Policy Generation with Reinforcement Learning. International journal of IoT, 5(1), 190-211. https://doi.org/10.55640/ijiot-05-01-10
