Interpreting the Policy-Gradient Control Evidence Base: Reproducible Evaluation, Stochastic Perturbations, and Model-Free Control

Keywords

policy optimization
rollback update
gradient noise
synthetic control
reproducible experiment

Abstract

The evidence base for policy-gradient control spans Kl-Regularized Stochastic Policy Gradient for Stable Continuous Control in Off-Policy Reinforcement…, while related work on Sustainable Reservoir Operation and Control Using a Deep Reinforcement Learning Policy… and Control Randomisation Approach for Policy Gradient and Application to Reinforcement Learning… broadens the methodological context. The discussion uses reproducible evaluation, stochastic perturbations, and model-free control as analytical lenses. Rather than ranking reported results, it examines which claims remain comparable across tasks, datasets, and operating conditions. By aligning terminology and evidence requirements, the article offers a more defensible basis for future empirical work. The final recommendations focus on traceable data, bounded claims, and evaluation under meaningful operating conditions.

References

Zhao, R., Xu, D., Jian, S., Tan, T., Sun, X., & Zhang, W. (2023). Quadratic exponential decrease roll-back: An efficient gradient update mechanism in proximal policy optimization. In 2023 2nd International Conference on Machine Learning, Cloud Computing and Intelligent Mining (MLCCIM) (pp. 65-70). IEEE. https://doi.org/10.1109/MLCCIM60412.2023.00015

de Miguel Gomez, A., & Toosi, F. (2021). Continuous Parameter Control in Genetic Algorithms using Policy Gradient Reinforcement Learning. Proceedings of the 13th International Joint Conference on Computational Intelligence, 115-122. https://doi.org/10.5220/0010643500003063

wani, M.-A., Mohd., M., Khanday, H.-A., & Bhat, W.-A. (2025). Kl-Regularized Stochastic Policy Gradient for Stable Continuous Control in Off-Policy Reinforcement Learning. . https://doi.org/10.2139/ssrn.5357499

Yang, X., Zhang, H., & Wang, Z. (2021). Policy Gradient Reinforcement Learning for Parameterized Continuous-Time Optimal Control. 2021 33rd Chinese Control and Decision Conference (CCDC), 59-64. https://doi.org/10.1109/ccdc52312.2021.9602755

Liu, G., Chen, G., & Huang, V. (2023). Policy ensemble gradient for continuous control problems in deep reinforcement learning. Neurocomputing, 548, 126381. https://doi.org/10.1016/j.neucom.2023.126381

Sadeghi Tabas, S., & Samadi, V. (2022). Sustainable Reservoir Operation and Control Using a Deep Reinforcement Learning Policy Gradient Method. . https://doi.org/10.5194/egusphere-egu22-13011

Naha, A., & Dey, S. (2025). Policy Gradient-based Reinforcement Learning for LQG Control with Chance Constraints. 2025 European Control Conference (ECC), 364-371. https://doi.org/10.23919/ecc65951.2025.11186950

Bao, Y., & Velni, J.-M. (2021). Model-free Control Design Using Policy Gradient Reinforcement Learning in LPV Framework. 2021 European Control Conference (ECC), 150-155. https://doi.org/10.23919/ecc54610.2021.9655004

Denkert, R., Pham, H., & Warin, X. (2025). Control Randomisation Approach for Policy Gradient and Application to Reinforcement Learning in Optimal Switching. Applied Mathematics & Optimization, 91(1). https://doi.org/10.1007/s00245-024-10207-5

Liu, Y., Gou, L., Xue, Y., & Qiao, M. (2025). Optimization of Aero-engine Control Laws Based on Deep Deterministic Policy Gradient Reinforcement Learning Algorithm. 2025 16th International Conference on Mechanical and Aerospace Engineering (ICMAE), 116-121. https://doi.org/10.1109/icmae66341.2025.11277007