Rethinking Evaluation in Policy-Gradient Control through Optimization Dynamics and Policy Ensembles

Keywords

policy optimization
rollback update
gradient noise
synthetic control
reproducible experiment

Abstract

The reference set connects Quadratic exponential decrease roll-back: An efficient gradient update mechanism in proximal… with studies of Policy Gradient Reinforcement Learning for Parameterized Continuous-Time Optimal Control and Policy Gradient-based Reinforcement Learning for LQG Control with Chance Constraints, offering several competing ways to frame policy-gradient control. The analysis follows optimization dynamics, gradient stability, and policy ensembles, tracing points of convergence as well as differences in terminology, measurement, and experimental design. This makes hidden assumptions visible without introducing unsupported performance claims. The contribution is a structured reading of the existing evidence, not a new synthetic benchmark claim. It provides criteria for selecting methods, interpreting metrics, and designing follow-up studies that can be independently checked.

References

Zhao, R., Xu, D., Jian, S., Tan, T., Sun, X., & Zhang, W. (2023). Quadratic exponential decrease roll-back: An efficient gradient update mechanism in proximal policy optimization. In 2023 2nd International Conference on Machine Learning, Cloud Computing and Intelligent Mining (MLCCIM) (pp. 65-70). IEEE. https://doi.org/10.1109/MLCCIM60412.2023.00015

de Miguel Gomez, A., & Toosi, F. (2021). Continuous Parameter Control in Genetic Algorithms using Policy Gradient Reinforcement Learning. Proceedings of the 13th International Joint Conference on Computational Intelligence, 115-122. https://doi.org/10.5220/0010643500003063

wani, M.-A., Mohd., M., Khanday, H.-A., & Bhat, W.-A. (2025). Kl-Regularized Stochastic Policy Gradient for Stable Continuous Control in Off-Policy Reinforcement Learning. . https://doi.org/10.2139/ssrn.5357499

Yang, X., Zhang, H., & Wang, Z. (2021). Policy Gradient Reinforcement Learning for Parameterized Continuous-Time Optimal Control. 2021 33rd Chinese Control and Decision Conference (CCDC), 59-64. https://doi.org/10.1109/ccdc52312.2021.9602755

Liu, G., Chen, G., & Huang, V. (2023). Policy ensemble gradient for continuous control problems in deep reinforcement learning. Neurocomputing, 548, 126381. https://doi.org/10.1016/j.neucom.2023.126381

Sadeghi Tabas, S., & Samadi, V. (2022). Sustainable Reservoir Operation and Control Using a Deep Reinforcement Learning Policy Gradient Method. . https://doi.org/10.5194/egusphere-egu22-13011

Naha, A., & Dey, S. (2025). Policy Gradient-based Reinforcement Learning for LQG Control with Chance Constraints. 2025 European Control Conference (ECC), 364-371. https://doi.org/10.23919/ecc65951.2025.11186950

Bao, Y., & Velni, J.-M. (2021). Model-free Control Design Using Policy Gradient Reinforcement Learning in LPV Framework. 2021 European Control Conference (ECC), 150-155. https://doi.org/10.23919/ecc54610.2021.9655004

Denkert, R., Pham, H., & Warin, X. (2025). Control Randomisation Approach for Policy Gradient and Application to Reinforcement Learning in Optimal Switching. Applied Mathematics & Optimization, 91(1). https://doi.org/10.1007/s00245-024-10207-5

Liu, Y., Gou, L., Xue, Y., & Qiao, M. (2025). Optimization of Aero-engine Control Laws Based on Deep Deterministic Policy Gradient Reinforcement Learning Algorithm. 2025 16th International Conference on Mechanical and Aerospace Engineering (ICMAE), 116-121. https://doi.org/10.1109/icmae66341.2025.11277007