Abstract
The reference set connects Quadratic exponential decrease roll-back: An efficient gradient update mechanism in proximal… with studies of Policy Gradient Reinforcement Learning for Parameterized Continuous-Time Optimal Control and Policy Gradient-based Reinforcement Learning for LQG Control with Chance Constraints, offering several competing ways to frame policy-gradient control. The analysis follows optimization dynamics, gradient stability, and policy ensembles, tracing points of convergence as well as differences in terminology, measurement, and experimental design. This makes hidden assumptions visible without introducing unsupported performance claims. The contribution is a structured reading of the existing evidence, not a new synthetic benchmark claim. It provides criteria for selecting methods, interpreting metrics, and designing follow-up studies that can be independently checked.
References
Zhao, R., Xu, D., Jian, S., Tan, T., Sun, X., & Zhang, W. (2023). Quadratic exponential decrease roll-back: An efficient gradient update mechanism in proximal policy optimization. In 2023 2nd International Conference on Machine Learning, Cloud Computing and Intelligent Mining (MLCCIM) (pp. 65-70). IEEE. https://doi.org/10.1109/MLCCIM60412.2023.00015
de Miguel Gomez, A., & Toosi, F. (2021). Continuous Parameter Control in Genetic Algorithms using Policy Gradient Reinforcement Learning. Proceedings of the 13th International Joint Conference on Computational Intelligence, 115-122. https://doi.org/10.5220/0010643500003063
wani, M.-A., Mohd., M., Khanday, H.-A., & Bhat, W.-A. (2025). Kl-Regularized Stochastic Policy Gradient for Stable Continuous Control in Off-Policy Reinforcement Learning. . https://doi.org/10.2139/ssrn.5357499
Yang, X., Zhang, H., & Wang, Z. (2021). Policy Gradient Reinforcement Learning for Parameterized Continuous-Time Optimal Control. 2021 33rd Chinese Control and Decision Conference (CCDC), 59-64. https://doi.org/10.1109/ccdc52312.2021.9602755
Liu, G., Chen, G., & Huang, V. (2023). Policy ensemble gradient for continuous control problems in deep reinforcement learning. Neurocomputing, 548, 126381. https://doi.org/10.1016/j.neucom.2023.126381
Sadeghi Tabas, S., & Samadi, V. (2022). Sustainable Reservoir Operation and Control Using a Deep Reinforcement Learning Policy Gradient Method. . https://doi.org/10.5194/egusphere-egu22-13011
Naha, A., & Dey, S. (2025). Policy Gradient-based Reinforcement Learning for LQG Control with Chance Constraints. 2025 European Control Conference (ECC), 364-371. https://doi.org/10.23919/ecc65951.2025.11186950
Bao, Y., & Velni, J.-M. (2021). Model-free Control Design Using Policy Gradient Reinforcement Learning in LPV Framework. 2021 European Control Conference (ECC), 150-155. https://doi.org/10.23919/ecc54610.2021.9655004
Denkert, R., Pham, H., & Warin, X. (2025). Control Randomisation Approach for Policy Gradient and Application to Reinforcement Learning in Optimal Switching. Applied Mathematics & Optimization, 91(1). https://doi.org/10.1007/s00245-024-10207-5
Liu, Y., Gou, L., Xue, Y., & Qiao, M. (2025). Optimization of Aero-engine Control Laws Based on Deep Deterministic Policy Gradient Reinforcement Learning Algorithm. 2025 16th International Conference on Mechanical and Aerospace Engineering (ICMAE), 116-121. https://doi.org/10.1109/icmae66341.2025.11277007
