UGTC: Uncertainty-Gated Temporal Credit --- A Modular Advantage Estimator for Actor-Critic RL
Keywords:
credit assignment, uncertainty estimation, reinforcement learning, temporal credit, adaptive lambda, ensemble methodsAbstract
The Generalized Advantage Estimator fixes a single bias--variance tradeoff ($\lambda$) for every state and training phase. We introduce UGTC (Uncertainty-Gated Temporal Credit), a plug-in module that replaces the advantage estimator with a state-dependent blend of a fast critic ($\lambda=0.80$) and a slow critic ensemble ($\lambda=0.99$, $M=3$), gated by ensemble disagreement. UGTC composes with PPO, TD3, SAC, and DreamerV3 via backbone-specific insertion points. Across six benchmarks (64 tasks, 10 seeds, bootstrap 95% CIs), UGTC-PPO converges $2.7\times$ faster on Hopper ($1.9\times$ wall-clock); UGTC-DreamerV3 peaks at $466.7$ vs. $191.8$ on MetaWorld ML45 ($2.4\times$); UGTC-PPO improves Procgen Hard by $+19.9\%$. We report two negative results---Crafter ($-23.3\%$) and Ant peak return ($-14\%$ vs. TD3)---with deep analysis showing the Ant regression stems from persistent conservative bias in high-dimensional state spaces with spatially uniform critic reliability. A systematic ablation including meta-gradient $\lambda$, REDQ, a learned gate, and a per-state $\lambda$ network confirms that calibrated epistemic uncertainty is the critical ingredient. A held-out sensitivity analysis on Procgen and ML45 validates that defaults transfer without retuning. We derive a practical applicability criterion (reward density $\times$ critic reliability variation, $R^2=0.85$) predicting when uncertainty-gated credit helps versus hurts.
References
Anschel, O., Baram, N., & Shimkin, N. (2017). Averaged-DQN: Variance reduction and stabilization for deep RL. In Proc. of ICML.
Bellemare, M. G., Dabney, W., & Munos, R. (2017). A distributional perspective on RL. In Proc. of ICML.
Chen, X., Wang, C., Zhou, Z., & Ross, K. (2021). Randomized ensembled double Q-learning. In Proc. of ICLR.
Dabney, W., Rowland, M., Bellemare, M. G., & Munos, R. (2018). Distributional RL with quantile regression. In Proc. of AAAI.
Fujimoto, S., Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods. In Proc. of ICML.
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic. In Proc. of ICML.
Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering diverse domains through world models. arXiv:2301.04104.
Kuznetsov, A., Shvechikov, P., Grishin, A., & Vetrov, D. (2020). Controlling overestimation bias with TQC. In Proc. of ICML.
Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Predictive uncertainty estimation using deep ensembles. In Proc. of NeurIPS.
Osband, I., Blundell, C., Pritzel, A., & Van Roy, B. (2016). Deep exploration via bootstrapped DQN. In Proc. of NeurIPS.
Peng, J., & Williams, R. J. (1996). Incremental multi-step Q-learning. Machine Learning, 22, 283–290.
Schulman, J., Moritz, P., Levine, S., Jordan, M., & Abbeel, P. (2016). High-dimensional continuous control using GAE. arXiv:1506.02438.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization. arXiv:1707.06347.
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
Tsitsiklis, J. N., & Van Roy, B. (1997). Analysis of TD learning with function approximation. IEEE TAC, 42(5), 674–690.
Xu, Z., van Hasselt, H., & Silver, D. (2018). Meta-gradient reinforcement learning. In Proc. of NeurIPS.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Yagiz Ekrem Dalar, Ömer F. AKSOY, Mutlu N. Sezer, Feyzi A. Salihoğlu, Ahmet Rifat Öztürk (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.