UGTC: Uncertainty-Gated Temporal Credit --- A Modular Advantage Estimator for Actor-Critic RL

Authors

Keywords:

credit assignment, uncertainty estimation, reinforcement learning, temporal credit, adaptive lambda, ensemble methods

Abstract

The Generalized Advantage Estimator fixes a single bias--variance tradeoff ($\lambda$) for every state and training phase. We introduce UGTC (Uncertainty-Gated Temporal Credit), a plug-in module that replaces the advantage estimator with a state-dependent blend of a fast critic ($\lambda=0.80$) and a slow critic ensemble ($\lambda=0.99$, $M=3$), gated by ensemble disagreement. UGTC composes with PPO, TD3, SAC, and DreamerV3 via backbone-specific insertion points. Across six benchmarks (64 tasks, 10 seeds, bootstrap 95% CIs), UGTC-PPO converges $2.7\times$ faster on Hopper ($1.9\times$ wall-clock); UGTC-DreamerV3 peaks at $466.7$ vs. $191.8$ on MetaWorld ML45 ($2.4\times$); UGTC-PPO improves Procgen Hard by $+19.9\%$. We report two negative results---Crafter ($-23.3\%$) and Ant peak return ($-14\%$ vs. TD3)---with deep analysis showing the Ant regression stems from persistent conservative bias in high-dimensional state spaces with spatially uniform critic reliability. A systematic ablation including meta-gradient $\lambda$, REDQ, a learned gate, and a per-state $\lambda$ network confirms that calibrated epistemic uncertainty is the critical ingredient. A held-out sensitivity analysis on Procgen and ML45 validates that defaults transfer without retuning. We derive a practical applicability criterion (reward density $\times$ critic reliability variation, $R^2=0.85$) predicting when uncertainty-gated credit helps versus hurts.

References

Anschel, O., Baram, N., & Shimkin, N. (2017). Averaged-DQN: Variance reduction and stabilization for deep RL. In Proc. of ICML.

Bellemare, M. G., Dabney, W., & Munos, R. (2017). A distributional perspective on RL. In Proc. of ICML.

Chen, X., Wang, C., Zhou, Z., & Ross, K. (2021). Randomized ensembled double Q-learning. In Proc. of ICLR.

Dabney, W., Rowland, M., Bellemare, M. G., & Munos, R. (2018). Distributional RL with quantile regression. In Proc. of AAAI.

Fujimoto, S., Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods. In Proc. of ICML.

Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic. In Proc. of ICML.

Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering diverse domains through world models. arXiv:2301.04104.

Kuznetsov, A., Shvechikov, P., Grishin, A., & Vetrov, D. (2020). Controlling overestimation bias with TQC. In Proc. of ICML.

Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Predictive uncertainty estimation using deep ensembles. In Proc. of NeurIPS.

Osband, I., Blundell, C., Pritzel, A., & Van Roy, B. (2016). Deep exploration via bootstrapped DQN. In Proc. of NeurIPS.

Peng, J., & Williams, R. J. (1996). Incremental multi-step Q-learning. Machine Learning, 22, 283–290.

Schulman, J., Moritz, P., Levine, S., Jordan, M., & Abbeel, P. (2016). High-dimensional continuous control using GAE. arXiv:1506.02438.

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization. arXiv:1707.06347.

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Tsitsiklis, J. N., & Van Roy, B. (1997). Analysis of TD learning with function approximation. IEEE TAC, 42(5), 674–690.

Xu, Z., van Hasselt, H., & Silver, D. (2018). Meta-gradient reinforcement learning. In Proc. of NeurIPS.

Downloads

Published

09.10.2026

How to Cite

UGTC: Uncertainty-Gated Temporal Credit --- A Modular Advantage Estimator for Actor-Critic RL. (2026). Ulysseus Young Explorers in Science, 1(1), 1-11. https://ojs.website.tuke.sk/TUKE/article/view/33