arXiv · 1507.07147
True Online Emphatic TD($\lambda$): Quick Reference and Implementation Guide
Abstract
This document is a guide to the implementation of true online emphatic TD($\lambda$), a model-free temporal-difference algorithm for learning to make long-term predictions which combines the emphasis idea (Sutton, Mahmood & White 2015) and the true-online idea (van Seijen & Sutton 2014). The setting used here includes linear function approximation, the possibility of off-policy training, and all the generality of general value functions, as well as the emphasis algorithm's notion of "interest".
Explore related subjects
Keep this discovery
Richard S. Sutton. 2015-07-25. True Online Emphatic TD($\lambda$): Quick Reference and Implementation Guide. https://arxiv.org/abs/1507.07147
Cite the original work for its findings. Save a collection to share your selection of sources.