arXiv · 1706.08100
Specifying Non-Markovian Rewards in MDPs Using LDL on Finite Traces (Preliminary Version)
Abstract
In Markov Decision Processes (MDPs), the reward obtained in a state depends on the properties of the last state and action. This state dependency makes it difficult to reward more interesting long-term behaviors, such as always closing a door after it has been opened, or providing coffee only following a request. Extending MDPs to handle such non-Markovian reward function was the subject of two previous lines of work, both using variants of LTL to specify the reward function and then compiling the new model back into a Markovian model. Building upon recent progress in the theories of temporal logics over finite traces, we adopt LDLf for specifying non-Markovian rewards and provide an elegant automata construction for building a Markovian model, which extends that of previous work and offers strong minimality and compositionality guarantees.
Explore related subjects
Keep this discovery
Ronen Brafman, Giuseppe De Giacomo, Fabio Patrizi. 2017-06-25. Specifying Non-Markovian Rewards in MDPs Using LDL on Finite Traces (Preliminary Version). https://arxiv.org/abs/1706.08100
Cite the original work for its findings. Save a collection to share your selection of sources.