arXiv · 2503.05042
Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning
Abstract
Automata-conditioned reinforcement learning (RL) has given promising results for learning multi-task policies capable of performing temporally extended objectives given at runtime, done by pretraining and freezing automata embeddings prior to training the downstream policy. However, no theoretical guarantees were given. This work provides a theoretical framework for the automata-conditioned RL problem and shows that it is probably approximately correct learnable. We then present a technique for learning provably correct automata embeddings, guaranteeing optimal multi-task policy learning. Our experimental evaluation confirms these theoretical results.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Beyazit Yalcinkaya, Niklas Lauffer, Marcell Vazquez-Chanlatte, Sanjit A. Seshia. 2025-03-06. Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning. https://arxiv.org/abs/2503.05042
Cite the original work for its findings. Save a collection to share your selection of sources.