arXiv · 2610.03506
Optimality of delayed reward attainment under uncertainty
Abstract
Decision making under uncertainty requires balancing the value of additional information against the cost of delaying the decision until that information becomes available. This challenge is especially prevalent for embodied decisions such as those faced by animals moving between food sources and robots selecting navigation targets, where movement through space shapes both the quality of information and the cost of acting. We consider a version of this problem, where the decision is on choosing among rewards of uncertain value, and those rewards can only be attained through changing the decision maker's state sufficiently. Our focus is on interactions between reward learning and state evolution, restated as a deterministic optimal control problem with two terminal alternatives, one with a known reward and the other learned gradually through noisy observations. We demonstrate that delayed reward attainment is often optimal in a variety of systems, especially so when the alternatives are initially difficult to distinguish and the cost of waiting is relatively low. For a subset of problems, we also derive sufficient conditions that allow shortening the planning horizon a priori, guaranteeing that further observations would not affect the value function. Finally, when observation quality depends on the state, we show that it is often optimal for a controller to evolve the state strategically toward more informative sensing locations before making a decision, in line with the behavior broadly observed in biological and engineered systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chuwei Wang, Alexander Vladimirsky, Anastasia Bizyaeva. 2026-10-02. Optimality of delayed reward attainment under uncertainty. https://arxiv.org/abs/2610.03506
Cite the original work for its findings. Save a collection to share your selection of sources.