arXiv · 1911.03107
Intelligent buses in a loop service: Emergence of no-boarding and holding strategies
Abstract
We study how $N$ intelligent buses serving a loop of $M$ bus stops learn a \emph{no-boarding strategy} and a \emph{holding strategy} by reinforcement learning. The high level no-boarding and holding strategies emerge from the low level actions of \emph{stay} or \emph{leave} when a bus is at a bus stop and everyone who wishes to alight has done so. A reward that encourages the buses to strive towards a staggered phase difference amongst them whilst picking up people allows the reinforcement learning process to converge to an optimal Q-table within a reasonable amount of simulation time. It is remarkable that this emergent behaviour of intelligent buses turns out to minimise the average waiting time of commuters, in various setups where buses have identical natural frequency, or different natural frequencies during busy as well as lull periods. Cooperative actions are also observed, e.g. the buses learn to \emph{unbunch}.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vee-Liem Saw, Luca Vismara, Lock Yue Chew. 2019-11-08. Intelligent buses in a loop service: Emergence of no-boarding and holding strategies. https://arxiv.org/abs/1911.03107
Cite the original work for its findings. Save a collection to share your selection of sources.