SearcharxivSearch

arXiv subjects

Michael Hunter

Publications and source records attributed to Michael Hunter.

8 recordsLinked to original sources

Explainable Reinforcement Learning for Adaptive Traffic Signal Control

Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastructure like traffic control, the opaque, black-box nature of deep RL models poses challenges for transportation agency acceptance, regulatory compliance, operational trust, troubleshooting, and fine-tuning. To bridge this gap between high-performance optimization and human-comprehensible interpretability, this effort introduces a novel, explainable entity centric RL framework for safe and transparent traffic signal control. Rather than processing traffic states through monolithic, flat vectors, the proposed architecture disaggregates real-time intersection observations into distinct, high-dimensional lane entities and phase temporal configurations to inherently preserve the structural topology and geometric configurations of the intersection. Relational dependencies and inter-lane conflicts are dynamically extracted via a dual-stage attention network featuring sequential multi-head cross-attention and self-attention blocks. This design yields a real time affinity matrix that quantifies the direct influence of signal phases on specific approach volumes and queues, providing full visual and analytical interpretability. To ensure strict operational reliability, a deterministic action-masking interface is integrated directly into the Proximal Policy Optimization pipeline, explicitly blocking invalid phase transitions to guarantee absolute compliance with established signal timing and safety constraints. Evaluated in a microscopic simulation environment, outperforms state-of-the-art baselines in delay minimization. More importantly, the emergent attention weights align precisely with established traffic engineering principles, offering an auditable, trust-enabling, and deployable architecture for next-generation adaptive traffic control systems.

cs.AI

Emergency Vehicle Preemption Strategies using Machine Learning to Optimize Traffic Operations

Emergency response vehicles (ERVs), such as fire trucks, operate to save lives and mitigate property damage. Emergency vehicle preemption (EVP) is typically implemented to provide the right-of-way to ERVs by giving green signals as they approach signalized intersections along their routes. EVP operations are usually optimized to minimize ERV delay. This study seeks to reduce delay experienced by other vehicles in the network while keeping ERV travel time near its optimum. A machine learning-based EVP strategy, termed MLEVP, is developed to determine EVP trigger times at multiple downstream intersections using real-time sensor data, including vehicle detections, signal indications, and ERV location. MLEVP proactively clears downstream traffic queues to reduce ERV response time while limiting delay on conflicting traffic movements. In the case study, MLEVP is developed using a calibrated microscopic simulation of a signalized corridor testbed in PTV Vissim. The EVP problem is formulated as a regression problem and solved using machine learning models trained on data generated from the simulation. Results demonstrate that the proposed algorithm can produce near-optimal ERV travel times while minimizing impacts on conflicting traffic.

cs.CE

Evaluating the Robustness of Reinforcement Learning based Adaptive Traffic Signal Control

Reinforcement learning (RL) has attracted increasing interest for adaptive traffic signal control due to its model-free ability to learn control policies directly from interaction with the traffic environment. However, several challenges remain before RL-based signal control can be considered ready for field deployment. Many existing studies rely on simplified signal timing structures, robustness of trained models under varying traffic demand conditions remains insufficiently evaluated, and runtime efficiency continues to pose challenges when training RL algorithms in traffic microscopic simulation environments. This study formulates an RL-based signal control algorithm capable of representing a full eight-phase ring-barrier configuration consistent with field signal controllers. The algorithm is trained and evaluated under varying traffic demand conditions and benchmarked against state-of-the-practice actuated signal control (ASC). To assess robustness, experiments are conducted across multiple traffic volumes and origin-destination (O-D) demand patterns with varying levels of structural similarity. To improve training efficiency, a distributed asynchronous training architecture is implemented that enables parallel simulation across multiple computing nodes. Results from a case study intersection show that the proposed RL-based signal control significantly outperforms optimized ASC, reducing average delay by 11-32% across movements. A model trained on a single O-D pattern generalizes well to similar unseen demand patterns but degrades under substantially different demand conditions. In contrast, a model trained on diverse O-D patterns demonstrates strong robustness, consistently outperforming ASC even under highly dissimilar unseen demand scenarios.

cs.LG

Adaptive Traffic Signal Control based on Multi-Agent Reinforcement Learning. Case Study on a simulated real-world corridor

Previous studies that have formulated multi-agent reinforcement learning (RL) algorithms for adaptive traffic signal control have primarily used value-based RL methods. However, recent literature has shown that policy-based methods may perform better in partially observable environments. Additionally, RL methods remain largely untested for real-world normally signal timing plans because of the simplifying assumptions common in the literature. The current study attempts to address these gaps and formulates a multi-agent proximal policy optimization (MA-PPO) algorithm to implement adaptive and coordinated traffic control along an arterial corridor. The formulated MA-PPO has a centralized-critic architecture under a centralized training and decentralized execution framework. Agents are designed to allow selection and implementation of up to eight signal phases, as commonly implemented in field controllers. The formulated algorithm is tested on a simulated real-world seven intersection corridor. The speed of convergence for each agent was found to depend on the size of the action space, which depends on the number and sequence of signal phases. The performance of the formulated MA-PPO adaptive control algorithm is compared with the field implemented actuated-coordinated signal control (ASC), modeled using PTV-Vissim-MaxTime software in the loop simulation (SILs). The trained MA-PPO performed significantly better than the ASC for all movements. Compared to ASC the MA-PPO showed 2% and 24% improvements in travel time in the primary and secondary coordination directions, respectively. For cross streets movements MA-PPO also showed significant crossing time reductions. Volume sensitivity experiments revealed that the formulated MA-PPO demonstrated good stability, robustness, and adaptability to changes in traffic demand.

cs.MA

Integrating Transit Signal Priority into Multi-Agent Reinforcement Learning based Traffic Signal Control

This study integrates Transit Signal Priority (TSP) into multi-agent reinforcement learning (MARL) based traffic signal control. The first part of the study develops adaptive signal control based on MARL for a pair of coordinated intersections in a microscopic simulation environment. The two agents, one for each intersection, are centrally trained using a value decomposition network (VDN) architecture. The trained agents show slightly better performance compared to coordinated actuated signal control based on overall intersection delay at v/c of 0.95. In the second part of the study the trained signal control agents are used as background signal controllers while developing event-based TSP agents. In one variation, independent TSP agents are formulated and trained under a decentralized training and decentralized execution (DTDE) framework to implement TSP at each intersection. In the second variation, the two TSP agents are centrally trained under a centralized training and decentralized execution (CTDE) framework and VDN architecture to select and implement coordinated TSP strategies across the two intersections. In both cases the agents converge to the same bus delay value, but independent agents show high instability throughout the training process. For the test runs, the two independent agents reduce bus delay across the two intersections by 22% compared to the no TSP case while the coordinated TSP agents achieve 27% delay reduction. In both cases, there is only a slight increase in delay for a majority of the side street movements.

cs.AI

Adaptive Transit Signal Priority based on Deep Reinforcement Learning and Connected Vehicles in a Traffic Microsimulation Environment

Model free reinforcement learning (RL) provides a potential alternative to earlier formulations of adaptive transit signal priority (TSP) algorithms based on mathematical programming that require complex and nonlinear objective functions. This study extends RL - based traffic control to include TSP. Using a microscopic simulation environment and connected vehicle data, the study develops and tests a TSP event-based RL agent that assumes control from another developed RL - based general traffic signal controller. The TSP agent assumes control when transit buses enter the dedicated short-range communication (DSRC) zone of the intersection. This agent is shown to reduce the bus travel time by about 21%, with marginal impacts to general traffic at a saturation rate of 0.95. The TSP agent also shows slightly better bus travel time compared to actuated signal control with TSP. The architecture of the agent and simulation is selected considering the need to improve simulation run time efficiency.

cs.LG

Virus-host interactions shape viral dispersal giving rise to distinct classes of travelling waves in spatial expansions

Reaction-diffusion waves have long been used to describe the growth and spread of populations undergoing a spatial range expansion. Such waves are generally classed as either pulled, where the dynamics are driven by the very tip of the front and stochastic fluctuations are high, or pushed, where cooperation in growth or dispersal results in a bulk-driven wave in which fluctuations are suppressed. These concepts have been well studied experimentally in populations where the cooperation leads to a density-dependent growth rate. By contrast, relatively little is known about experimental populations that exhibit density-dependent dispersal. Using bacteriophage T7 as a test organism, we present novel experimental measurements that demonstrate that the diffusion of phage T7, in a lawn of host E. coli, is hindered by steric interactions with host bacteria cells. The coupling between host density, phage dispersal and cell lysis caused by viral infection results in an effective density-dependent diffusion coefficient akin to cooperative behavior. Using a system of reaction-diffusion equations, we show that this effect can result in a transition from a pulled to pushed expansion. Moreover, we find that a second, independent density-dependent effect on phage dispersal spontaneously emerges as a result of the viral incubation period, during which phage is trapped inside the host unable to disperse. Additional stochastic agent-based simulations reveal that lysis time dramatically affects the rate of diversity loss in viral expansions. Taken together, our results indicate both that bacteriophage can be used as a controllable laboratory population to investigate the impact of density-dependent dispersal on evolution, and that the genetic diversity and adaptability of expanding viral populations could be much greater than is currently assumed.

physics.bio-ph

Interaction between Nearly Hard Colloidal Spheres at an Oil-Water Interface

We show that the interaction potential between sterically stabilized, nearly hard-sphere [poly(methylmethacrylate)-poly(lauryl methacrylate) (PMMA-PLMA)] colloids at a water-oil interface has a negligible unscreened-dipole contribution, suggesting that models previously developed for charged particles at liquid interfaces are not necessarily applicable to sterically stabilized particles. Interparticle potentials, $U(r)$, are extracted from radial distribution functions [$g(r)$, measured by fluorescence microscopy] via Ornstein-Zernike inversion and via a reverse Monte Carlo scheme. The results are then validated by particle tracking in a blinking optical trap. Using a Bayesian model comparison, we find that our PMMA-PLMA data is better described by a screened monopole only rather than a functional form having a screened monopole plus an unscreened dipole term. We postulate that the long range repulsion we observe arises mainly through interactions between neutral holes on a charged interface, i.e., the charge of the liquid interface cannot, in general, be ignored. In agreement with this interpretation, we find that the interaction can be tuned by varying salt concentration in the aqueous phase. Inspired by recent theoretical work on point charges at dielectric interfaces, which we explain is relevant here, we show that a screened $\frac{1}{r^2}$ term can also be used to fit our data. Finally, we present measurements for poly(methyl methacrylate)-poly(12-hydroxystearic acid) (PMMA-PHSA) particles at a water-oil interface. These suggest that, for PMMA-PHSA particles, there is an additional contribution to the interaction potential. This is in line with our optical-tweezer measurements for PMMA-PHSA colloids in bulk oil, which indicate that they are slightly charged.

cond-mat.soft