Searcharxiv⌕ Search

arXiv subjects

Andreas A. Malikopoulos

Publications and source records attributed to Andreas A. Malikopoulos.

At least 19 recordsLinked to original sources

Predictive Rerouting of Connected and Automated Vehicles Using Traffic and Charging Demand Forecasts

In this paper, we consider the problem of predictive rerouting of connected and automated vehicles (CAVs) in mixed traffic with electric vehicle charging demand. We provide a framework that combines a diffusion convolutional recurrent neural network (DCRNN) with a routing policy that accounts for congestion, CO2 emissions, route length, and charging demand. The DCRNN uses historical network observations to forecast traffic conditions and charging demand. These forecasts are then used to evaluate feasible alternative routes for eligible CAVs. A route change is accepted when the alternative preserves connectivity to the original destination and improves the prescribed route cost. We evaluate the proposed framework in SUMO under controlled traffic disruptions at five CAV penetration levels, ranging from 5% to 45%. We compare its performance with K-shortest-path routing, predictive-density routing, V2X proactive routing, and a reference scenario without rerouting. In the considered scenarios, the proposed framework reduces the average travel-time index by approximately 1.8% relative to the reference scenario and achieves the lowest average travel-time index, highest average speed, and lowest aggregate CO2 emissions among the active routing methods. Across all five penetration levels, it accepts 41 route changes, compared with 146 for K-shortest-path routing and 153 for V2X proactive routing. The reference scenario retains lower aggregate emissions and distance traveled, illustrating the tradeoff between congestion reduction and the additional travel associated with rerouting.

cs.NI↗

Preference-Adaptive Control in Autonomous Driving

In this paper, we present a preference-adaptive receding-horizon control framework for autonomous driving that accounts for passenger preferences and motion-sickness susceptibility. We formulate a finite-horizon optimal control problem with adaptive weights for speed, acceleration comfort, and motion sickness, while maintaining a fixed weight for collision risk. We predict motion sickness online using an individualized model on the Motion Illness Symptoms Classifi- cation (MISC) scale and update the preference weights offline from emotion-derived pairwise comparisons using Bayesian inference. A deterministic safety supervisor checks the planned trajectory and modifies the control command when necessary. We evaluate the framework using three simulated passenger profiles under three motion-sickness susceptibility levels. The learned weights yield distinct closed-loop behaviors, and the mean evaluation emotion score improves in seven of nine scenarios. No collisions occur in the full-method experiments, while the safety supervisor intervenes in 1.621% of the learning frames.

eess.SY↗

Closed-Loop Refinement and Execution for Learned Driving Planners

Learning-based driving planners are usually trained and evaluated in open loop against logged trajectories. In closed loop, a trajectory with small displacement error can still stall the vehicle, steer it into a conflict with surrounding agents, or be executed with abrupt braking. We introduce Closed-Loop Refinement and Execution (CLRE), a hierarchical receding-horizon control framework designed to mitigate these failure modes while leaving the upstream planner frozen and adding no new learned model. The upper layer treats the nominal trajectory as a reference and solves a finite-horizon optimal control problem that trades route progress against interaction with predicted agents. Solving it from several initializations gives a candidate set, and a prediction-conditioned oriented-bounding-box (OBB) feasibility test retains only candidates whose minimum predicted OBB clearance over the horizon meets a threshold. The lower layer executes the lowest-cost survivor, or a route-centerline backup when none remains, through the tracking controller supplied with the planner, augmented by a range-based speed bound and a saturated proportional braking law. In closed-loop simulation on the 220-route Bench2Drive validation set with VAD as the upstream planner, CLRE raises the driving score from 42.26 to 55.89 and route completion from 55.69 to 71.51, and reduces collision events from 117 to 97.

eess.SY↗

Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss

In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups. Each group uses a recurrent predictor to estimate unavailable interaction information due to communication loss. A higher-level policy then generates a compact coordination reference that conditions the local control policy within each group. A predictive safety filter evaluates and modifies the proposed controls when they violate safety constraints. Simulation results show improved task completion under communication loss, reduced communication growth as the team size increases, and safe operation in the tested scenarios.

eess.SY↗

Safety-Constrained Optimal Control and Hamiltonian-Gradient Learning for Systems with Unknown Dynamics

We present a framework for safety-constrained optimal control when the plant dynamics are unknown and only an approximate model is available. We apply Pontryagin's Minimum Principle to a model-based problem whose cost penalizes model-plant state mismatch and establish conditions under which the constrained model and plant Hamiltonians have the same minimizer. These conditions show that optimality can be recovered by preserving control-relevant Hamiltonian structure rather than identifying the complete plant dynamics. Motivated by this observation, we develop a fully online method that uses integral trajectory data and policy iteration to learn the dynamics-dependent factors of the plant Hamiltonian gradient directly. Experiments on robotic cruise control demonstrate constrained model-plant control equivalence and online Hamiltonian-gradient reconstruction.

eess.SY↗

A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.

math.OC↗

Real-time autonomous magnetic microrobot navigation across dynamic and biologically relevant environments

Autonomous microrobots could enable minimally invasive interventions in confined biological environments, but their operation requires real-time navigation among moving obstacles and environmental disturbances. Here we present a closed-loop framework for autonomous magnetic microrobot navigation that separates long-range geometric planning from short-range reactive control. An analytic geometry-based planner continuously generates collision-free global routes from microscopy images, while rule-based or reinforcement-learning local controllers respond to moving obstacles before returning control to the replanned global route. The planner produces short paths with substantially lower computation time than representative graph-, sampling- and geometry-based planners, and the integrated framework operates within the video-rate control budget. We demonstrate microrobot autonomous navigation in static and dynamic microfluidic environments, including avoidance of moving particles and human red blood cells, navigation with and against fluid flow, and repeated long-duration target transitions. We further demonstrate magnetic microrobot actuation in ex vivo ocular and vascularized cortical tissues and photoacoustic detection of magnetically induced microrobot dynamics in post-mortem tissue. These results establish a modular framework for autonomous magnetic microrobot navigation in complex biological environments.

cs.RO↗

When Does Selfishness Align with Team Goals? A Structural Analysis of Equilibrium and Optimality

This paper investigates the relationship between the team-optimal solution and the Nash equilibrium (NE) to assess the impact of self-interested decisions on team performance. In classical team decision problems, team members typically act cooperatively towards a common objective to achieve a team-optimal solution. However, in practice, members may behave selfishly by prioritizing their goals, resulting in an NE under a non-cooperative game. To study this misalignment, we develop a parameterized model for team and game problems, where game parameters represent each individual's deviation from the team objective. The study begins by exploring the consistency and deviation between the NE and the team-optimal solution under fixed game parameters. We provide a necessary and sufficient condition for any NE to be a team optimum, along with establishing an upper bound to measure their difference when this consistency fails. We then study how to steer the NE toward the team-optimal solution by adjusting game parameters in an incomplete-information leader--follower setting, where the leader observes only equilibrium responses rather than the exact lower-level game structure. To address this challenge, we develop a two-stage learned-response intervention framework: the leader first learns a behaviorally consistent lower-level model from observed equilibria and then computes the intervention over the learned response map using bilevel hypergradient optimization, followed by convergence analysis and simulation validation.

math.OC↗

Route Recommendations for Traffic Management Under Learned Partial Driver Compliance

In this paper, we aim to mitigate congestion in traffic management systems by guiding travelers along system-optimal (SO) routes. However, we recognize that most theoretical approaches assume perfect driver compliance, which often does not reflect reality, as drivers tend to deviate from recommendations to fulfill their personal objectives. Therefore, we propose a route recommendation framework that explicitly learns partial driver compliance and optimizes traffic flow under realistic adherence. We first compute an SO edge flow through flow optimization techniques. Next, we train a compliance model based on historical driver decisions to capture individual responses to our recommendations. Finally, we formulate a stochastic optimization problem that minimizes the gap between the target SO flow and the realized flow under conditions of imperfect adherence. Our simulations conducted on a grid network reveal that our approach significantly reduces travel time compared to baseline strategies, demonstrating the practical advantage of incorporating learned compliance into traffic management.

eess.SY↗

Accelerating Time-Optimal Trajectory Planning for Connected and Automated Vehicles with Graph Neural Networks

In this paper, we present a learning-based framework that accelerates time- and energy-optimal trajectory planning for connected and automated vehicles (CAVs) using graph neural networks (GNNs). We formulate the multi-agent coordination problem encountered in traffic scenarios as a cooperative trajectory planning problem that minimizes travel time, subject to motion primitives derived from energy-optimal solutions. The performance of this framework can be further improved through replanning at each time step, enabling the system to incorporate newly observed information. To achieve real-time execution, we employ a graph isomorphism network with edge features (GINEConv) to learn the solutions of the time-optimal trajectory planning problem from offline-generated data. The trained model produces online predictions that serve as warm-starts for numerical optimization, thereby enabling rapid computation of minimal exit times and the associated feasible trajectories. This learning-to-warm-start approach substantially reduces computation time while preserving the control performance of the time- and energy-optimal trajectory planning framework.

eess.SY↗

Spatiotemporal Forecasting of Incidents and Congestion with Implications for Sustainable Traffic Control

Urban traffic anomalies, such as collisions and disruptions, threaten the safety, efficiency, and sustainability of transportation systems. In this paper, we present a simulation-based framework for modeling, detecting, and predicting such anomalies in urban networks. Using the Simulation of Urban MObility (SUMO) platform, we generate reproducible rear-end and intersection crash scenarios with matched baselines, enabling controlled experimentation and comparative evaluation. We record vehicle-level travel time, speed, and emissions for both edge- and network-level analysis. Building on this dataset, we develop a hybrid forecasting architecture that combines bidirectional long short-term memory networks with a diffusion convolutional recurrent neural network to capture temporal dynamics and spatial dependencies. Our simulation studies on the Broadway corridor in New York City demonstrate the framework's ability to reproduce consistent incident conditions, quantify their effects, and provide accurate multi-horizon traffic forecasts. Our results highlight the value of combining controlled anomaly generation with deep predictive models to support reproducible evaluation and sustainable traffic management.

eess.SY↗

Integrated Routing and Intersection Control for Mixed Traffic

The rapid development of cyber-physical systems is driving a transition toward mixed traffic environments comprising both human-driven and connected and automated vehicles (CAVs). This shift presents a unique opportunity to leverage the efficient operation of CAVs to improve overall network throughput. This paper introduces a hierarchical framework designed to bridge macroscopic routing optimization at the network level with microscopic vicinity control at signalized intersections. The upper layer utilizes aggregated traffic information to provide proactive routing guidance for CAVs, aiming to minimize total travel time. The lower layer leverages local vehicle states to jointly optimize traffic light phases and individual CAV trajectories, aiming to reduce intersection crossing delays and optimize energy consumption, respectively. The effectiveness of the proposed framework is validated through SUMO on the Sioux Falls benchmark network. Results demonstrate that the integration of these macroscopic and microscopic layers yields significantly better performance compared to applying either layer in isolation, significantly improving network throughput and reducing congestion.

eess.SY↗

Closed-Form Characterization of Constrained Double-Integrator Optimal Control

We consider the energy-optimal control problem for double-integrator systems subject to state and control constraints, with fixed terminal time and free terminal speed. When the constraints become active, the optimal trajectory consists of a combination of bang, unconstrained, and coast arcs, whose switching instants must be computed explicitly. In this paper, we derive closed-form expressions for the switching times of all admissible profiles, including both constrained and unconstrained arcs, reducing the computation in each case to explicit algebraic equations. In contrast to prior work, we classify all possible combinations of arcs, including special cases, and provide the specific conditions under which each case arises. Furthermore, we prove that when the initial unconstrained trajectory violates both speed and control constraints, the optimal solution follows a predetermined bang-affine-coast profile, enabling direct identification of the optimal trajectory without intermediate feasibility checks.

eess.SY↗

Cross-fitted Proximal Learning for Model-Based Reinforcement Learning

Model-based reinforcement learning is attractive for sequential decision-making because it explicitly estimates reward and transition models and then supports planning through simulated rollouts. In offline settings with hidden confounding, however, models learned directly from observational data may be biased. This challenge is especially pronounced in partially observable systems, where latent factors may jointly affect actions, rewards, and future observations. Recent work has shown that policy evaluation in such confounded partially observable Markov decision processes (POMDPs) can be reduced to estimating reward-emission and observation-transition bridge functions satisfying conditional moment restrictions (CMRs). In this paper, we study the statistical estimation of these bridge functions. We formulate bridge learning as a CMR problem with nuisance objects given by a conditional mean embedding and a conditional density. We then develop a $K$-fold cross-fitted extension of the existing two-stage bridge estimator. The proposed procedure preserves the original bridge-based identification strategy while using the available data more efficiently than a single sample split. We also derive an oracle-comparator bound for the cross-fitted estimator and decompose the resulting error into a Stage I term induced by nuisance estimation and a Stage II term induced by empirical averaging.

cs.LG↗

An Online Learning Approach for Two-Player Zero-Sum Linear Quadratic Games

In this paper, we present an online learning approach for two-player zero-sum linear quadratic games with unknown dynamics. We develop a framework combining regularized least squares model estimation, high probability confidence sets, and surrogate model selection to maintain a regular model for policy updates. We apply a shrinkage step at each episode to identify a surrogate model in the region where the generalized algebraic Riccati equation admits a stabilizing saddle point solution. We then establish regret analysis on algorithm convergence, followed by a numerical example to illustrate the convergence performance and verify the regret analysis.

eess.SY↗

Rollout-Based Charging Scheduling for Electric Truck Fleets in Large Transportation Networks

In this paper, we investigate the charging scheduling optimization problem for large electric truck fleets operating with dedicated charging infrastructure. A central coordinator jointly determines the charging sequence and power allocation of each truck to minimize the total operational cost of the fleet. The problem is inherently combinatorial and nonlinear due to the coupling between discrete sequencing decisions and continuous charging control, rendering exact optimization intractable for real-time implementation. To address this challenge, we propose a rollout-based dynamic programming framework built upon an inner-outer two-layer structure, which decouples ordering decisions from the schedule optimization, thus enabling efficient policy evaluation and approximation. The proposed method achieves near-optimal solutions with polynomial-time complexity and adapts to dynamic arrivals and time-varying electricity prices. Simulation studies show that the rollout-based approach significantly outperforms conventional heuristics with high computational efficiency, demonstrating its effectiveness and practical applicability for real-time charging management in large-scale transportation networks.

eess.SY↗

Cooperative Detour Planning for Dual-Task Drone Fleets

As Urban air mobility scales, commercial drone fleets offer a compelling, yet underexplored opportunity to function as mobile sensor networks for real-time urban traffic monitoring. In this paper, we propose a decentralized framework that enables drone fleets to simultaneously execute delivery tasks and observe network traffic conditions. We model the urban environment with dynamic information values associated with road segments, which accumulate traffic condition uncertainty over time and are reset upon drone visitation. This problem is formulated as a mixed-integer linear programming problem where drones maximize the traffic information reward while respecting the maximum detour for each delivery and the battery budget of each drone. Unlike centralized approaches that are computationally heavy for large fleets, our method focuses on dynamic local clustering. When drones enter communication range, they exchange their belief in traffic status and transition from isolated path planning to a local joint optimization mode, resolving coupled constraints to obtain replanned paths for each drone, respectively. Simulation results built on the real city network of Barcelona, Spain, demonstrate that, compared to a shortest-path policy that ignores the traffic monitoring task, our proposed method better utilizes the battery and detour budget to explore the city area and obtain adequate traffic information; and, thanks to its decentralized manner, this ``meet-and-merge" strategy achieves near-global optimality in network coverage with significantly reduced computation overhead compared to the centralized baseline.

eess.SY↗

Communication-Efficient Distributed Learning with Differential Privacy

We address nonconvex learning problems over undirected networks. In particular, we focus on the challenge of designing an algorithm that is both communication-efficient and that guarantees the privacy of the agents' data. The first goal is achieved through a local training approach, which reduces communication frequency. The second goal is achieved by perturbing gradients during local training, specifically through gradient clipping and additive noise. We prove that the resulting algorithm converges to a stationary point of the problem within a bounded distance. Additionally, we provide theoretical privacy guarantees within a differential privacy framework that ensure agents' training data cannot be inferred from the trained model shared over the network. We show the algorithm's superior performance on a classification task under the same privacy budget, compared with state-of-the-art methods.

cs.LG↗