SearcharxivSearch

arXiv subjects

Christopher H. Onder

Publications and source records attributed to Christopher H. Onder.

10 recordsLinked to original sources

Bridging RL and MPC for mixed-integer optimal control with application to Formula 1 race strategies

We propose a hybrid reinforcement learning (RL) and model predictive control (MPC) framework for mixed-integer optimal control, where discrete variables enter the cost and dynamics but not the constraints. Existing hierarchical approaches learn a policy only for the discrete action space, leaving continuous optimization to MPC. Unlike these methods, we train an RL agent on the full hybrid action space, such that the learned critic approximates the Q-function of the underlying Markov decision process. During deployment, the RL actor is rolled out over the prediction horizon to parametrize an integer-free nonlinear MPC through the discrete action sequence and provide a continuous warm-start. The learned critic can additionally serve as a terminal cost to capture long-term performance. We prove recursive feasibility, and validate the framework on a Formula 1 race strategy problem, where an ablation study identifies the contribution of each learned component. The hybrid method achieves near-optimal performance relative to an offline mixed-integer nonlinear program benchmark, outperforming a standalone RL agent. Moreover, the hybrid scheme enables adaptation to unseen disturbances through modular MPC extensions at zero retraining cost.

eess.SY

Game Theory in Formula 1: From Physical to Strategic Interactions

This paper presents an optimization framework to model multi-agent racing dynamics. By incorporating physically accurate interaction models and accounting for the optimal responses of competing agents, our approach reveals strategic behaviors typical of motorsport. Aerodynamic wake effects, trajectory optimization, and energy management are captured and evaluated on a representative case study, based on a Formula 1 scenario. We describe the minimum lap time problem with two agents as either a Nash or a Stackelberg game, and by employing the Karush-Kuhn-Tucker conditions during the problem formulation, we recover the structure of a nonlinear program. In addition, we introduce an algorithm to refine local Stackelberg solutions, using the Nash costs as upper bounds. The resulting strategies are analyzed through case studies. We examine the impact of slipstreaming on trajectory selection in corners, straights, and high-speed sections, while also identifying optimal overtaking locations based on energy allocation strategies. Exploiting the structural similarities of the game formulations, we are able to compare symmetric and hierarchical strategies to analyze competitive racing dynamics. The proposed methodology closes the gap between theoretical game theory and practical applications, with relevance in multi-agent systems with coupled nonlinear dynamics.

eess.SY

Learning-based Multi-agent Race Strategies in Formula 1

In Formula 1, race strategies are adapted according to evolving race conditions and competitors' actions. This paper proposes a reinforcement learning approach for multi-agent race strategy optimization. Agents learn to balance energy management, tire degradation, aerodynamic interaction, and pit-stop decisions. Building on a pre-trained single-agent policy, we introduce an interaction module that accounts for the behavior of competitors. The combination of the interaction module and a self-play training scheme generates competitive policies, and agents are ranked based on their relative performance. Results show that the agents adapt pit timing, tire selection, and energy allocation in response to opponents, achieving robust and consistent race performance. Because the framework relies only on information available during real races, it can support race strategists' decisions before and during races.

cs.AI

Towards Learning-Based Formula 1 Race Strategies

This paper presents two complementary frameworks to optimize Formula 1 race strategies, jointly accounting for energy allocation, tire wear and pit stop timing. First, the race scenario is modeled using lap time maps and a dynamic tire wear model capturing the main trade-offs arising during a race. Then, we solve the problem by means of a mixed-integer nonlinear program that handles the integer nature of the pit stop decisions. The same race scenario is embedded into a reinforcement learning environment, on which an agent is trained. Providing fast inference at runtime, this method is suited to improve human decision-making during real races. The learned policy's suboptimality is assessed with respect to the optimal solution, both in a nominal scenario and with an unforeseen disturbance. In both cases, the agent achieves approximately 5s of suboptimality on 1.5h of race time, mainly attributable to the different energy allocation strategy. This work lays the foundations for learning-based race strategies and provides a benchmark for future developments.

eess.SY

Game-theoretic Energy Management Strategies With Interacting Agents in Formula 1

This paper presents an interaction-aware energy management optimization framework for Formula 1 racing. The considered scenario involves two agents and a drag reduction model. Strategic interactions between the agents are captured by a Stackelberg game formulated as a bilevel program. To address the computational challenges associated with bilevel optimization, the problem is reformulated as a single-level nonlinear program employing the Karush-Kuhn-Tucker conditions. The proposed framework contributes towards the development of new energy management and allocation strategies, caused by the presence of another agent. For instance, it provides valuable insights on how to redistribute the energy in order to optimally exploit the wake effect, showcasing a notable difference with the behavior studied in previous works. Robust energy allocations can be identified to reduce the lap time loss associated with unexpected choices of the other agent. It allows to recognize the boundary conditions for the interaction to become relevant, impacting the system's behavior, and to assess if overtaking is possible and beneficial. Overall, the framework provides a comprehensive approach for a two-agent Formula 1 racing problem with strategic interactions, offering physically intuitive and practical results.

eess.SY

Optimization of the Energy-Comfort Trade-Off of HVAC Systems in Electric City Buses Based on a Steady-State Model

The electrification of public transport vehicles offers the potential to relieve city centers of pollutant and noise emissions. Furthermore, electric buses have lower life-cycle greenhouse gas (GHG) emissions than diesel buses, particularly when operated with sustainably produced electricity. However, the heating, ventilation, and air-conditioning (HVAC) system can consume a significant amount of energy, thus limiting the achievable driving range. In this paper, we address the HVAC system in an electric city bus by analyzing the trade-off between the energy consumption and the thermal comfort of the passengers. We do this by developing a dynamic thermal model for the bus, which we simplify by considering it to be in steady state. We introduce a method that is able to quickly optimize the steady-state HVAC system inputs for a large number of samples representative of a year-round operation. A comparison between the results from the steady-state optimization approach and a dynamic simulation reveals small deviations in both the HVAC system power demand and achieved thermal comfort. Thus, the approximation of the system performance with a steady-state model is justified. We present two case studies to demonstrate the practical relevance of the approach. First, we show how the method can be used to compare different HVAC system designs based on a year-round performance evaluation. Second, we show how the method can be used to extract setpoints for online controllers that achieve close-to-optimal performance without any predictive information. In conclusion, this study shows that a steady-state analysis of the HVAC systems of an electric city bus is a valuable approach to evaluate and optimize its performance.

eess.SY

Highly Efficient Year-Round Energy and Comfort Optimization of HVAC Systems in Electric City Buses

In this paper, we present a novel approach to perform highly efficient numerical simulations of the heating, ventilation, and air-conditioning (HVAC) system of an electric city bus. The models for this simulation are based on the assumption of a steady-state operation. We show two approaches to obtain the minimum energy requirement for a certain thermal comfort criterion under specific ambient conditions. Due to the computationally efficient approach developed, we can evaluate the model on a large dataset of 7500 scenarios in various ambient conditions to estimate the year-round performance of the system subject to different comfort requirements. Compared to a heating strategy based on positive temperature coefficient (PTC) elements, we can thus show that a heat pump (HP) can reduce the annual mean power consumption by up to 60%. Ceiling-mounted radiant heating elements complementing a PTC heating system can reduce the annual mean power consumption by up to 10%, while they cannot improve the energy efficiency when used in conjunction with a HP. Finally, a broad sensitivity study reveals the fact that improving the HP's heating-mode coefficient of performance (COP) manifests the largest leverage in terms of mean annual power consumption. Moreover, the annual energy expenditure for cooling are around eight times smaller than those for heating. The case study considered thus reveals that the advantages of improving the COP of the cooling mode are significantly lower.

eess.SY

ZTBus: A Large Dataset of Time-Resolved City Bus Driving Missions

This paper presents the Zurich Transit Bus (ZTBus) dataset, which consists of data recorded during driving missions of electric city buses in Zurich, Switzerland. The data was collected over several years on two trolley buses as part of multiple research projects. It involves more than a thousand missions across all seasons, each mission usually covering a full day of operation. The ZTBus dataset contains detailed information on the vehicle's power demand, propulsion system, odometry, global position, ambient temperature, door openings, number of passengers, dispatch patterns within the public transportation network, etc. All signals are synchronized in time and include an absolute timestamp in tabular form. The dataset can be used as a foundation for a variety of studies and analyses. For example, the data can serve as a basis for simulations to estimate the performance of different public transit vehicle types, or to evaluate and optimize control strategies of hybrid electric vehicles. Furthermore, numerous influencing factors on vehicle operation, such as traffic, passenger volume, etc., can be analyzed in detail.

eess.SY

Low-level Online Control of the Formula 1 Power Unit with Feedforward Cylinder Deactivation

Since 2014, the Fédération Internationale de l'Automobile has prescribed a parallel hybrid powertrain for the Formula 1 race cars. The complex low-level interactions between the thermal and the electrical part represent a non-trivial and challenging system to be controlled online. We present a novel controller architecture composed of a supervisory controller for the energy management, a feedforward cylinder deactivation controller, and a track region-dependent low-level nonlinear model predictive controller to optimize the engine actuators. Except for the nonlinear model predictive controller, the proposed controller subsystems are computationally inexpensive and are real time capable. The framework is tested and validated in a simulation environment for several realistic scenarios disturbed by driver actions or grip conditions on the track. In particular, we analyze how the control architecture deals with an unexpected gearshift trajectory during an acceleration phase. Further, we demonstrate how an increased maximum velocity trajectory impacts the online low-level controller. Our results show a suboptimality over an entire lap with respect to the benchmark solution of 49 ms and 64 ms, respectively, which we deem acceptable. Compared to the same control architecture with full knowledge of the disturbances, the suboptimality amounted to only 2 ms and 17 ms. For all case studies we show that the cylinder deactivation capability decreases the suboptimality by 7 to 8 ms.

eess.SY

On the Interaction between Autonomous Mobility-on-Demand and Public Transportation Systems

In this paper we study models and coordination policies for intermodal Autonomous Mobility-on-Demand (AMoD), wherein a fleet of self-driving vehicles provides on-demand mobility jointly with public transit. Specifically, we first present a network flow model for intermodal AMoD, where we capture the coupling between AMoD and public transit and the goal is to maximize social welfare. Second, leveraging such a model, we design a pricing and tolling scheme that allows to achieve the social optimum under the assumption of a perfect market with selfish agents. Finally, we present a real-world case study for New York City. Our results show that the coordination between AMoD fleets and public transit can yield significant benefits compared to an AMoD system operating in isolation.

eess.SY