SearcharxivSearch

arXiv subjects

Thomas Tournaire

Publications and source records attributed to Thomas Tournaire.

2 recordsLinked to original sources

V-Max: A Reinforcement Learning Framework for Autonomous Driving

Learning-based decision-making has the potential to enable generalizable Autonomous Driving (AD) policies, reducing the engineering overhead of rule-based approaches. Imitation Learning (IL) remains the dominant paradigm, benefiting from large-scale human demonstration datasets, but it suffers from inherent limitations such as distribution shift and imitation gaps. Reinforcement Learning (RL) presents a promising alternative, yet its adoption in AD remains limited due to the lack of standardized and efficient research frameworks. To this end, we introduce V-Max, an open research framework providing all the necessary tools to make RL practical for AD. V-Max is built on Waymax, a hardware-accelerated AD simulator designed for large-scale experimentation. We extend it using ScenarioNet's approach, enabling the fast simulation of diverse AD datasets.

cs.LG

Optimal control policies for resource allocation in the Cloud: comparison between Markov decision process and heuristic approaches

We consider an auto-scaling technique in a cloud system where virtual machines hosted on a physical node are turned on and off depending on the queue's occupation (or thresholds), in order to minimise a global cost integrating both energy consumption and performance. We propose several efficient optimisation methods to find threshold values minimising this global cost: local search heuristics coupled with aggregation of Markov chain and with queues approximation techniques to reduce the execution time and improve the accuracy. The second approach tackles the problem with a Markov Decision Process (MDP) for which we proceed to a theoretical study and provide theoretical comparison with the first approach. We also develop structured MDP algorithms integrating hysteresis properties. We show that MDP algorithms (value iteration, policy iteration) and especially structured MDP algorithms outperform the devised heuristics, in terms of time execution and accuracy. Finally, we propose a cost model for a real scenario of a cloud system to apply our optimisation algorithms and show their relevance.

math.OC