SearcharxivSearch

arXiv subjects

Iman Fadakar

Publications and source records attributed to Iman Fadakar.

2 recordsLinked to original sources

SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

Multi-agent interaction is a fundamental aspect of autonomous driving in the real world. Despite more than a decade of research and development, the problem of how to competently interact with diverse road users in diverse scenarios remains largely unsolved. Learning methods have much to offer towards solving this problem. But they require a realistic multi-agent simulator that generates diverse and competent driving interactions. To meet this need, we develop a dedicated simulation platform called SMARTS (Scalable Multi-Agent RL Training School). SMARTS supports the training, accumulation, and use of diverse behavior models of road users. These are in turn used to create increasingly more realistic and diverse interactions that enable deeper and broader research on multi-agent interaction. In this paper, we describe the design goals of SMARTS, explain its basic architecture and its key features, and illustrate its use through concrete multi-agent experiments on interactive scenarios. We open-source the SMARTS platform and the associated benchmark tasks and evaluation metrics to encourage and empower research on multi-agent learning for autonomous driving. Our code is available at https://github.com/huawei-noah/SMARTS.

cs.MA

Adaptive Hessian Estimation Based Extremum Localization

In this paper we study continuous time adaptive extremum localization of an arbitrary quadratic function $F(\cdot)$ based on Hessian estimation, using measured the signal intensity by a sensory agent. The function $F(\cdot)$ represents a signal field as a result of a source located at the maximum point of $F(\cdot)$ and is decreasing as moving away from the source location. Stability of the proposed adaptive estimation and localization scheme is analyzed and the Hessian parameter and location estimates are shown to asymptotically converge to the true values. Moreover, the stability and convergence properties of algorithm are shown to be robust to drift in the extremum location. Simulation test results are displayed to verify the established properties of the proposed scheme as well as robustness to signal measurement noise.

math.OC