SearcharxivSearch

arXiv subjects

Pedram Rabiee

Publications and source records attributed to Pedram Rabiee.

8 recordsLinked to original sources

Safe Exploration in Reinforcement Learning: Training Backup Control Barrier Functions with Zero Training Time Safety Violations

This paper introduces the reinforcement learning backup shield (RLBUS), an algorithm that guarantees safe exploration in reinforcement learning (RL) by incorporating backup control barrier functions (BCBFs). RLBUS constructs an implicit control forward invariant subset of the safe set using multiple backup policies, ensuring safety in the presence of input constraints. While traditional BCBFs often result in conservative control forward-invariant sets due to the design of backup controllers, RLBUS addresses this limitation by leveraging model-free RL to train an additional backup policy, which enlarges the identified control forward invariant subset of the safe set. This approach enables the exploration of larger regions in the state space with zero safety violations during training. The effectiveness of RLBUS is demonstrated on an inverted pendulum example, where the expanded invariant set allows for safe exploration over a broader state space, enhancing performance without compromising safety.

eess.SY

Guaranteed-Safe MPPI Through Composite Control Barrier Functions for Efficient Sampling in Multi-Constrained Robotic Systems

We present a new guaranteed-safe model predictive path integral (GS-MPPI) control algorithm that enhances sample efficiency in nonlinear systems with multiple safety constraints. The approach use a composite control barrier function (CBF) along with MPPI to ensure all sampled trajectories are provably safe. We first construct a single CBF constraint from multiple safety constraints with potentially differing relative degrees, using it to create a safe closed-form control law. This safe control is then integrated into the system dynamics, allowing MPPI to optimize over exclusively safe trajectories. The method not only improves computational efficiency but also addresses the myopic behavior often associated with CBFs by incorporating long-term performance considerations. We demonstrate the algorithm's effectiveness through simulations of a nonholonomic ground robot subject to position and speed constraints, showcasing safety and performance.

eess.SY

Soft-Minimum and Soft-Maximum Barrier Functions for Safety with Actuation Constraints

This paper presents two new control approaches for guaranteed safety (remaining in a safe set) subject to actuator constraints (the control is in a convex polytope). The control signals are computed using real-time optimization, including linear and quadratic programs subject to affine constraints, which are shown to be feasible. The first control method relies on a soft-minimum barrier function that is constructed using a finite-time-horizon prediction of the system trajectories under a known backup control. The main result shows that the control is continuous and satisfies the actuator constraints, and a subset of the safe set is forward invariant under the control. Next, we extend this method to allow from multiple backup controls. This second approach relies on a combined soft-maximum/soft-minimum barrier function, and it has properties similar to the first. We demonstrate these controls on numerical simulations of an inverted pendulum and a nonholonomic ground robot.

eess.SY

A Closed-Form Control for Safety Under Input Constraints Using a Composition of Control Barrier Functions

We present a closed-form optimal control that satisfies both safety constraints (i.e., state constraints) and input constraints (e.g., actuator limits) using a composition of multiple control barrier functions (CBFs). This main contribution is obtained through the combination of several ideas. First, we present a method for constructing a single relaxed control barrier function (R-CBF) from multiple CBFs, which can have different relative degrees. The construction relies on a log-sum-exponential soft-minimum function and yields an R-CBF whose zero-superlevel set is a subset of the intersection of the zero-superlevel sets of all CBFs used in the composition. Next, we use the soft-minimum R-CBF to construct a closed-form control that is optimal with respect to a quadratic cost subject to the safety constraints. Finally, we use the soft-minimum R-CBF to develop a closed-form optimal control that not only guarantees safety but also respects input constraints. The key elements in developing this novel control include: the introduction of the control dynamics, which allow the input constraints to be transformed into controller-state constraints; the use of the soft-minimum R-CBF to compose multiple safety and input CBFs, which have different relative degrees; and the development of a desired surrogate control (i.e., a desired input to the control dynamics). We demonstrate these new control approaches in simulation on a nonholonomic ground robot.

eess.SY

Composition of Control Barrier Functions With Differing Relative Degrees for Safety Under Input Constraints

This paper presents a new approach for guaranteed safety subject to input constraints (e.g., actuator limits) using a composition of multiple control barrier functions (CBFs). First, we present a method for constructing a single CBF from multiple CBFs, which can have different relative degrees. This construction relies on a soft minimum function and yields a CBF whose $0$-superlevel set is a subset of the union of the $0$-superlevel sets of all the CBFs used in the construction. Next, we extend the approach to systems with input constraints. Specifically, we introduce control dynamics that allow us to express the input constraints as CBFs in the closed-loop state (i.e., the state of the system and the controller). The CBFs constructed from input constraints do not have the same relative degree as the safety constraints. Thus, the composite soft-minimum CBF construction is used to combine the input-constraint CBFs with the safety-constraint CBFs. Finally, we present a feasible real-time-optimization control that guarantees that the state remains in the $0$-superlevel set of the composite soft-minimum CBF. We demonstrate these approaches on a nonholonomic ground robot example.

eess.SY

Implementation of Linear Parameter Varying System to Investigate the Impact of Varying Flow Rate on the Lithium-ion Batteries Thermal Management System Performance

Battery thermal management system is an indispensable part of the electric vehicles working with Lithium-ion batteries. Accordingly, lithium-ion batteries modeling, battery heat generation, and thermal management are the main focus of researchers and car manufacturers. To fulfill the need of manufacturers in the design process, a faster model than time-consuming Computational Fluid Dynamics models (CFD) is required. Reduced Order Models (ROM) address this requirement to maintain the accuracy of CFD models while could be compiled faster. Linear Time Invariant (LTI) reduced order model has been used in the literature; however, due to the limitation of LTI system, considering the constant flow rate for the cooling fluid, a Linear Parameter Varying system with three scheduling parameters was developed in this study. It is shown that LPV system results could fit accurately to CFD results in conditions that LTI system cannot maintain accuracy. Moreover, it is shown that applying varying water flow rates could result in a smoother temperature profile.

eess.SY

Soft-Minimum Barrier Functions for Safety-Critical Control Subject to Actuation Constraints

This paper presents a new control approach for guaranteed safety (remaining in a safe set) subject to actuator constraints (the control is in a convex polytope). The control signals are computed using real-time optimization, including linear and quadratic programs subject to affine constraints, which are shown to be feasible. The control method relies on a new soft-minimum barrier function that is constructed using a finite-time-horizon prediction of the system trajectories under a known backup control. The main result shows that: (i) the control is continuous and satisfies the actuator constraints, and (ii) a subset of the safe set is forward invariant under the control. We also demonstrate this control on numerical simulations of an inverted pendulum and a double-integrator ground robot.

eess.SY

The Impact of Reference-Command Preview on Human-in-the-Loop Control Behavior

This article presents results from an experiment in which 44 human subjects interact with a dynamic system to perform 40 trials of a command-following task. The reference command is unpredictable and different on each trial, but all subjects have the same sequence of reference commands for the 40 trials. The subjects are divided into 4 groups of 11 subjects. One group performs the command-following task without preview of the reference command, and the other 3 groups are given preview of the reference command for different time lengths into the future (0.5 s, 1 s, 1.5 s). A subsystem identification algorithm is used to obtain best-fit models of each subject's control behavior on each trial. The time- and frequency-domain performance, as well as the identified models of the control behavior for the 4 groups are examined to investigate the effects of reference-command preview. The results suggest that preview tends to improve performance by allowing the subjects to compensate for sensory time delay and approximate the inverse dynamics in feedforward. However, too much preview may decrease performance by degrading the ability to use the correct phase lead in feedforward.

eess.SY