Learning State-Action Control Barrier Functions for Model-Free Control under Constraints
Ensuring constraint satisfaction for learning-based control is a critical challenge, especially in the model-free case. While safety filters address this challenge in the model-based setting, they typically rely on predictive models derived from physics or data. This reliance limits their applicability for advanced model-free learning control methods. To address this gap, we propose a new optimization-based control framework that determines safe control inputs directly from data. The benefit of the framework is that both the objective function and the safety constraint can be updated using model-free learning algorithms, enabling performance optimization while maintaining safety. As a key component, the concept of direct data-driven safety filters (3DSF) is first proposed. The framework employs a novel safety certificate, called the state-action control barrier function (SACBF). We present two approaches for synthesizing SACBFs, including learning from an expert controller and reinforcement learning. We develop a robustness framework that guarantees safety and recursive feasibility for the expert-guided approach under learning errors. The proposed control framework bridges the gap between model-free learning-based control and constrained control, by decoupling performance optimization from safety enforcement. Simulations on vehicle control illustrate the superior performance regarding constraint satisfaction and task achievement compared to model-based methods and reward shaping.