SearcharxivSearch

arXiv subjects

Tom Lefebvre

Publications and source records attributed to Tom Lefebvre.

At least 19 recordsLinked to original sources

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. Modern data-driven methods mitigate uncertainty through large-scale training and computation, and degrade significantly in performance with limited number of training samples. By contrast, classical model-based controllers are computationally efficient and reliable, but their limited ability to represent task-relevant uncertainty can hinder performance in contact-rich interactions. In this work, we propose to expand the capabilities of model-based manipulation control through more flexible uncertainty modeling that retains performance while exactly adapting to uncertainty. Our approach casts the manipulation problem as a distributionally robust control optimization and proposes a novel deterministic formulation based on Stein variational inference that preserves performance while explicitly modeling task-sensitive parameter uncertainty. As a result, the derived controllers are more aware of task sensitivities to uncertainty, yielding high reliability without compromising performance. Experimental results demonstrate up to 3$\times$ improved robustness across a range of contact-rich manipulation tasks under broad parametric uncertainty, outperforming existing model-based control methods.

cs.RO

Gaussian Process Dual MPC using Active Inference: An Autonomous Vehicle Usecase

Designing controllers under uncertainty requires balancing the need to explore system dynamics with the requirement to maintain reliable control performance. Dual control addresses this challenge by selecting actions that both regulate the system and actively gather informative data. This paper investigates the use of the Active Inference framework, grounded in the Free Energy Principle, for developing a dual model-predictive controller (MPC). To identify and quantify uncertainty, we introduce an online sparse semi-parametric Gaussian Process model that combines the flexibility of nonparametric with the efficiency of parametric learning for real-time updates. By applying the expected free energy functional to this adaptive probabilistic model, we derive an MPC objective that incorporates an information-theoretic term, which captures uncertainty arising from both the learned model and measurement noise. This formulation leads to a stochastic optimal control problem for dual controller design, which is solved using a novel dynamic-programming-based method. Simulation results on a vehicle use case demonstrate that the proposed algorithm enhances autonomous driving control performance across different settings and scenarios.

math.OC

Unifying Entropy Regularization in Optimal Control: From and Back to Classical Objectives via Iterated Soft Policies and Path Integral Solutions

This paper develops a unified perspective on several optimal control formulations through the lens of Kullback-Leibler (KL) regularization. We propose a central problem that separates the KL penalties on policies and transitions with independent weights, thus generalizing the standard trajectory-level KL-regularization used in probabilistic optimal control. This umbrella formulation recovers various control problems: the classical Stochastic Optimal Control (SOC), Risk-Sensitive Stochastic Optimal Control (RSOC), and their policy-based KL-regularized counterparts, termed soft-policy SOC and RSOC, which yield tractable surrogates. Beyond being regularized variants, these soft-policy formulations majorize the original SOC and RSOC, thus, iterating their solutions recovers the original objectives. We further identify a synchronized case of soft-policy RSOC where the policy and transition KL weights coincide, yielding a linear Bellman operator, path-integral solution, and compositionality -- extending these computationally favourable properties to a broad class of control problems.

math.OC

How to Capture Human Preference: Commissioning of a Robotic Use-Case via Preferential Bayesian Optimisation

The popularity of Bayesian Optimization (BO) to automate or support the commissioning of engineering systems is rising. Conventional BO, however, relies on the availability of a scalar objective function. The latter is often difficult to define and rarely captures the nuanced judgement of expert operators in industrial settings. Preferential Bayesian Optimization (PBO) addresses this limitation by relying solely on pairwise preference feedback of a human expert, so-called duels. In this paper, we study PBO's capacity to commission a particular setup where a manipulator needs to push a block towards a target position. We benchmark state-of-the-art algorithms in both simulations and in the real world. Our results confirm that PBO can commission the set-up to the satisfaction of an expert operator whilst relying solely on binary preference feedback. To evaluate to what extend the same result can be achieved using conventional BO we investigate the experts decision consistency against an expert-designed cost function. Our study reveals that the experts fail to define a cost function that is in full agreement with their own decision process as witnessed in the PBO experiments. We then show that the auxiliary cost function that is constructed as a by-product of the PBO algorithms outperforms the expert-designed cost function in terms of decision consistency. Furthermore we demonstrate that this cost function can be used with conventional BO algorithms in an effort to reproduce the optimal design. This proofs the preference based cost function captures the experts' preferences perhaps more effectively than the experts could articulate preference themselves. In conclusion, we discuss downsides and propose directions for future research.

eess.SY

Differential Flatness of Quasi-Static Slider-Pusher Models with Applications in Control

This paper investigates the dynamic properties of planar slider-pusher systems as a motion primitive in manipulation tasks. To that end, we construct a differential kinematic model deriving from the limit surface approach under the quasi-static assumption and with negligible contact friction. The quasi-static model applies to generic slider shapes and circular pusher geometries, enabling a differential kinematic representation of the system. From this model, we analyze differential flatness - a property advantageous for control synthesis and planning - and find that slider-pusher systems with polygon sliders and circular pushers exhibit flatness with the centre of mass as a flat output. Leveraging this property, we propose two control strategies for trajectory tracking: a cascaded quasi-static feedback strategy and a dynamic feedback linearization approach. We validate these strategies through closed-loop simulations incorporating perturbed models and input noise, as well as experimental results using a physical setup with a finger-like pusher and vision-based state detection. The real-world experiments confirm the applicability of the simulation gains, highlighting the potential of the proposed methods for

eess.SY

Dual Control Reference Generation for Optimal Pick-and-Place Execution under Payload Uncertainty

This work addresses the problem of robot manipulation tasks under unknown dynamics, such as pick-and-place tasks under payload uncertainty, where active exploration and(/for) online parameter adaptation during task execution are essential to enable accurate model-based control. The problem is framed as dual control seeking a closed-loop optimal control problem that accounts for parameter uncertainty. We simplify the dual control problem by pre-defining the structure of the feedback policy to include an explicit adaptation mechanism. Then we propose two methods for reference trajectory generation. The first directly embeds parameter uncertainty in robust optimal control methods that minimize the expected task cost. The second method considers minimizing the so-called optimality loss, which measures the sensitivity of parameter-relevant information with respect to task performance. We observe that both approaches reason over the Fisher information as a natural side effect of their formulations, simultaneously pursuing optimal task execution. We demonstrate the effectiveness of our approaches for a pick-and-place manipulation task. We show that designing the reference trajectories whilst taking into account the control enables faster and more accurate task performance and system identification while ensuring stable and efficient control.

cs.RO

Probabilistic Latent Variable Modeling for Dynamic Friction Identification and Estimation

Precise identification of dynamic models in robotics is essential to support control design, friction compensation, output torque estimation, etc. A longstanding challenge remains in the identification of friction models for robotic joints, given the numerous physical phenomena affecting the underlying friction dynamics which result into nonlinear characteristics and hysteresis behaviour in particular. These phenomena proof difficult to be modelled and captured accurately using physical analogies alone. This has motivated researchers to shift from physics-based to data-driven models. Currently, these methods are still limited in their ability to generalize effectively to typical industrial robot deployement, characterized by high- and low-velocity operations and frequent direction reversals. Empirical observations motivate the use of dynamic friction models but these remain particulary challenging to establish. To address the current limitations, we propose to account for unidentified dynamics in the robot joints using latent dynamic states. The friction model may then utilize both the dynamic robot state and additional information encoded in the latent state to evaluate the friction torque. We cast this stochastic and partially unsupervised identification problem as a standard probabilistic representation learning problem. In this work both the friction model and latent state dynamics are parametrized as neural networks and integrated in the conventional lumped parameter dynamic robot model. The complete dynamics model is directly learned from the noisy encoder measurements in the robot joints. We use the Expectation-Maximisation (EM) algorithm to find a Maximum Likelihood Estimate (MLE) of the model parameters. The effectiveness of the proposed method is validated in terms of open-loop prediction accuracy in comparison with baseline methods, using the Kuka KR6 R700 as a test platform.

cs.RO

Introducing DAIMYO: a first-time-right dynamic design architecture and its application to tail-sitter UAS development

In recent years, there has been a notable evolution in various multidisciplinary design methodologies for dynamic systems. Among these approaches, a noteworthy concept is that of concurrent conceptual and control design or co-design. This approach involves the tuning of feedforward and/or feedback control strategies in conjunction with the conceptual design of the dynamic system. The primary aim is to discover integrated solutions that surpass those attainable through a disjointed or decoupled approach. This concurrent design paradigm exhibits particular promise in the context of hybrid unmanned aerial systems (UASs), such as tail-sitters, where the objectives of versatility (driven by control considerations) and efficiency (influenced by conceptual design) often present conflicting demands. Nevertheless, a persistent challenge lies in the potential disparity between the theoretical models that underpin the design process and the real-world operational environment, the so-called reality gap. Such disparities can lead to suboptimal performance when the designed system is deployed in reality. To address this issue, this paper introduces DAIMYO, a novel design architecture that incorporates a high-fidelity environment, which emulates real-world conditions, into the procedure in pursuit of a `first-time-right' design. The outcome of this innovative approach is a design procedure that yields versatile and efficient UAS designs capable of withstanding the challenges posed by the reality gap.

math.OC

Optimal Path Planning of Airborne Wind Energy Systems with a Flexible Tether

In this work, we establish an optimal control framework for airborne wind energy systems (AWESs) with flexible tethers. The AWES configuration, consisting of a six-degree-of-freedom aircraft, a flexible tether, and a winch, is formulated as an index-1 differential-algebraic system of equations (DAE). We achieve this by adopting a minimal coordinate representation that uses Euler angles to characterize the aircraft's attitude and employing a quasi-static approach for the tether. The presented method contrasts with other recent optimization studies that use an index-3 DAE approach. By doing so, our approach avoids related inconsistency condition problems. We use a homotopy strategy to solve the optimal control problem that ultimately generates optimal trajectories of the AWES with a flexible tether. We furthermore compare with a rigid tether model by investigating the resulting mechanical powers and tether forces. Simulation results demonstrate the efficacy of the presented methodology and the necessity to incorporate the flexibility of the tether when solving the optimal control problem.

math.OC

Deterministic Trajectory Optimization through Probabilistic Optimal Control

In this article, we discuss two algorithms tailored to discrete-time deterministic finite-horizon nonlinear optimal control problems or so-called deterministic trajectory optimization problems. Both algorithms can be derived from an emerging theoretical paradigm that we refer to as probabilistic optimal control. The paradigm reformulates stochastic optimal control as an equivalent probabilistic inference problem and can be viewed as a generalisation of the former. The merit of this perspective is that it allows to address the problem using the Expectation-Maximization algorithm. It is shown that the application of this algorithm results in a fixed point iteration of probabilistic policies that converge to the deterministic optimal policy. Two strategies for policy evaluation are discussed, using state-of-the-art uncertainty quantification methods resulting into two distinct algorithms. The algorithms are structurally closest related to the differential dynamic programming algorithm and related methods that use sigma-point methods to avoid direct gradient evaluations. The main advantage of the algorithms is an improved balance between exploration and exploitation over the iterations, leading to improved numerical stability and accelerated convergence. These properties are demonstrated on different nonlinear systems.

math.OC

Flatness-based MPC using B-splines transcription with application to a Pusher-Slider System

This work discusses the use of model predictive control (MPC) for the manipulation of a pusher-slider system. In particular we leverage the differential flatness of the pusher-slider in combination with a B-splines transcription to address the computational demand that is typically associated to real-time implementation of an MPC controller. We demonstrate the flatness based B-spline MPC controller in simulation and compare it to a standard MPC implementation approach using direct multiple shooting. We evaluate the computational advantage of the flatness based MPC empirically and document computational acceleration up to 65%.

math.OC

Differential Flatness of Slider-Pusher Systems for Constrained Time Optimal Collision Free Path Planning

In this work we show that the differential kinematics of slider-pusher systems are differentially flat assuming quasi-static behaviour and frictionless contact. Second we demonstrate that the state trajectories are invariant to time-differential transformations of the path parametrizing coordinate. For one this property allows to impose arbitrary velocity profiles on the slider without impacting the geometry of the state trajectory. This property implies that certain path planning problems may be decomposed approximately into a strictly geometric path planning and an auxiliary throughput speed optimization problem. Building on these insights we elaborate a numerical approach tailored to constrained time optimal collision free path planning and apply it to the slider-pusher system.

math.DS

A Supervisory Learning Control Framework for Autonomous & Real-time Task Planning for an Underactuated Cooperative Robotic task

We introduce a framework for cooperative manipulation, applied on an underactuated manipulation problem. Two stationary robotic manipulators are required to cooperate in order to reposition an object within their shared work space. Control of multi-agent systems for manipulation tasks cannot rely on individual control strategies with little to no communication between the agents that serve the common objective through swarming. Instead a coordination strategy is required that queries subtasks to the individual agents. We formulate the problem in a Task And Motion Planning (TAMP) setting, while considering a decomposition strategy that allows us to treat the task and motion planning problems separately. We solve the supervisory planning problem offline using deep Reinforcement Learning techniques resulting into a supervisory policy capable of coordinating the two manipulators into a successful execution of the pick-and-place task. Additionally, a benefit of solving the task planning problem offline is the possibility of real-time (re)planning, demonstrating robustness in the event of subtask execution failure or on-the-fly task changes. The framework achieved zero-shot deployment on the real setup with a success rate that is higher than 90%.

cs.RO

Probabilistic Control and Majorization of Optimal Control

Probabilistic control design is founded on the principle that a rational agent attempts to match modelled with an arbitrary desired closed-loop system trajectory density. The framework was originally proposed as a tractable alternative to traditional optimal control design, parametrizing desired behaviour through fictitious transition and policy densities and using the information projection as a proximity measure. In this work we introduce an alternative parametrization of desired closed-loop behaviour and explore alternative proximity measures between densities. It is then illustrated how the associated probabilistic control problems solve into uncertain or probabilistic policies. Our main result is to show that the probabilistic control objectives majorize conventional, stochastic and risk sensitive, optimal control objectives. This observation allows us to identify two probabilistic fixed point iterations that converge to the deterministic optimal control policies establishing an explicit connection between either formulations. Further we demonstrate that the risk sensitive optimal control formulation is also technically equivalent to a Maximum Likelihood estimation problem on a probabilistic graph model where the notion of costs is directly encoded into the model. The associated treatment of the estimation problem is then shown to coincide with the moment projected probabilistic control formulation. That way optimal decision making can be reformulated as an iterative inference problem. Based on these insights we discuss directions for algorithmic development.

cs.LG

Information-Theoretic Policy Learning from Partial Observations with Fully Informed Decision Makers

In this work we formulate and treat an extension of the Imitation from Observations problem. Imitation from Observations is a generalisation of the well-known Imitation Learning problem where state-only demonstrations are considered. In our treatment we extend the scope of Imitation from Observations to feature-only demonstrations which could arguably be described as partial observations. Therewith we mean that the full state of the decision makers is unknown and imitation must take place on the basis of a limited set of features. We set out for methods that extract an executable policy directly from those features which, in the literature, would be referred to as Behavioural Cloning methods. Our treatment combines elements from probability and information theory and draws connections with entropy regularized Markov Decision Processes.

eess.SY

Adaptive control of a mechatronic system using constrained residual reinforcement learning

We propose a simple, practical and intuitive approach to improve the performance of a conventional controller in uncertain environments using deep reinforcement learning while maintaining safe operation. Our approach is motivated by the observation that conventional controllers in industrial motion control value robustness over adaptivity to deal with different operating conditions and are suboptimal as a consequence. Reinforcement learning on the other hand can optimize a control signal directly from input-output data and thus adapt to operational conditions, but lacks safety guarantees, impeding its use in industrial environments. To realize adaptive control using reinforcement learning in such conditions, we follow a residual learning methodology, where a reinforcement learning algorithm learns corrective adaptations to a base controller's output to increase optimality. We investigate how constraining the residual agent's actions enables to leverage the base controller's robustness to guarantee safe operation. We detail the algorithmic design and propose to constrain the residual actions relative to the base controller to increase the method's robustness. Building on Lyapunov stability theory, we prove stability for a broad class of mechatronic closed-loop systems. We validate our method experimentally on a slider-crank setup and investigate how the constraints affect the safety during learning and optimality after convergence.

eess.SY

Entropy Regularised Deterministic Optimal Control: From Path Integral Solution to Sample-Based Trajectory Optimisation

Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the optimal policy can be expressed in terms of an expectation over stochastic paths. By estimating the expectation with Monte Carlo sampling and reinterpreting the process as exploration noise, a stochastic search algorithm is obtained tailored to (deterministic) trajectory optimisation. For the purpose of future algorithmic development, it is essential to properly understand the underlying theoretical foundations that allow for a principled derivation of such methods. In this paper we make a connection between entropy regularisation in optimisation and deterministic optimal control. We then show that the optimal policy is given by a belief function rather than a deterministic function. The policy belief is governed by a Bayesian-type update where the likelihood can be expressed in terms of a conditional expectation over paths induced by a prior policy. Our theoretical investigation firmly roots sample based trajectory optimisation in the larger family of control as inference. It allows us to justify a number of heuristics that are common in the literature and motivate a number of new improvements that benefit convergence.

cs.RO

Risk Sensitive Path Integral Control for Infinite Horizon Problem Formulations

Path Integral Control methods were developed for stochastic optimal control covering a wide class of finite horizon formulations with control affine nonlinear dynamics. Characteristic for this class is that the HJB equation is linear and consequently the value function can be expressed as a conditional expectation of the exponentially weighted cost-to-go evaluated over trajectories with uncontrolled system dynamics, hence the name. Subsequently it was shown that under the same assumptions Path Integral Control generalises to finite horizon risk sensitive stochastic optimal control problems. Here we study whether the HJB of infinite horizon formulations can be made linear as well. Our interest in infinite horizon formulations is motivated by the stationarity of the associated value function and their inherent dynamic stability seeking nature. Technically a stationary value function may ease the solution of the associated linear HJB. Second we argue this may offer an interesting starting point for off-linear Reinforcement Learning applications. We show formally that the discounted and average cost formulations are respectively intractable and tractable.

math.OC