SearcharxivSearch

arXiv subjects

Joshua Aurand

Publications and source records attributed to Joshua Aurand.

8 recordsLinked to original sources

Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning

Legged robots are well-suited for navigating terrains inaccessible to wheeled robots, making them ideal for applications in search and rescue or space exploration. However, current control methods often struggle to generalize across diverse, unstructured environments. This paper introduces a novel framework for agile locomotion of legged robots by combining multi-expert distillation with reinforcement learning (RL) fine-tuning to achieve robust generalization. Initially, terrain-specific expert policies are trained to develop specialized locomotion skills. These policies are then distilled into a unified foundation policy via the DAgger algorithm. The distilled policy is subsequently fine-tuned using RL on a broader terrain set, including real-world 3D scans. The framework allows further adaptation to new terrains through repeated fine-tuning. The proposed policy leverages depth images as exteroceptive inputs, enabling robust navigation across diverse, unstructured terrains. Experimental results demonstrate significant performance improvements over existing methods in synthesizing multi-terrain skills into a single controller. Deployment on the ANYmal D robot validates the policy's ability to navigate complex environments with agility and robustness, setting a new benchmark for legged robot locomotion.

cs.RO

Close-Proximity Satellite Operations through Deep Reinforcement Learning and Terrestrial Testing Environments

With the increasingly congested and contested space environment, safe and effective satellite operation has become increasingly challenging. As a result, there is growing interest in autonomous satellite capabilities, with common machine learning techniques gaining attention for their potential to address complex decision-making in the space domain. However, the "black-box" nature of many of these methods results in difficulty understanding the model's input/output relationship and more specifically its sensitivity to environmental disturbances, sensor noise, and control intervention. This paper explores the use of Deep Reinforcement Learning (DRL) for satellite control in multi-agent inspection tasks. The Local Intelligent Network of Collaborative Satellites (LINCS) Lab is used to test the performance of these control algorithms across different environments, from simulations to real-world quadrotor UAV hardware, with a particular focus on understanding their behavior and potential degradation in performance when deployed beyond the training environment.

cs.RO

Assessing Autonomous Inspection Regimes: Active Versus Passive Satellite Inspection

This paper addresses the problem of satellite inspection, where one or more satellites (inspectors) are tasked with imaging or inspecting a resident space object (RSO) due to potential malfunctions or anomalies. Inspection strategies are often reduced to a discretized action space with predefined waypoints, facilitating tractability in both classical optimization and machine learning based approaches. However, this discretization can lead to suboptimal guidance in certain scenarios. This study presents a comparative simulation to explore the tradeoffs of passive versus active strategies in multi-agent missions. Key factors considered include RSO dynamic mode, state uncertainty, unmodeled entrance criteria, and inspector motion types. The evaluation is conducted with a focus on fuel utilization and surface coverage. Building on a Monte-Carlo based evaluator of passive strategies and a reinforcement learning framework for training active inspection policies, this study investigates conditions under which passive strategies, such as Natural Motion Circumnavigation (NMC), may perform comparably to active strategies like Reinforcement Learning based waypoint transfers.

cs.RO

Stability Analysis of Deep Reinforcement Learning for Multi-Agent Inspection in a Terrestrial Testbed

The design and deployment of autonomous systems for space missions require robust solutions to navigate strict reliability constraints, extended operational duration, and communication challenges. This study evaluates the stability and performance of a hierarchical deep reinforcement learning (DRL) framework designed for multi-agent satellite inspection tasks. The proposed framework integrates a high-level guidance policy with a low-level motion controller, enabling scalable task allocation and efficient trajectory execution. Experiments conducted on the Local Intelligent Network of Collaborative Satellites (LINCS) testbed assess the framework's performance under varying levels of fidelity, from simulated environments to a cyber-physical testbed. Key metrics, including task completion rate, distance traveled, and fuel consumption, highlight the framework's robustness and adaptability despite real-world uncertainties such as sensor noise, dynamic perturbations, and runtime assurance (RTA) constraints. The results demonstrate that the hierarchical controller effectively bridges the sim-to-real gap, maintaining high task completion rates while adapting to the complexities of real-world environments. These findings validate the framework's potential for enabling autonomous satellite operations in future space missions.

cs.RO

The Safe Trusted Autonomy for Responsible Space Program

The Safe Trusted Autonomy for Responsible Space (STARS) program aims to advance autonomy technologies for space by leveraging machine learning technologies while mitigating barriers to trust, such as uncertainty, opaqueness, brittleness, and inflexibility. This paper presents the achievements and lessons learned from the STARS program in integrating reinforcement learning-based multi-satellite control, run time assurance approaches, and flexible human-autonomy teaming interfaces, into a new integrated testing environment for collaborative autonomous satellite systems. The primary results describe analysis of the reinforcement learning multi-satellite control and run time assurance algorithms. These algorithms are integrated into a prototype human-autonomy interface using best practices from human-autonomy trust literature, however detailed analysis of the effectiveness is left to future work. References are provided with additional detailed results of individual experiments.

eess.SY

Epstein-Zin Utility Maximization on a Random Horizon

This paper solves the consumption-investment problem under Epstein-Zin preferences on a random horizon. In an incomplete market, we take the random horizon to be a stopping time adapted to the market filtration, generated by all observable, but not necessarily tradable, state processes. Contrary to prior studies, we do not impose any fixed upper bound for the random horizon, allowing for truly unbounded ones. Focusing on the empirically relevant case where the risk aversion and the elasticity of intertemporal substitution are both larger than one, we characterize the optimal consumption and investment strategies using backward stochastic differential equations with superlinear growth on unbounded random horizons. This characterization, compared with the classical fixed-horizon result, involves an additional stochastic process that serves to capture the randomness of the horizon. As demonstrated in two concrete examples, changing from a fixed horizon to a random one drastically alters the optimal strategies.

q-fin.MF

Exposure-Based Multi-Agent Inspection of a Tumbling Target Using Deep Reinforcement Learning

As space becomes more congested, on orbit inspection is an increasingly relevant activity whether to observe a defunct satellite for planning repairs or to de-orbit it. However, the task of on orbit inspection itself is challenging, typically requiring the careful coordination of multiple observer satellites. This is complicated by a highly nonlinear environment where the target may be unknown or moving unpredictably without time for continuous command and control from the ground. There is a need for autonomous, robust, decentralized solutions to the inspection task. To achieve this, we consider a hierarchical, learned approach for the decentralized planning of multi-agent inspection of a tumbling target. Our solution consists of two components: a viewpoint or high-level planner trained using deep reinforcement learning and a navigation planner handling point-to-point navigation between pre-specified viewpoints. We present a novel problem formulation and methodology that is suitable not only to reinforcement learning-derived robust policies, but extendable to unknown target geometries and higher fidelity information theoretic objectives received directly from sensor inputs. Operating under limited information, our trained multi-agent high-level policies successfully contextualize information within the global hierarchical environment and are correspondingly able to inspect over 90% of non-convex tumbling targets, even in the absence of additional agent attitude control.

cs.RO

Mortality and Healthcare: a Stochastic Control Analysis under Epstein-Zin Preferences

This paper studies optimal consumption, investment, and healthcare spending under Epstein-Zin preferences. Given consumption and healthcare spending plans, Epstein-Zin utilities are defined over an agent's random lifetime, partially controllable by the agent as healthcare reduces mortality growth. To the best of our knowledge, this is the first time Epstein-Zin utilities are formulated on a controllable random horizon, via an infinite-horizon backward stochastic differential equation with superlinear growth. A new comparison result is established for the uniqueness of associated utility value processes. In a Black-Scholes market, the stochastic control problem is solved through the related Hamilton-Jacobi-Bellman (HJB) equation. The verification argument features a delicate containment of the growth of the controlled morality process, which is unique to our framework, relying on a combination of probabilistic arguments and analysis of the HJB equation. In contrast to prior work under time-separable utilities, Epstein-Zin preferences facilitate calibration. The model-generated mortality closely approximates actual mortality data in the US and UK; moreover, the efficacy of healthcare can be calibrated and compared between the two countries.

q-fin.MF