Searcharxiv⌕ Search

arXiv subjects

Isaac E. Weintraub

Publications and source records attributed to Isaac E. Weintraub.

At least 19 recordsLinked to original sources

Combining Reinforcement Learning with Arc-search Interior-Point Method for Path Planning

Path planning in environments containing obstacles has numerous practical applications. The problem is challenging because it is inherently nonlinear and nonconvex. Consequently, a variety of techniques have been developed to address this problem, among which machine learning and optimal control (or optimization) have emerged as two prominent approaches. In general, machine learning methods do not require a high-fidelity model, and a trained agent can often generate a feasible path in real time. However, the resulting path is not necessarily optimal with respect to performance objectives such as minimizing path length or travel time. In contrast, optimal control and optimization methods typically rely on high-fidelity models and often require computational effort that may not satisfy real-time constraints. Nevertheless, these methods are more likely to produce optimal or near-optimal solutions. To overcome the limitations of each approach while exploiting their respective strengths, this paper proposes a framework that combines reinforcement learning with an arc-search interior-point method for path planning. Numerical simulations demonstrate that the proposed approach effectively integrates the real-time decision-making capability of reinforcement learning with the optimization performance of the arc-search interior-point method, resulting in improved path-planning performance.

math.OC↗

Path Planning Using Deep Deterministic Policy Gradient: A Reinforcement Learning Approach

Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge because the problem is nonlinear and nonconvex even in simplest scenarios. While traditional optimal control methods can be used to find ideal paths, the computational time is often too slow for real-time decision-making. To solve this challenge, we propose a method based on Deep Deterministic Policy Gradient (DDPG) and model the threat as possibly multiple circular 'no-go' zones. A mission is regarded as a failure if the vehicle enters this restricted zone at any time or does not reach a neighborhood of the destination. The DDPG agent is trained through trial and error in a simulated environment, learning a direct mapping from its current state (position and heading) to a series of feasible actions that guide the agent to safely reach its destination. The reword function has three parts: (a) an attractive field centered at the final destination, (b) some repulsive fields centered at the origins of circular obstacles, and (c) a penalty of control energy consumption (the magnitude of heading change) that indirectly in favor for straight path. The DDPG trains the agent using these incentives to find the largest possible set of starting points wherein a safe path to the destination is guaranteed. This provides critical information for mission planning, showing beforehand whether a task is achievable from a given starting point, assisting pre-mission planning activities. The approach is validated in simulation. A comparison between the DDPG method and a traditional optimal control (pseudo-spectral) method is carried out. The results show that the learning-based agent produces effective paths while being significantly faster, making it a better fit for real-time applications.

cs.RO↗

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find ideal paths, the computational time is often too slow for real-time decision-making. To solve this challenge, we propose a method based on Deep Deterministic Policy Gradient (DDPG) and model the threat as a simple, circular `no-go' zone. A mission failure is claimed if the vehicle enters this `no-go' zone at any time or does not reach a neighborhood of the destination. The DDPG agent is trained to learn a direct mapping from its current state (position and velocity) to a series of feasible actions that guide the agent to safely reach its goal. A reward function and two neural networks, critic and actor, are used to describe the environment and guide the control efforts. The DDPG trains the agent to find the largest possible set of starting points (``feasible set'') wherein a safe path to the goal is guaranteed. This provides critical information for mission planning, showing beforehand whether a task is achievable from a given starting point, assisting pre-mission planning activities. The approach is validated in simulation. A comparison between the DDPG method and a traditional optimal control (pseudo-spectral) method is carried out. The results show that the learning-based agent may produce effective paths while being significantly faster, making it a better fit for real-time applications. However, there are areas (``infeasible set'') where the DDPG agent cannot find paths to the destination, and the paths in the feasible set may not be optimal. These preliminary results guide our future research: (1) improve the reward function to enlarge the DDPG feasible set, (2) examine the feasible set obtained by the pseudo-spectral method, and (3) investigate the arc-search IPM method for the path planning problem.

math.OC↗

Safe Navigation in the Presence of Range-Limited Pursuers

This paper examines the degree to which an evader seeking a safe and efficient path to a target location can benefit from increasing levels of knowledge regarding one or more range-limited pursuers seeking to intercept it. Unlike previous work, this research considers the time of flight of the pursuers actively attempting interception. It is shown that additional knowledge allows the evader to safely steer closer to the threats, shortening paths without accepting additional risk of capture. A control heuristic is presented, suitable for real-time implementation, which capitalizes on all knowledge available to the evader.

math.OC↗

Probabilistic Weapon Engagement Zones for a Turn Constrained Pursuer

Curve-straight probabilistic engagement zones (CSPEZ) quantify the spatial regions an evader should avoid to reduce capture risk from a turn-rate-limited pursuer following a curve-straight path with uncertain parameters including position, heading, velocity, range, and maximum turn rate. This paper presents methods for generating evader trajectories that minimize capture risk under such uncertainty. We first derive an analytic solution for the deterministic curve-straight basic engagement zone (CSBEZ), then extend this formulation to a probabilistic framework using four uncertainty-propagation approaches: Monte Carlo sampling, linearization, quadratic approximation, and neural-network regression. We evaluate the accuracy and computational cost of each approximation method and demonstrate how CSPEZ constraints can be integrated into a trajectory-optimization algorithm to produce safe paths that explicitly account for pursuer uncertainty.

cs.RO↗

Engagement Zones for a Turn Constrained Pursuer

This work derives two basic engagement zone models, describing regions of potential risk or capture for a mobile vehicle by a pursuer. The pursuer is modeled as having turn-constraints rather than simple motion. Turn-only (C-Paths) and turn-straight (CS-Paths) paths are considered for the pursuer of limited range. Following the derivation, a simulation of a vehicle avoiding the pursuer's engagement zone is provided.

math.OC↗

Min-Time Escape of a Dubins Car from a Polygon

A turn constrained vehicle is initially located inside a polygon region and desires to escape in minimum time. First, the method of characteristics is used to describe the time-optimal strategies for reaching a line of infinite length. Next, the approach is extended to polygons constructed of a series of line segments. Using this construction technique, the min-time path to reach each edge is obtained; the resulting minimum of the set of optimal trajectories is then selected for escaping the polygon.

eess.SY↗

Multi Agent Pathfinding for Noise Restricted Hybrid Fuel Unmanned Aerial Vehicles

Multi Agent Path Finding (MAPF) seeks the optimal set of paths for multiple agents from respective start to goal locations such that no paths conflict. We address the MAPF problem for a fleet of hybrid-fuel unmanned aerial vehicles which are subject to location-dependent noise restrictions. We solve this problem by searching a constraint tree for which the subproblem at each node is a set of shortest path problems subject to the noise and fuel constraints and conflict zone avoidance. A labeling algorithm is presented to solve this subproblem, including the conflict zones which are treated as dynamic obstacles. We present the experimental results of the algorithms for various graph sizes and number of agents.

math.OC↗

Pursuit-Evasion on a Sphere and When It Can Be Considered Flat

In classical works on a planar differential pursuit-evasion game with a faster pursuer, the intercept point resulting from the equilibrium strategies lies on the Apollonius circle. This property was exploited for the construction of the equilibrium strategies for two faster pursuers against one evader. Extensions for planar multiple-pursuer single-evader scenarios have been considered. We study a pursuit-evasion game on a sphere and the relation of the equilibrium intercept point to the Apollonius domain on the sphere. The domain is a generalization of the planar Apollonius circle set. We find a condition resulting in the intercept point belonging to the Apollonius domain, which is the characteristic of the planar game solution. Finally, we use this characteristic to discuss pursuit and evasion strategies in the context of two pursuers and a single slower evader on the sphere and illustrate it using numerical simulations.

math.OC↗

Sampling-Based Risk-Aware Path Planning Around Dynamic Engagement Zones

Existing methods for avoiding dynamic engagement zones (EZs) and minimizing risk leverage the calculus of variations to obtain optimal paths. While such methods are deterministic, they scale poorly as the number of engagement zones increases. Furthermore, optimal-control based strategies are sensitive to initial guesses and often converge to local, rather than global, minima. This paper presents a novel sampling-based approach to obtain a feasible flight plan for a Dubins vehicle to reach a desired location in a bounded operating region in the presence of a large number of engagement zones. The dynamic EZs are coupled to the vehicle dynamics through its heading angle. Thus, the dynamic two-dimensional obstacles in the (x,y) plane can be transformed into three-dimensional static obstacles in a lifted (x,y,ψ) space. This insight is leveraged in the formulation of a Rapidly-exploring Random Tree (RRT*) algorithm. The algorithm is evaluated with a Monte Carlo experiment that randomizes EZ locations to characterize the success rate and average path length as a function of the number of EZs and as the computation time made available to the planner is increased.

math.OC↗

Noise Aware Path Planning and Power Management of Hybrid Fuel UAVs

Hybrid fuel Unmanned Aerial Vehicles (UAV), through their combination of multiple energy sources, offer several advantages over the standard single fuel source configuration, the primary one being increased range and efficiency. Multiple power or fuel sources also allow the distinct pitfalls of each source to be mitigated while exploiting the advantages within the mission or path planning. We consider here a UAV equipped with a combustion engine-generator and battery pack as energy sources. We consider the path planning and power-management of this platform in a noise-aware manner. To solve the path planning problem, we first present the Mixed Integer Linear Program (MILP) formulation of the problem. We then present and analyze a label-correcting algorithm, for which a pseudo-polynomial running time is proven. Results of extensive numerical testing are presented which analyze the performance and scalability of the labeling algorithm for various graph structures, problem parameters, and search heuristics. It is shown that the algorithm can solve instances on graphs as large as twenty thousand nodes in only a few seconds.

math.OC↗

Virtual Target Selection for a Multiple-Pursuer Multiple-Evader Scenario

This paper considers an M-pursuer N-evader scenario involving virtual targets. The virtual targets serve as an intermediary target for the pursuers, allowing the pursuers to delay their final assignment to the evaders. However, upon reaching the virtual target, the pursuers must decide which evader to capture. It is assumed that there are more pursuers than evaders and that the pursuers are faster than the evaders. The objective is two-part: first, assign each pursuer to a virtual target and evader such that the pursuer team's energy is minimized, and second, choose the virtual targets' locations for this minimization problem. The approach taken is to consider the Apollonius geometry between each pursuer's virtual target location and each evader. Using the constructed Apollonius circles, the pursuer's travel distance and maneuver at a virtual target are obtained. These metrics serve as a gauge for the total energy required to capture a particular evader and are used to solve the joint virtual target selection and pursuer-evader assignment problem. This paper provides a mathematical definition of this problem, the solution approach taken, and an example.

math.OC↗

Basic Engagement Zones

This paper establishes a more formal definition for an Engagement Zone (EZ) and derives some basic EZs associated with fundamental engagement models associated with pursuit-evasion and turret-evasion. The basic EZs presented in this paper capture the most salient aspects of the Pursuer-Agent and Turret-Agent engagements: namely the geometry of the aspect angle and the relative differences in capability (i.e., maximum speeds, range, etc.). One of the main advantages of utilizing EZs for path planning is that they encode an overall desire for Agent to go somewhere without requiring an aggressive maneuver or active evasion should the Pursuer or Turret begin its pursuit. It is shown that there is some advantage, in terms of time savings, in EZ-based navigation around a single range-limited Pursuer as compared with circumnavigating the capturability region.

math.OC↗

Optimal Trajectories for Multiple-UAS Simultaneous Target Acquisition with Obstacle Avoidance

This work develops feasible path trajectories for a coordinated strike with multiple aircraft in a constrained environment. Using direct orthogonal collocation methods, the two-point boundary value optimal control problem is transcribed into a nonlinear programming problem. A coordinate transformation is performed on the state variables to leverage the benefits of a simplex discretization of the search domain. Applying these techniques allows each path constraint to be removed from the feasible search space, eliminating computationally expensive, nonlinear constraint equations and problem specific parameters from the optimal control formulation. Heuristic search techniques are used to determine a Dubins path solution through the space to seed the optimal control solver. In the scenario, three aircraft are initiated in separate directions and are required to avoid all constrained regions while simultaneously arriving at the target location, each with a different viewing angle. A focus of this work is to reduce computation times for optimal control solvers such that real-time solutions can be implemented onboard small unmanned aircraft systems. Analysis of the problem examines optimal flight paths through simplex corridors, velocity and heading vectors, control vectors of acceleration and heading rate, and objective times for minimum time flight.

math.OC↗

Development of Linear Battery Model for Path Planning with Mixed Integer Linear Programming: Simulated and Experimental Validation

Mixed Integer Linear Programs (MILPs) are often used in the path planning of both ground and aerial vehicles. Such a formulation of the path planning problem requires a linear objective function and constraints, limiting the fidelity of the the tracking of vehicle states. One such parameter is the state of charge of the battery used to power the vehicle. Accurate battery state estimation requires nonlinear differential equations to be solved. This state estimation is important in path planning to ensure flyable paths, however when using MILPs to formulate the path planning problem these nonlinear equations cannot be implemented. Poor accuracy in battery estimation during the path planning runs the risk of the planned path being feasible by the estimation model but in reality will deplete the battery to a critical level. To the end of higher accuracy battery estimation within a MILP, we present here a simple linear battery model which predicts the change in state-of-charge (SOC) of a battery given a power draw and duration. This model accounts for changes in battery voltage due to applied electrical load and changes in battery SOC. The battery model is presented and then tested against alternate battery models in numerical and in experimental tests. Further, the effect the proposed linear model has over a simpler SOC estimation on the time-to-solve a resource constrained shortest path problem is also evaluated, implemented in two different algorithms. It is seen that the linear model performs well in battery state estimation while remaining implementable in a Linear Program or MILP, with little affect on the time-to-solve.

math.OC↗

Determining Follower Aircraft's Optimal Trajectory in Relation to a Dynamic Formation Ring

The specific objective of this paper is to develop a tool that calculates the optimal trajectory of the follower aircraft as it completes a formation rejoin, and then maintains the formation position, defined as a ring of points, until a fixed final time. The tool is designed to produce optimal trajectories for a variety of initial conditions and leader trajectories. Triple integrator dynamics are used to model the follower aircraft in three dimensions. Control is applied directly to the rate of acceleration. Both the follower's and leader's velocities and accelerations are bounded, as dictated by the aircraft's performance envelope. Lastly, a path constraint is used to ensure the follower avoids the leader's jet wash region. This optimal control problem is solved through numerical analysis using the direct orthogonal collocation solver GPOPS-II. Two leader trajectories are investigated, including a descending spiral and continuous vertical loops. Additionally, a study of the effect of various initial guesses is performed. All trajectories displayed a direct capture of the formation position, however changes in solver initial conditions demonstrate various behaviors in how the follower maintains the formation position. The developed tool has proven adequate to support future research in crafting real-time controllers capable of determining near-optimal trajectories.

math.OC↗

Surveillance of a Faster Fixed-Course Target

The maximum surveillance of a target which is holding course is considered, wherein an observer vehicle aims to maximize the time that a faster target remains within a fixed-range of the observer. This entails two coupled phases: an approach phase and observation phase. In the approach phase, the observer strives to make contact with the faster target, such that in the observation phase, the observer is able to maximize the time where the target remains within range. Using Pontryagin's Minimum Principle, the optimal control laws for the observer are found in closed-form. Example scenarios highlight various aspects of the engagement.

math.OC↗