SearcharxivSearch

arXiv subjects

Juho Bae

Publications and source records attributed to Juho Bae.

6 recordsLinked to original sources

PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by reward and transition similarity. In practice, the choice of the latent distance can strongly affect performance: using a fixed, pre-specified global norms (e.g., $\ell_p$ norms or other hand-designed metrics) may be overly restrictive to capture the behavioral distance. In contrast, unconstrained pairwise distances may admit degenerate solutions that drive the metric loss down without improving the representation. To address this gap, we introduce **PAMD: Pairwise Adaptive Mahalanobis Distance**, which parameterizes a positive-definite, pair-conditioned metric for measuring latent state similarity. PAMD is a simple plug-in for existing bisimulation-based methods, offering a more expressive yet structured alternative to fixed, pre-specified latent distances. We empirically validate our method on visual MuJoCo continuous-control tasks, where final performance of several recent bisimulation-based RL algorithms is substantially improved when equipped with the distance we propose.

cs.AI

Feedback Integrators: Non-Asymptotic Invariance for One-Step Methods and Gain Selection under Euler Discretization

For dynamical systems evolving on a manifold and admitting first integrals, standard one-step numerical methods generally cause the discrete trajectory to drift off the manifold and the numerical values of the first integrals to deviate from their prescribed values. Feedback integrators address this by extending the dynamics to an ambient Euclidean space and adding a feedback term that drives the numerical trajectory toward the set satisfying both the manifold constraint and the prescribed values of the first integrals. Existing theory, however, has two limitations: it remains asymptotic, guaranteeing only eventual entrance into an attractor containing the desired set, and it does not explain how the feedback gain should be chosen. In this paper, we first close the former gap for general one-step methods by proving positive invariance of arbitrarily small sublevel neighborhoods of the feedback Lyapunov function for sufficiently small step sizes. We then specialize to Euler discretization and analyze how the feedback gain enters the Taylor-based error bound. In this setting, we characterize a range of scaled gains that guarantee positive invariance for sufficiently small step sizes and identify the scaling that minimizes the Taylor-based upper bound. We further propose adaptive gain-selection rules under Euler discretization, including both stepwise and periodically updated variants, and establish corresponding boundedness guarantees for the resulting discrete trajectories. These results identify Euler discretization as the first setting in which gain selection for feedback integrators closes in explicit form, whereas extensions to general higher-order one-step methods remain genuinely method-dependent. Numerical experiments on free rigid body motion in $\operatorname{SO}(3)$, the Kepler problem, and a perturbed Kepler problem with rotational symmetry support the analysis.

math.NA

Learning Safety-Compatible Observers for Unknown Systems

This paper presents a data-driven approach for jointly learning a robust full-state observer and its robustness certificate for systems with unknown dynamics. Leveraging incremental input-to-state stability (delta ISS) notions, we jointly learn a delta ISS Lyapunov function that serves as the robustness certificate and prove practical convergence of the estimation error under standard fidelity assumptions on the learned models. This renders the observer safety-compatible: they can be consumed by certificate-based safe controllers so that, when the controller tolerates bounded estimation error, the controller's certificate remains valid under output feedback. We further extend the approach to interconnected systems via the small-gain theorem, yielding a distributed observer design framework. We validate the approach on a variety of nonlinear systems.

eess.SY

Fixed-Time State Transfer via Pontryagin Extremals

This paper concerns the problem of fixed-time transition between two states of nonlinear systems (i.e., the point-to-point steering problem). We propose a formulation applicable to a broad class of nonlinear systems, and show that it is theoretically complete in the sense that it admits a solution if and only if the target state is reachable from the initial state. When the target state is reachable, we prove that a solution can always be constructed by concatenation of two Pontryagin extremals, one generated by the original dynamics $f$ from the initial state and the other generated by the inverted dynamics $-f$ from the terminal state. This allows the problem to be formulated as a two-point boundary value problem (TPBVP) of extremals, where the solution existence to the formulated TPBVP is equivalent to that of the original problem. The theoretical developments are applied to curves with prescribed curvature bounds in $\mathbb{R}^3$, thereby extending the recent works on Dubins car to dimension three. We prove that to construct a curvature-bounded path in $\mathbb{R}^3$ with prescribed length and boundary conditions, it suffices to consider the trajectories that are concatenations of CSC, CCC, their subsegments, and H, where C denotes a circular arc with maximum curvature, S a straight line segment, and H a certain class of helicoidal arcs with constant curvature. Numerical demonstrations are conducted on a nonlinear dynamics example, and on curvature-bounded paths in $\mathbb{R}^2$ and $\mathbb{R}^3$.

math.OC

On the Reachability of 3-Dimensional Paths with a Prescribed Curvature Bound

This paper presents the reachability analysis of curves in $\mathbb{R}^3$ with a prescribed curvature bound. Based on Pontryagin Maximum Principle, we leverage the existing knowledge on the structure of solutions to minimum-time problems, or Markov-Dubins problem, to reachability considerations. Based on this development, two types of reachability are discussed. First, we prove that any boundary point of the reachability set, with the directional component taken into account as well as geometric coordinates, can be reached via curves of H, CSC, CCC, or their respective subsegments, where H denotes a helicoidal arc, C a circular arc with maximum curvature, and S a straight segment. Second, we show that the reachability set when directional component is not considered\textemdash{}the position reachability set\textemdash{}is simply a solid of revolution of its two-dimensional counterpart, the Dubins car. These findings extend the developments presented in literature on Dubins car into spatial curves in $\mathbb{R}^3$.

math.OC

Constraint-Aware Mesh Refinement Method by Reachability Set Envelope of Curvature Bounded Paths

This paper presents an enhanced direct-method-based approach for the real-time solution of optimal control problems to handle path constraints, such as obstacles. The principal contributions of this work are twofold: first, the existing methods for constructing reachability sets in the literature are extended to derive the envelope of these sets, which determines the region swept by all feasible trajectories between adjacent sample points. Second, we propose a novel method to guarantee constraint violation-free between discrete states in two dimensions through mesh refinement approach. To illustrate the effectiveness of the proposed methodology, numerical simulations are conducted on real-time path planning for fixed-wing unmanned aerial vehicles.

eess.SY