SearcharxivSearch

arXiv subjects

Kenta Hoshino

Publications and source records attributed to Kenta Hoshino.

11 recordsLinked to original sources

Contract-Based Decomposition of Temporal Logic Specifications for Networked Systems under Arbitrary Partitions

Computational complexity is an inherent limitation of formal synthesis for networked systems, and decomposing the global specification into local ones relaxes this limitation at the cost of conservatism. Since the granularity of the partition governs this trade-off, it is reasonable to treat the partition as a design variable, which calls for local specifications that remain correct for every partition. To this end, this paper gives each agent a local specification, written as an assume-guarantee contract that the agent can establish from local information. We first derive a necessary and sufficient condition for these contracts to decompose the global specification under a given partition. Building on this, we then present a condition under which the decomposition is correct for every partition, so that the partition becomes a free design variable. For linear dynamics and signal temporal logic formulas with affine predicates, we further synthesize a controller for each coalition by a tube-based approach. Finally, simulations on a network of input-coupled tanks show how the choice of partition trades computational cost against conservatism.

eess.SY

World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

World Action Models (WAMs) aim to control robots by stochastically generating visual futures and then decoding actions, but empirical observations indicate that the results can strongly depend on which future is selected. We propose World-Coherent-Decoding (WCD), a self-verifying test-time planning framework that treats WAM rollouts as falsifiable future--action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using internal generative signals: flow-based video surprisal for visual plausibility and action path effort for action-generation stability. After execution, the realized observation audits the selected imagination, yielding an imagination--reality mismatch that trains a lightweight online predictor for future candidate selection. Thus, WCD converts delayed self-verification into pre-execution reliability estimation without updating the backbone model. On RoboTwin 2.0, WCD improves Hard success under limited randomized-scene supervision from $55.80\%$ to $60.90\%$, with a $+16.43$ gains on Horizon-3 tasks, and shows qualitative robustness on real Franka visual-shift tests. These results highlight a simple principle: test-time scaling for WAMs depends less on sampling more futures than on selecting reliable ones.

cs.CV

Nonlinear receding-horizon differential game for drone racing along a three-dimensional path

Drone racing requires high-speed navigation through three-dimensional paths, posing significant challenges in control engineering. Existing control methods lack a feedback control framework that simultaneously addresses nonlinear drone dynamics and multi-agent competitive interactions, such as overtaking or obstructing opponents. To overcome this limitation, this study proposes a game-theoretic control framework, the nonlinear receding-horizon differential game (NRHDG), for competitive drone racing. NRHDG accounts explicitly for adversarial behavior by predicting and countering an opponent's worst-case behavior in real time. It extends standard nonlinear model predictive control (NMPC), which typically assumes a fixed opponent model. First, we develop a novel path-following formulation based on projection-point dynamics, eliminating the need for computationally expensive distance minimization during online control. Second, we propose a potential function that enables each drone to dynamically switch between overtaking and obstructing maneuvers, depending on the race situation. Third, we establish new performance metrics to evaluate NRHDG against NMPC across racing scenarios. Simulation results demonstrate that NRHDG outperforms NMPC in both overtaking and obstructing performance. Specifically, for randomly generated initial conditions and different levels of speed advantage for the rear-start drone, the 95\% confidence intervals for the arc-length-based mean performance differences excluded zero, indicating statistically significant advantages of NRHDG over NMPC in both overtaking and obstructing.

eess.SY

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form

Designing a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust constrained MDP (RCMDP), where an optimal policy minimizes cumulative cost while satisfying constraints in the worst-case scenario across a set of environments. We first prove that the conventional policy gradient approach to the Lagrangian max-min formulation can become trapped in suboptimal solutions. This occurs when its inner minimization encounters a sum of conflicting gradients from the objective and constraint functions. To address this, we leverage the epigraph form of the RCMDP problem, which resolves the conflict by selecting a single gradient from either the objective or the constraints. Building on the epigraph form, we propose a bisection search algorithm with a policy gradient subroutine and prove that it identifies an $\varepsilon$-optimal policy in an RCMDP with $\tilde{\mathcal{O}}(\varepsilon^{-4})$ robust policy evaluations.

cs.LG

Stochastic Safety-critical Control Compensating Safety Probability for Marine Vessel Tracking

A marine vessel is a nonlinear system subject to irregular disturbances such as wind and waves, which cause tracking errors between the nominal and actual trajectories. In this study, a nonlinear vessel maneuvering model that includes a tracking controller is formulated and then controlled using a linear approximation around the nominal trajectory. The resulting stochastic linearized system is analyzed using a stochastic zeroing control barrier function (ZCBF). A stochastic safety compensator is designed to ensure probabilistic safety, and its effectiveness is verified through numerical simulations.

eess.SY

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well understood in the tabular setting, theoretical results for function approximation remain scarce. This paper closes the gap by proposing an RL algorithm for linear CMDPs that achieves $\tilde{\mathcal{O}}(\sqrt{K})$ regret with an episode-wise zero-violation guarantee. Furthermore, our method is computationally efficient, scaling polynomially with problem-dependent parameters while remaining independent of the state space size. Our results significantly improve upon recent linear CMDP algorithms, which either violate the constraint or incur exponential computational costs.

cs.LG

Safety-Critical Control for Discrete-time Stochastic Systems with Flexible Safe Bounds using Affine and Quadratic Control Barrier Functions

This paper presents a safe controller synthesis of discrete-time stochastic systems using Control Barrier Functions (CBFs). The proposed condition allows the design of a safe controller synthesis that ensures system safety while avoiding the conservative bounds of safe probabilities. In particular, this study focuses on the design of CBFs that provide flexibility in the choice of functions to obtain tighter bounds on the safe probabilities. Numerical examples demonstrate the effectiveness of the approach.

eess.SY

Physics-Informed Representation and Learning: Control and Risk Quantification

Optimal and safety-critical control are fundamental problems for stochastic systems, and are widely considered in real-world scenarios such as robotic manipulation and autonomous driving. In this paper, we consider the problem of efficiently finding optimal and safe control for high-dimensional systems. Specifically, we propose to use dimensionality reduction techniques from a comparison theorem for stochastic differential equations together with a generalizable physics-informed neural network to estimate the optimal value function and the safety probability of the system. The proposed framework results in substantial sample efficiency improvement compared to existing methods. We further develop an autoencoder-like neural network to automatically identify the low-dimensional features of the system to enhance the ease of design for system integration. We also provide experiments and quantitative analysis to validate the efficacy of the proposed method. Source code is available at https://github.com/jacobwang925/path-integral-PINN.

eess.SY

Control Barrier Functions for Stochastic Systems and Safety-critical Control Designs

In recent years, the analysis of a control barrier function has received considerable attention because it is helpful for the safety-critical control required in many control application problems. While the extension of the analysis to a stochastic system studied by many researchers, it remains a challenging issue. In this paper, we consider sufficient conditions for reciprocal and zeroing control barrier functions ensuring safety with probability one and design a control law using the functions. Then, we propose another version of a stochastic zeroing control barrier function to evaluate a probability of a sample path staying in a safe set and confirm the convergence of a specific expectation related to the attractiveness of a safe set. We also show a way of deisgning a safety-critical control law based on our stochastic zeroing control barrier function. Finally, we confirm the validity of the proposed control design and the analysis using the control barrier functions via simple examples with their numerical simulation.

math.OC

Benchmarking Actor-Critic Deep Reinforcement Learning Algorithms for Robotics Control with Action Constraints

This study presents a benchmark for evaluating action-constrained reinforcement learning (RL) algorithms. In action-constrained RL, each action taken by the learning system must comply with certain constraints. These constraints are crucial for ensuring the feasibility and safety of actions in real-world systems. We evaluate existing algorithms and their novel variants across multiple robotics control environments, encompassing multiple action constraint types. Our evaluation provides the first in-depth perspective of the field, revealing surprising insights, including the effectiveness of a straightforward baseline approach. The benchmark problems and associated code utilized in our experiments are made available online at github.com/omron-sinicx/action-constrained-RL-benchmark for further research and development.

cs.LG

Non-Smooth Stochastic Lyapunov Functions With Weak Extension of Viscosity Solutions

This paper proposes a notion of viscosity weak supersolutions to build a bridge between stochastic Lyapunov stability theory and viscosity solution theory. Different from ordinary differential equations, stochastic differential equations can have the origins being stable despite having no smooth stochastic Lyapunov functions (SLFs). The feature naturally requires that the related Lyapunov equations are illustrated via viscosity solution theory, which deals with non-smooth solutions to partial differential equations. This paper claims that stochastic Lyapunov stability theory needs a weak extension of viscosity supersolutions, and the proposed viscosity weak supersolutions describe non-smooth SLFs ensuring a large class of the origins being noisily (asymptotically) stable and (asymptotically) stable in probability. The contribution of the non-smooth SLFs are confirmed by a few examples; especially, they ensure that all the linear-quadratic-Gaussian (LQG) controlled systems have the origins being noisily asymptotically stable for any additive noises.

math.OC