SearcharxivSearch

arXiv subjects

Yohei Hosoe

Publications and source records attributed to Yohei Hosoe.

10 recordsLinked to original sources

Linear Stochastic Systems with i.i.d. uncertainties: Exact Covariance Characterization, Stability Analysis and State-feedback Design

This paper studies linear discrete-time systems affected by independent and identically distributed (i.i.d.) multiplicative uncertainties and additive noise. It establishes the main links between covariance recursions, the spectral properties of associated Kronecker-based matrices, and mean-square stability, and exploits these links to derive tractable conditions for controller synthesis. We first derive a deterministic covariance recursion within the tube-based Stochastic Model Predictive Control (SMPC) framework using a Kronecker product based matrix augmentation. For linear stochastic systems with multiplicative uncertainty and without additive noise, we show that the full-space matrix representation arising from the covariance recursion has the same spectral radius as its symmetric-space counterpart. Combined with the existing symmetric-space characterization, this establishes that Schur stability of the full-space augmented matrix is equivalent to mean-square stability. For state-feedback design, we propose new sufficient Linear Matrix Inequality (LMI) conditions that are numerically more tractable owing to their reduced size compared with the conventional necessary and sufficient conditions. Numerical tests illustrate the usefulness of the covariance characterization for recursively estimating the covariance without relying on sampling-based methods. We also assess the computational burden of the proposed LMI conditions and their conservatism relative to the necessary and sufficient ones.

eess.SY

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well understood in the tabular setting, theoretical results for function approximation remain scarce. This paper closes the gap by proposing an RL algorithm for linear CMDPs that achieves $\tilde{\mathcal{O}}(\sqrt{K})$ regret with an episode-wise zero-violation guarantee. Furthermore, our method is computationally efficient, scaling polynomially with problem-dependent parameters while remaining independent of the state space size. Our results significantly improve upon recent linear CMDP algorithms, which either violate the constraint or incur exponential computational costs.

cs.LG

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form

Designing a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust constrained MDP (RCMDP), where an optimal policy minimizes cumulative cost while satisfying constraints in the worst-case scenario across a set of environments. We first prove that the conventional policy gradient approach to the Lagrangian max-min formulation can become trapped in suboptimal solutions. This occurs when its inner minimization encounters a sum of conflicting gradients from the objective and constraint functions. To address this, we leverage the epigraph form of the RCMDP problem, which resolves the conflict by selecting a single gradient from either the objective or the constraints. Building on the epigraph form, we propose a bisection search algorithm with a policy gradient subroutine and prove that it identifies an $\varepsilon$-optimal policy in an RCMDP with $\tilde{\mathcal{O}}(\varepsilon^{-4})$ robust policy evaluations.

cs.LG

Benchmarking Actor-Critic Deep Reinforcement Learning Algorithms for Robotics Control with Action Constraints

This study presents a benchmark for evaluating action-constrained reinforcement learning (RL) algorithms. In action-constrained RL, each action taken by the learning system must comply with certain constraints. These constraints are crucial for ensuring the feasibility and safety of actions in real-world systems. We evaluate existing algorithms and their novel variants across multiple robotics control environments, encompassing multiple action constraint types. Our evaluation provides the first in-depth perspective of the field, revealing surprising insights, including the effectiveness of a straightforward baseline approach. The benchmark problems and associated code utilized in our experiments are made available online at github.com/omron-sinicx/action-constrained-RL-benchmark for further research and development.

cs.LG

H2 Performance Analysis and Synthesis for Discrete-Time Linear Systems with Dynamics Determined by an i.i.d. Process

This paper is concerned with H2 control of discrete-time linear systems with dynamics determined by an independent and identically distributed (i.i.d.) process. A definition of H2 norm is first discussed for the class of systems. Then, a linear matrix inequality (LMI) condition is derived for the associated performance analysis, which is tractable in the sense of numerical computation. The results about analysis are also extended toward state-feedback controller synthesis.

eess.SY

Stochastic Aperiodic Control of Networked Systems with i.i.d. Time-Varying Communication Delays

This paper studies stochastic aperiodic stabilization of a networked control system (NCS) consisting of a continuous-time plant and a discrete-time controller. The plant and the controller are assumed to be connected by communication channels with i.i.d. time-varying delays. The delays are theoretically not required to be bounded even when the plant is unstable in the deterministic sense. In our NCS, the sampling interval is supposed to be determined directly by such communication delays. A necessary and sufficient inequality condition is presented for designing a state-feedback controller stabilizing the NCS at sampling points in a stochastic sense. The results are also illustrated numerically.

eess.SY

Contraction Analysis of Discrete-time Stochastic Systems

In this paper, we develop a novel contraction framework for stability analysis of discrete-time nonlinear systems with parameters following stochastic processes. For general stochastic processes, we first provide a sufficient condition for uniform incremental exponential stability (UIES) in the first moment with respect to a Riemannian metric. Then, focusing on the Euclidean distance, we present a necessary and sufficient condition for UIES in the second moment. By virtue of studying general stochastic processes, we can readily derive UIES conditions for special classes of processes, e.g., i.i.d. processes and Markov processes, which is demonstrated as selected applications of our results.

eess.SY

On Second-Moment Stability of Discrete-Time Linear Systems with General Stochastic Dynamics

This paper provides a new unified framework for second-moment stability of discrete-time linear systems with stochastic dynamics. Relations of notions of second-moment stability are studied for the systems with general stochastic dynamics, and associated Lyapunov inequalities are derived. Any type of stochastic process can be dealt with as a special case in our framework for determining system dynamics, and our results together with assumptions (i.e., restrictions) on the process immediately lead us to stability conditions for the corresponding special stochastic systems. As a demonstration of usefulness of such a framework, three selected applications are also provided.

eess.SY

Distribution Modeling and Stabilization Control for Discrete-Time Linear Random Dynamical Systems Using Ensemble Kalman Filter

This paper studies an output feedback stabilization control framework for discrete-time linear systems with stochastic dynamics determined by an independent and identically distributed (i.i.d.) process. The controller is constructed with an ensemble Kalman filter (EnKF) and a feedback gain designed with our earlier result about state feedback control. The EnKF is also used for modeling the distribution behind the system, which is required in the feedback gain synthesis. The effectiveness of our control framework is demonstrated with numerical experiments. This study will become the first step toward the realization of learning type control using our stochastic systems control theory.

eess.SY

Equivalent Stability Notions, Lyapunov Inequality, and Its Application in Discrete-Time Linear Systems with Stochastic Dynamics Determined by an i.i.d. Process

This paper is concerned with stability analysis and synthesis for discrete-time linear systems with stochastic dynamics. Equivalence is first proved for three stability notions under some key assumptions on the randomness behind the systems. In particular, we use the assumption that the stochastic process determining the system dynamics is independent and identically distributed (i.i.d.) with respect to the discrete time. Then, a Lyapunov inequality condition is derived for stability in a necessary and sufficient sense. Although our Lyapunov inequality will involve decision variables contained in the expectation operation, an idea is provided to solve it as a standard linear matrix inequality; the idea also plays an important role in state feedback synthesis based on the Lyapunov inequality. Motivating numerical examples are further discussed as an application of our approach.

eess.SY