SearcharxivSearch

arXiv subjects

Do Wan Kim

Publications and source records attributed to Do Wan Kim.

9 recordsLinked to original sources

Relaxed Conditions for Parameterized Linear Matrix Inequality in the Form of Nested Fuzzy Summations

The aim of this study is to investigate less conservative conditions for parameterized linear matrix inequalities (PLMIs) that are formulated as nested fuzzy summations. Such PLMIs are commonly encountered in stability analysis and control design problems for Takagi-Sugeno (T-S) fuzzy systems. Utilizing the weighted inequality of arithmetic and geometric means (AM-GM inequality), we develop new, less conservative linear matrix inequalities for the PLMIs. This methodology enables us to efficiently handle the product of membership functions that have intersecting indices. Through empirical case studies, we demonstrate that our proposed conditions produce less conservative results compared to existing approaches in the literature.

math.OC

On the Local Quadratic Stability of T-S Fuzzy Systems in the Vicinity of the Origin

The main goal of this paper is to introduce new local stability conditions for continuous-time Takagi-Sugeno (T-S) fuzzy systems. These stability conditions are based on linear matrix inequalities (LMIs) in combination with quadratic Lyapunov functions. Moreover, they integrate information on the membership functions at the origin and effectively leverage the linear structure of the underlying nonlinear system in the vicinity of the origin. As a result, the proposed conditions are proved to be less conservative compared to existing methods using fuzzy Lyapunov functions in the literature. Moreover, we establish that the proposed methods offer necessary and sufficient conditions for the local exponential stability of T-S fuzzy systems. The paper also includes discussions on the inherent limitations associated with fuzzy Lyapunov approaches. To demonstrate the theoretical results, we provide comprehensive examples that elucidate the core concepts and validate the efficacy of the proposed conditions.

eess.SY

Control Theoretic Analysis of Temporal Difference Learning

The goal of this manuscript is to conduct a controltheoretic analysis of Temporal Difference (TD) learning algorithms. TD-learning serves as a cornerstone in the realm of reinforcement learning, offering a methodology for approximating the value function associated with a given policy in a Markov Decision Process. Despite several existing works that have contributed to the theoretical understanding of TD-learning, it is only in recent years that researchers have been able to establish concrete guarantees on its statistical efficiency. In this paper, we introduce a finite-time, control-theoretic framework for analyzing TD-learning, leveraging established concepts from the field of linear systems control. Consequently, this paper provides additional insights into the mechanics of TD learning and the broader landscape of reinforcement learning, all while employing straightforward analytical tools derived from control theory.

cs.AI

Continuous-Time Distributed Dynamic Programming for Networked Multi-Agent Markov Decision Processes

The main goal of this paper is to investigate continuous-time distributed dynamic programming (DP) algorithms for networked multi-agent Markov decision problems (MAMDPs). In our study, we adopt a distributed multi-agent framework where individual agents have access only to their own rewards, lacking insights into the rewards of other agents. Moreover, each agent has the ability to share its parameters with neighboring agents through a communication network, represented by a graph. We first introduce a novel distributed DP, inspired by the distributed optimization method of Wang and Elia. Next, a new distributed DP is introduced through a decoupling process. The convergence of the DP algorithms is proved through systems and control perspectives. The study in this paper sets the stage for new distributed temporal different learning algorithms.

eess.SY

Relaxed Conditions for Parameterized Linear Matrix Inequality in the Form of Double Sum

The aim of this study is to investigate less conservative conditions for a parameterized linear matrix inequality (PLMI) expressed in the form of a double convex sum. This type of PLMI frequently appears in T-S fuzzy control system analysis and design problems. In this letter, we derive new, less conservative linear matrix inequalities (LMIs) for the PLMI by employing the proposed sum relaxation method based on Young's inequality. The derived LMIs are proven to be less conservative than the existing conditions related to this topic in the literature. The proposed technique is applicable to various stability analysis and control design problems for T-S fuzzy systems, which are formulated as solving the PLMIs in the form of a double convex sum. Furthermore, examples is provided to illustrate the reduced conservatism of the derived LMIs.

eess.SY

Finite-Time Accuracy of Temporal-Difference Learning Under Schur-Stable Recursions

Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision process under a fixed policy. While a substantial body of work has established its convergence and stability properties, more recent efforts have focused on its statistical efficiency through finite-time error bounds. In this paper, we advance this line of research by developing a new finite-time error analysis for tabular TD learning that directly exploits a discrete-time stochastic linear system representation and leverages Schur stability of the associated matrices. Beyond the specific bounds obtained, the proposed framework provides a reusable template for analyzing TD learning and related RL algorithms, and it offers control-theoretic insights that may guide future developments in finite-sample RL theory.

cs.LG

Data-Driven Control Design with LMIs and Dynamic Programming

The goal of this paper is to develop data-driven control design and evaluation strategies based on linear matrix inequalities (LMIs) and dynamic programming. We consider deterministic discrete-time LTI systems, where the system model is unknown. We propose efficient data collection schemes from the state-input trajectories together with data-driven LMIs to design state-feedback controllers for stabilization and linear quadratic regulation (LQR) problem. In addition, we investigate theoretically guaranteed exploration schemes to acquire valid data from the trajectories under different scenarios. In particular, we prove that as more and more data is accumulated, the collected data becomes valid for the proposed algorithms with higher probability. Finally, data-driven dynamic programming algorithms with convergence guarantees are then discussed.

math.OC

Multi-Objective LQG Design with Primal-Dual Method

The goal of this paper is to study a multi-objective linear quadratic Gaussian (LQG) control problem. In particular, we consider an optimal control problem minimizing a quadratic cost over a finite time horizon for linear stochastic systems subject to control energy constraints. To solve the problem, we suggest an efficient bisection line search algorithm which is computationally efficient compared to other approaches such as the semidefinite programming. The main idea is to use the Lagrangian function and Karush-Kuhn-Tucker (KKT) optimality conditions to solve the constrained optimization problem. The Lagrange multiplier is searched using the bisection line search. Numerical examples are given to demonstrate the effectiveness of the proposed methods.

math.OC

Quantitative estimates for stress concentration of the Stokes flow between adjacent circular cylinders

When two inclusions with high contrast material properties are located close to each other in a homogeneous medium, stress may become arbitrarily large in the narrow region between them. In this paper, we investigate such stress concentration in the two-dimensional Stokes flow when inclusions are the two-dimensional cross sections of circular cylinders of the same radii and the background velocity field is linear. We construct two vector-valued functions which completely capture the singular behavior of the stress and derive an asymptotic representation formula for the stress in terms of these functions as the distance between the two cylinders tends to zero. We then show, using the representation formula, that the stress always blows up by proving that either the pressure or the shear stress component of the stress tensor blows up. The blow-up rate is shown to be $δ^{-1/2}$, where $δ$ is the distance between the two cylinders. To our best knowledge, this work is the first to rigorously derive the asymptotic solution in the narrow region for the Stokes flow.

math.AP