SearcharxivSearch

arXiv subjects

Sean Meyn

Publications and source records attributed to Sean Meyn.

At least 37 records · Page 2Linked to original sources

Kullback-Leibler-Quadratic Optimal Control

This paper presents approaches to mean-field control, motivated by distributed control of multi-agent systems. Control solutions are based on a convex optimization problem, whose domain is a convex set of probability mass functions (pmfs). The main contributions follow: 1. Kullback-Leibler-Quadratic (KLQ) optimal control is a special case, in which the objective function is composed of a control cost in the form of Kullback-Leibler divergence between a candidate pmf and the nominal, plus a quadratic cost on the sequence of marginals. Theory in this paper extends prior work on deterministic control systems, establishing that the optimal solution is an exponential tilting of the nominal pmf. Transform techniques are introduced to reduce complexity of the KLQ solution, motivated by the need to consider time horizons that are much longer than the inter-sampling times required for reliable control. 2. Infinite-horizon KLQ leads to a state feedback control solution with attractive properties. It can be expressed as either state feedback, in which the state is the sequence of marginal pmfs, or an open loop solution is obtained that is more easily computed. 3. Numerical experiments are surveyed in an application of distributed control of residential loads to provide grid services, similar to utility-scale battery storage. The results show that KLQ optimal control enables the aggregate power consumption of a collection of flexible loads to track a time-varying reference signal, while simultaneously ensuring each individual load satisfies its own quality of service constraints. Keywords: Mean field games, distributed control, Markov decision processes, Demand Dispatch.

math.OC

High-Impedance Non-Linear Fault Detection via Eigenvalue Analysis with low PMU Sampling Rates

This technique holds several advantages over contemporary techniques: It utilizes technology that is already deployed in the field, it offers a significant degree of generality, and so far it has displayed a very high-level of sensitivity without sacrificing accuracy. Validation is performed in the form of simulations based in the IEEE 13 Node System and non-linear fault models. Test results are encouraging, indicating potential for real-life applications.

eess.SY

Uncertainty Error Modeling for Non-Linear State Estimation With Unsynchronized SCADA and $μ$PMU Measurements

Distribution systems of the future smart grid require enhancements to the reliability of distribution system state estimation (DSSE) in the face of low measurement redundancy, unsynchronized measurements, and dynamic load profiles. Micro phasor measurement units ($μ$PMUs) facilitate co-synchronized measurements with high granularity, albeit at an often prohibitively expensive installation cost. Supervisory control and data acquisition (SCADA) measurements can supplement $μ$PMU data, although they are received at a slower sampling rate. Further complicating matters is the uncertainty associated with load dynamics and unsynchronized measurements-not only are the SCADA and $μ$PMU measurements not synchronized with each other, but the SCADA measurements themselves are received at different time intervals with respect to one another. This paper proposes a non-linear state estimation framework which models dynamic load uncertainty error by updating the variances of the unsynchronized measurements, leading to a time-varying system of weights in the weighted least squares state estimator. Case studies are performed on the 33-Bus Distribution System in MATPOWER, using Ornstein-Uhlenbeck stochastic processes to simulate dynamic load conditions.

eess.SY

High Impedance Fault Detection Through Quasi-Static State Estimation: A Parameter Error Modeling Approach

This paper presents a model for detecting high-impedance faults (HIFs) using parameter error modeling and a two-step per-phase weighted least squares state estimation (SE) process. The proposed scheme leverages the use of phasor measurement units and synthetic measurements to identify per-phase power flow and injection measurements which indicate a parameter error through $χ^2$ Hypothesis Testing applied to the composed measurement error (CME). Although current and voltage waveforms are commonly analyzed for high-impedance fault detection, wide-area power flow and injection measurements, which are already inherent to the SE process, also show promise for real-world high-impedance fault detection applications. The error distributions after detection share the measurement function error spread observed in proven parameter error diagnostics and can be applied to HIF identification. Further, this error spread related to the HIF will be clearly discerned from measurement error. Case studies are performed on the 33-Bus Distribution System in Simulink.

eess.SY

Sufficient Exploration for Convex Q-learning

In recent years there has been a collective research effort to find new formulations of reinforcement learning that are simultaneously more efficient and more amenable to analysis. This paper concerns one approach that builds on the linear programming (LP) formulation of optimal control of Manne. A primal version is called logistic Q-learning, and a dual variant is convex Q-learning. This paper focuses on the latter, while building bridges with the former. The main contributions follow: (i) The dual of convex Q-learning is not precisely Manne's LP or a version of logistic Q-learning, but has similar structure that reveals the need for regularization to avoid over-fitting. (ii) A sufficient condition is obtained for a bounded solution to the Q-learning LP. (iii) Simulation studies reveal numerical challenges when addressing sampled-data systems based on a continuous time model. The challenge is addressed using state-dependent sampling. The theory is illustrated with applications to examples from OpenAI gym. It is shown that convex Q-learning is successful in cases where standard Q-learning diverges, such as the LQR problem.

math.OC

Model-Free Characterizations of the Hamilton-Jacobi-Bellman Equation and Convex Q-Learning in Continuous Time

Convex Q-learning is a recent approach to reinforcement learning, motivated by the possibility of a firmer theory for convergence, and the possibility of making use of greater a priori knowledge regarding policy or value function structure. This paper explores algorithm design in the continuous time domain, with finite-horizon optimal control objective. The main contributions are (i) Algorithm design is based on a new Q-ODE, which defines the model-free characterization of the Hamilton-Jacobi-Bellman equation. (ii) The Q-ODE motivates a new formulation of Convex Q-learning that avoids the approximations appearing in prior work. The Bellman error used in the algorithm is defined by filtered measurements, which is beneficial in the presence of measurement noise. (iii) A characterization of boundedness of the constraint region is obtained through a non-trivial extension of recent results from the discrete time setting. (iv) The theory is illustrated in application to resource allocation for distributed energy resources, for which the theory is ideally suited.

math.OC

The Conditional Poincaré Inequality for Filter Stability

This paper is concerned with the problem of nonlinear filter stability of ergodic Markov processes. The main contribution is the conditional Poincaré inequality (PI), which is shown to yield filter stability. The proof is based upon a recently discovered duality which is used to transform the nonlinear filtering problem into a stochastic optimal control problem for a backward stochastic differential equation (BSDE). Based on these dual formalisms, a comparison is drawn between the stochastic stability of a Markov process and the filter stability. The latter relies on the conditional PI described in this paper, whereas the former relies on the standard form of PI.

math.PR

Reliable Power Grid: Long Overdue Alternatives to Surge Pricing

This paper takes a fresh look at the economic theory that is motivation for pricing models, such as critical peak pricing (CPP), or surge pricing, and the demand response models advocated by policy makers and in the power systems literature. The economic analysis in this paper begins with two premises: 1) a meaningful analysis requires a realistic model of stakeholder/consumer rationality, and 2) the relationship between electric power and the ultimate use of electricity are only loosely related in many cases. The most obvious examples are refrigerators and hot water heaters that consume power intermittently to maintain constant temperature. Based on a realistic model of user preferences, it is shown that the use of CPP and related pricing schemes will eventually destabilize the grid with increased participation. Analysis of this model also leads to a competitive equilibrium, along with a characterization of the dynamic prices in this equilibrium. However, we argue that these prices will not lead to a robust control solution that is acceptable to either grid operators or consumers. These findings are presented in this paper to alert policy makers of the risk of implementing real time prices to control our energy grid. Competitive equilibrium theory can only provide a caricature of a real-world market, since complexities such as sunk cost and risk are not included. The paper explains why these approximations are especially significant in the power industry. It concludes with policy recommendations to bolster the reliability of the power grid, with a focus on planning across different timescales and alternate approaches to leveraging demand-side flexibility in the grid.

math.OC

Aggregate capacity of TCLs with cycling constraints

Thermostatically Controlled Loads (TCLs) such as air conditioners and water heaters typically maintain their temperature within a preset range using on/off actuation. These types of loads are inherently flexible: many different power consumption trajectories exist that can keep the temperature within range. Decades of research has shown that flexible loads can provide valuable grid services. Quantifying the power and energy capacities of a collection of TCLs is a well-studied problem. However, most works focus on temperature constraints. In this work, we present a characterization of the capacity of a collection of TCLs that considers not only temperature, but also cycling and energy constraints. The characterization leads to a set of convex constraints. A grid operator can use this characterization to compute a feasible power consumption trajectory for an ensemble of TCLs that comes closest to what the operator needs to maintain demand-supply balance. Unlike prior attempts at capacity characterizations incorporating cycling constraints, our results are independent of the algorithm used to coordinate the TCLs.

eess.SY

Accelerating Optimization and Reinforcement Learning with Quasi-Stochastic Approximation

The ODE method has been a workhorse for algorithm design and analysis since the introduction of the stochastic approximation. It is now understood that convergence theory amounts to establishing robustness of Euler approximations for ODEs, while theory of rates of convergence requires finer analysis. This paper sets out to extend this theory to quasi-stochastic approximation, based on algorithms in which the "noise" is based on deterministic signals. The main results are obtained under minimal assumptions: the usual Lipschitz conditions for ODE vector fields, and it is assumed that there is a well defined linearization near the optimal parameter $θ^*$, with Hurwitz linearization matrix $A^*$. The main contributions are summarized as follows: (i) If the algorithm gain is $a_t=g/(1+t)^ρ$ with $g>0$ and $ρ\in(0,1)$, then the rate of convergence of the algorithm is $1/t^ρ$. There is also a well defined "finite-$t$" approximation: \[ a_t^{-1}\{Θ_t-θ^*\}=\bar{Y}+Ξ^{\mathrm{I}}_t+o(1) \] where $\bar{Y}\in\mathbb{R}^d$ is a vector identified in the paper, and $\{Ξ^{\mathrm{I}}_t\}$ is bounded with zero temporal mean. (ii) With gain $a_t = g/(1+t)$ the results are not as sharp: the rate of convergence $1/t$ holds only if $I + g A^*$ is Hurwitz. (iii) Based on the Ruppert-Polyak averaging of stochastic approximation, one would expect that a convergence rate of $1/t$ can be obtained by averaging: \[ Θ^{\text{RP}}_T=\frac{1}{T}\int_{0}^T Θ_t\,dt \] where the estimates $\{Θ_t\}$ are obtained using the gain in (i). The preceding sharp bounds imply that averaging results in $1/t$ convergence rate if and only if $\bar{Y}=\sf 0$. This condition holds if the noise is additive, but appears to fail in general. (iv) The theory is illustrated with applications to gradient-free optimization and policy gradient algorithms for reinforcement learning.

math.OC

Explicit Mean-Square Error Bounds for Monte-Carlo and Linear Stochastic Approximation

This paper concerns error bounds for recursive equations subject to Markovian disturbances. Motivating examples abound within the fields of Markov chain Monte Carlo (MCMC) and Reinforcement Learning (RL), and many of these algorithms can be interpreted as special cases of stochastic approximation (SA). It is argued that it is not possible in general to obtain a Hoeffding bound on the error sequence, even when the underlying Markov chain is reversible and geometrically ergodic, such as the M/M/1 queue. This is motivation for the focus on mean square error bounds for parameter estimates. It is shown that mean square error achieves the optimal rate of $O(1/n)$, subject to conditions on the step-size sequence. Moreover, the exact constants in the rate are obtained, which is of great value in algorithm design.

math.PR

Demand Dispatch with Heterogeneous Intelligent Loads

A distributed control architecture is presented that is intended to make a collection of heterogeneous loads appear to the grid operator as a nearly perfect battery. Local control is based on randomized decision rules advocated in prior research, and extended in this paper to any load with a discrete number of power states. Additional linear filtering at the load ensures that the input-output dynamics of the aggregate has a nearly flat input-output response: the behavior of an ideal, multi-GW battery system.

eess.SY

Model-Free Primal-Dual Methods for Network Optimization with Application to Real-Time Optimal Power Flow

This paper examines the problem of real-time optimization of networked systems and develops online algorithms that steer the system towards the optimal trajectory without explicit knowledge of the system model. The problem is modeled as a dynamic optimization problem with time-varying performance objectives and engineering constraints. The design of the algorithms leverages the online zero-order primal-dual projected-gradient method. In particular, the primal step that involves the gradient of the objective function (and hence requires networked systems model) is replaced by its zero-order approximation with two function evaluations using a deterministic perturbation signal. The evaluations are performed using the measurements of the system output, hence giving rise to a feedback interconnection, with the optimization algorithm serving as a feedback controller. The paper provides some insights on the stability and tracking properties of this interconnection. Finally, the paper applies this methodology to a real-time optimal power flow problem in power systems, and shows its efficacy on the IEEE 37-node distribution test feeder for reference power tracking and voltage regulation.

math.OC

State Space Collapse in Resource Allocation for Demand Dispatch

Demand dispatch is the science of extracting virtual energy storage through the automatic control of deferrable loads to provide balancing or regulation services to the grid, while maintaining consumer-end quality of service (QoS). The control of a large collection of heterogeneous loads is in part a resource allocation problem, since different classes of loads are valuable for different services. The goal of this paper is to unveil the structure of the optimal solution to the resource allocation problem and to investigate short term market implications. It is found that the marginal cost for each load class evolves on a two-dimensional subspace, spanned by an optimal costate process and its derivative. The resource allocation problem is recast to construct a dynamic competitive equilibrium model, in which the consumer utility is the negative of the cost of deviation from ideal QoS. It is found that a competitive equilibrium exists, with the equilibrium price equal to the negative of an optimal costate process. Moreover, the equilibrium price is different from what would be obtained based on the standard assumption that the consumer's utility is a function of power consumption.

eess.SY

Geometric Ergodicity in a Weighted Sobolev Space

For a discrete-time Markov chain $\{X(t)\}$ evolving on $\Re^\ell$ with transition kernel $P$, natural, general conditions are developed under which the following are established: 1. The transition kernel $P$ has a purely discrete spectrum, when viewed as a linear operator on a weighted Sobolev space $L_\infty^{v,1}$ of functions with norm, $$ \|f\|_{v,1} = \sup_{x \in \Re^\ell} \frac{1}{v(x)} \max \{|f(x)|, |\partial_1 f(x)|,\ldots,|\partial_\ell f(x)|\}, $$ where $v\colon \Re^\ell \to [1,\infty)$ is a Lyapunov function and $\partial_i:=\partial/\partial x_i$. 2. The Markov chain is geometrically ergodic in $L_\infty^{v,1}$: There is a unique invariant probability measure $π$ and constants $B<\infty$ and $δ>0$ such that, for each $f\in L_\infty^{v,1}$, any initial condition $X(0)=x$, and all $t\geq 0$: $$\Big| \text{E}_x[f(X(t))] - π(f)\Big| \le Be^{-δt}v(x),\quad \|\nabla \text{E}_x[f(X(t))] \|_2 \le Be^{-δt} v(x), $$ where $π(f)=\int fdπ$. 3. For any function $f\in L_\infty^{v,1}$ there is a function $h\in L_\infty^{v,1}$ solving Poisson's equation: \[ h-Ph = f-π(f). \] Part of the analysis is based on an operator-theoretic treatment of the sensitivity process that appears in the theory of Lyapunov exponents.

math.PR

Diffusion approximations and control variates for MCMC

A new methodology is presented for the construction of control variates to reduce the variance of additive functionals of Markov Chain Monte Carlo (MCMC) samplers. Our control variates are definedthrough the minimization of the asymptotic variance of the Langevin diffusion over a family of functions, which can be seen as a quadratic risk minimization procedure. The use of these control variates is theoretically justified. We show that the asymptotic variances of some well-known MCMC algorithms, including the Random Walk Metropolis and the (Metropolis) Unadjusted/Adjusted Langevin Algorithm, are close to the asymptotic variance of the Langevin diffusion. Several examples of Bayesian inference problems demonstrate that the corresponding reduction in the variance is significant.

stat.ME

Optimal Rate of Convergence for Quasi-Stochastic Approximation

The Robbins-Monro stochastic approximation algorithm is a foundation of many algorithmic frameworks for reinforcement learning (RL), and often an efficient approach to solving (or approximating the solution to) complex optimal control problems. However, in many cases practitioners are unable to apply these techniques because of an inherent high variance. This paper aims to provide a general foundation for "quasi-stochastic approximation," in which all of the processes under consideration are deterministic, much like quasi-Monte-Carlo for variance reduction in simulation. The variance reduction can be substantial, subject to tuning of pertinent parameters in the algorithm. This paper introduces a new coupling argument to establish optimal rate of convergence provided the gain is sufficiently large. These results are established for linear models, and tested also in non-ideal settings. A major application of these general results is a new class of RL algorithms for deterministic state space models. In this setting, the main contribution is a class of algorithms for approximating the value function for a given policy, using a different policy designed to introduce exploration.

math.OC

Optimal Matrix Momentum Stochastic Approximation and Applications to Q-learning

Acceleration is an increasingly common theme in the stochastic optimization literature. The two most common examples are Nesterov's method, and Polyak's momentum technique. In this paper two new algorithms are introduced for root finding problems: 1) PolSA is a root finding algorithm with specially designed matrix momentum, and 2) NeSA can be regarded as a variant of Nesterov's algorithm, or a simplification of PolSA. The PolSA algorithm is new even in the context of optimization (when cast as a root finding problem). The research surveyed in this paper is motivated by applications to reinforcement learning. It is well known that most variants of TD- and Q-learning may be cast as SA (stochastic approximation) algorithms, and the tools from general SA theory can be used to investigate convergence and bounds on convergence rate. In particular, the asymptotic variance is a common metric of performance for SA algorithms, and is also one among many metrics used in assessing the performance of stochastic optimization algorithms. There are two well known SA techniques that are known to have optimal asymptotic variance: the Ruppert-Polyak averaging technique, and stochastic Newton-Raphson (SNR). The former algorithm can have extremely bad transient performance, and the latter can be computationally expensive. It is demonstrated here that parameter estimates from the new PolSA algorithm couple with those of the ideal (but more complex) SNR algorithm. The new algorithm is thus a third approach to obtain optimal asymptotic covariance. These strong results require assumptions on the model. A linearized model is considered, and the noise is assumed to be a martingale difference sequence. Numerical results are obtained in a non-linear setting that is the motivation for this work: In PolSA implementations of Q-learning it is observed that coupling occurs with SNR in this non-ideal setting.

math.OC