SearcharxivSearch

arXiv subjects

Andrew Lamperski

Publications and source records attributed to Andrew Lamperski.

At least 19 recordsLinked to original sources

Non-Asymptotic Analysis of Classical Spectrum Estimators for $L$-mixing Time-series Data with Estimated Means

Spectral estimation is an important tool in time series analysis, with applications including economics, astronomy, and climatology. The asymptotic theory for non-parametric estimation is well-known but the development of non-asymptotic theory is still ongoing. Our recent work obtained the first non-asymptotic error bounds on the Bartlett and Welch methods with restrictive assumptions. In this work, we derive non-asymptotic error bounds for both Bartlett and Welch estimators for $L$-mixing time-series data with unknown means, and the results cover the special case with known zero means. The class of $L$-mixing processes contains common models in time series analysis, including autoregressive processes and measurements of geometrically ergodic Markov chains. Our new error bounds are of $O(\frac{1}{\sqrt{k}})$, where $k$ is the number of data segments used in the algorithm. Such bounds are the tightest among the existing work on non-asymptotic analysis of classical spectrum estimators with or without zero-mean assumptions.

math.ST

Bounds on Spectral Gaps for Non-Reversible Markov Chains with Applications to Temporal Difference Learning

This work is motivated by the analysis of temporal difference algorithms, where stability can be guaranteed by bounding the eigenvalues of an associated matrix derived from a, typically non-reversible, Markov kernel. We generalize the existing sufficient conditions for stability and show that the associated eigenvalues can be bounded in terms of the Dirichlet spectral gap of the Markov kernel. We derive a collection of methods for showing that non-reversible Markov chains have positive spectral gaps. We show that if a Markov chain has positive absolute spectral gap, then it has a positive Dirichlet spectral gap. In the case of discrete-time linear Gaussian systems, we give explicit bounds for both Dirichlet and absolute spectral gaps. Additionally, we present an example of a Markov chain which is $V$-uniformly geometrically ergodic but has zero Dirichlet spectral gap.

math.PR

Active Sensing Subserves Task-Level Control

Active sensing is traditionally defined as the expenditure of energy, typically in the form of movement, for obtaining information. Here, we propose that the combination of reliance on adaptive sensors, the linkage between movement and sensing, and task-level control inevitably gives rise to the emergence of active sensing movements. In this way, active sensing is not driven by sensory goals, such as minimizing uncertainty about the state, but rather is necessary for task-level control. This hypothesis, that active sensing subserves control, is supported by both empirical data from organisms and mathematical theory. Interestingly, active sensing behaviors often occur in discrete epochs, interspersed with goal-oriented behavior. This suggests that animals switch between two behavioral modes with distinct control policies, an `explore' mode in which animals produce dynamic movements to shape sensory feedback, and an `exploit' mode in which animals produce slower compensatory movements that are directly related to achieving task goals. This strategy for feedback control that relies on adaptive sensors, active sensing, and mode switching is not commonly used in engineered systems despite being ubiquitous in biology. Engineered systems comprising state-of-the-art sensors, actuators, and mechanical designs can outperform animals with respect to ``cost functions'' such as maximum force generation, precision, and speed. Nevertheless, animals routinely achieve robust, graceful behaviors that are currently unmatched by engineered systems, suggesting that current control systems are insufficient. These insights, expressed in the language of control theory, may be critical for improving robotic sensing and control.

q-bio.NC

Non-Asymptotic Error Bounds for Causally Conditioned Directed Information Rates of Gaussian Sequences

Directed information and its causally conditioned variations are often used to measure causal influences between random processes. In practice, these quantities must be measured from data. Non-asymptotic error bounds for these estimates are known for sequences over finite alphabets, but less is known for real-valued data. This paper examines the case in which the data are sequences of Gaussian vectors. We provide an explicit formula for causally conditioned directed information rate based on optimal prediction and define an estimator based on this formula. We show that our estimator gives an error of order $O\left(N^{-1/2}\log(N)\right)$ with high probability, where $N$ is the total sample size.

cs.IT

A Neural Network Algorithm for KL Divergence Estimation with Quantitative Error Bounds

Estimating the Kullback-Leibler (KL) divergence between random variables is a fundamental problem in statistical analysis. For continuous random variables, traditional information-theoretic estimators scale poorly with dimension and/or sample size. To mitigate this challenge, a variety of methods have been proposed to estimate KL divergences and related quantities, such as mutual information, using neural networks. The existing theoretical analyses show that neural network parameters achieving low error exist. However, since they rely on non-constructive neural network approximation theorems, they do not guarantee that the existing algorithms actually achieve low error. In this paper, we propose a KL divergence estimation algorithm using a shallow neural network with randomized hidden weights and biases (i.e. a random feature method). We show that with high probability, the algorithm achieves a KL divergence estimation error of $O(m^{-1/2}+T^{-1/3})$, where $m$ is the number of neurons and $T$ is both the number of steps of the algorithm and the number of samples.

cs.LG

Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential

Stochastic gradient descent (SGD) is the main algorithm behind a large body of work in machine learning. In many cases, constraints are enforced via projections, leading to projected stochastic gradient algorithms. In recent years, a large body of work has examined the convergence properties of projected SGD for non-convex losses in asymptotic and non-asymptotic settings. Strong quantitative guarantees are available for convergence measured via Moreau envelopes. However, these results cannot be compared directly with work on unconstrained SGD, since the Moreau envelope construction changes the gradient. Other common measures based on gradient mappings have the limitation that convergence can only be guaranteed if variance reduction methods, such as mini-batching, are employed. This paper presents an analysis of projected SGD for non-convex losses over compact convex sets. Convergence is measured via the distance of the gradient to the Goldstein subdifferential generated by the constraints. Our proposed convergence criterion directly reduces to commonly used criteria in the unconstrained case, and we obtain convergence without requiring variance reduction. We obtain results for data that are independent, identically distributed (IID) or satisfy mixing conditions ($L$-mixing). In these cases, we derive asymptotic convergence and $O(N^{-1/3})$ non-asymptotic bounds in expectation, where $N$ is the number of steps. In the case of IID sub-Gaussian data, we obtain almost-sure asymptotic convergence and high-probability non-asymptotic $O(N^{-1/5})$ bounds. In particular, these are the first non-asymptotic high-probability bounds for projected SGD with non-convex losses.

math.OC

Non-Asymptotic Analysis of Classical Spectrum Estimators with $L$-mixing Time-series Data

Spectral estimation is a fundamental problem for time series analysis, which is widely applied in economics, speech analysis, seismology, and control systems. The asymptotic convergence theory for classical, non-parametric estimators, is well-understood, but the non-asymptotic theory is still rather limited. Our recent work gave the first non-asymptotic error bounds on the well-known Bartlett and Welch methods, but under restrictive assumptions. In this paper, we derive non-asymptotic error bounds for a class of non-parametric spectral estimators, which includes the classical Bartlett and Welch methods, under the assumption that the data is an $L$-mixing stochastic process. A broad range of processes arising in time-series analysis, such as autoregressive processes and measurements of geometrically ergodic Markov chains, can be shown to be $L$-mixing. In particular, $L$-mixing processes can model a variety of nonlinear phenomena which do not satisfy the assumptions of our prior work. Our new error bounds for $L$-mixing processes match the error bounds in the restrictive settings from prior work up to logarithmic factors.

math.ST

Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications

Neural networks are widely used to approximate unknown functions in control. A common neural network architecture uses a single hidden layer (i.e. a shallow network), in which the input parameters are fixed in advance and only the output parameters are trained. The typical formal analysis asserts that if output parameters exist to approximate the unknown function with sufficient accuracy, then desired control performance can be achieved. A long-standing theoretical gap was that no conditions existed to guarantee that, for the fixed input parameters, required accuracy could be obtained by training the output parameters. Our recent work has partially closed this gap by demonstrating that if input parameters are chosen randomly, then for any sufficiently smooth function, with high-probability there are output parameters resulting in $O((1/m)^{1/2})$ approximation errors, where $m$ is the number of neurons. However, some applications, notably continuous-time value function approximation, require that the network approximates the both the unknown function and its gradient with sufficient accuracy. In this paper, we show that randomly generated input parameters and trained output parameters result in gradient errors of $O((\log(m)/m)^{1/2})$, and additionally, improve the constants from our prior work. We show how to apply the result to policy evaluation problems.

cs.LG

Approximation with Random Shallow ReLU Networks with Applications to Model Reference Adaptive Control

Neural networks are regularly employed in adaptive control of nonlinear systems and related methods of reinforcement learning. A common architecture uses a neural network with a single hidden layer (i.e. a shallow network), in which the weights and biases are fixed in advance and only the output layer is trained. While classical results show that there exist neural networks of this type that can approximate arbitrary continuous functions over bounded regions, they are non-constructive, and the networks used in practice have no approximation guarantees. Thus, the approximation properties required for control with neural networks are assumed, rather than proved. In this paper, we aim to fill this gap by showing that for sufficiently smooth functions, ReLU networks with randomly generated weights and biases achieve $L_{\infty}$ error of $O(m^{-1/2})$ with high probability, where $m$ is the number of neurons. It suffices to generate the weights uniformly over a sphere and the biases uniformly over an interval. We show how the result can be used to get approximations of required accuracy in a model reference adaptive control application.

math.OC

Non-Asymptotic Pointwise and Worst-Case Bounds for Classical Spectrum Estimators

Spectrum estimation is a fundamental methodology in the analysis of time-series data, with applications including medicine, speech analysis, and control design. The asymptotic theory of spectrum estimation is well-understood, but the theory is limited when the number of samples is fixed and finite. This paper gives non-asymptotic error bounds for a broad class of spectral estimators, both pointwise (at specific frequencies) and in the worst case over all frequencies. The general method is used to derive error bounds for the classical Blackman-Tukey, Bartlett, and Welch estimators. In particular, these are first non-asymptotic error bounds for Bartlett and Welch estimators.

math.ST

An algorithm for bilevel optimization with traffic equilibrium constraints: convergence rate analysis

Bilevel optimization with traffic equilibrium constraints plays an important role in transportation planning and management problems such as traffic control, transport network design, and congestion pricing. In this paper, we consider a double-loop gradient-based algorithm to solve such bilevel problems and provide a non-asymptotic convergence guarantee of $\mathcal{O}(K^{-1})+\mathcal{O}(λ^D)$ where $K$, $D$ are respectively the number of upper- and lower-level iterations, and $0<λ<1$ is a constant. Compared to existing literature, which either provides asymptotic convergence or makes strong assumptions and requires a complex design of step sizes, we establish convergence for choice of simple constant step sizes and considering fewer assumptions. The analysis techniques in this paper use concepts from the field of robust control and can potentially serve as a guiding framework for analyzing more general bilevel optimization algorithms.

math.OC

Function Approximation with Randomly Initialized Neural Networks for Approximate Model Reference Adaptive Control

Classical results in neural network approximation theory show how arbitrary continuous functions can be approximated by networks with a single hidden layer, under mild assumptions on the activation function. However, the classical theory does not give a constructive means to generate the network parameters that achieve a desired accuracy. Recent results have demonstrated that for specialized activation functions, such as ReLUs and some classes of analytic functions, high accuracy can be achieved via linear combinations of randomly initialized activations. These recent works utilize specialized integral representations of target functions that depend on the specific activation functions used. This paper defines mollified integral representations, which provide a means to form integral representations of target functions using activations for which no direct integral representation is currently known. The new construction enables approximation guarantees for randomly initialized networks for a variety of widely used activation functions.

math.OC

Constrained Langevin Algorithms with L-mixing External Random Variables

Langevin algorithms are gradient descent methods augmented with additive noise, and are widely used in Markov Chain Monte Carlo (MCMC) sampling, optimization, and machine learning. In recent years, the non-asymptotic analysis of Langevin algorithms for non-convex learning has been extensively explored. For constrained problems with non-convex losses over a compact convex domain with IID data variables, the projected Langevin algorithm achieves a deviation of $O(T^{-1/4} (\log T)^{1/2})$ from its target distribution [27] in $1$-Wasserstein distance. In this paper, we obtain a deviation of $O(T^{-1/2} \log T)$ in $1$-Wasserstein distance for non-convex losses with $L$-mixing data variables and polyhedral constraints (which are not necessarily bounded). This improves on the previous bound for constrained problems and matches the best-known bound for unconstrained problems.

cs.LG

Sufficient Conditions for Persistency of Excitation with Step and ReLU Activation Functions

This paper defines geometric criteria which are then used to establish sufficient conditions for persistency of excitation with vector functions constructed from single hidden-layer neural networks with step or ReLU activation functions. We show that these conditions hold when employing reference system tracking, as is commonly done in adaptive control. We demonstrate the results numerically on a system with linearly parameterized activations of this type and show that the parameter estimates converge to the true values with the sufficient conditions met.

math.OC

Wasserstein Contraction Bounds on Closed Convex Domains with Applications to Stochastic Adaptive Control

This paper is motivated by the problem of quantitatively bounding the convergence of adaptive control methods for stochastic systems to a stationary distribution. Such bounds are useful for analyzing statistics of trajectories and determining appropriate step sizes for simulations. To this end, we extend a methodology from (unconstrained) stochastic differential equations (SDEs) which provides contractions in a specially chosen Wasserstein distance. This theory focuses on unconstrained SDEs with fairly restrictive assumptions on the drift terms. Typical adaptive control schemes place constraints on the learned parameters and their update rules violate the drift conditions. To this end, we extend the contraction theory to the case of constrained systems represented by reflected stochastic differential equations and generalize the allowable drifts. We show how the general theory can be used to derive quantitative contraction bounds on a nonlinear stochastic adaptive regulation problem.

math.OC

Projected Stochastic Gradient Langevin Algorithms for Constrained Sampling and Non-Convex Learning

Langevin algorithms are gradient descent methods with additive noise. They have been used for decades in Markov chain Monte Carlo (MCMC) sampling, optimization, and learning. Their convergence properties for unconstrained non-convex optimization and learning problems have been studied widely in the last few years. Other work has examined projected Langevin algorithms for sampling from log-concave distributions restricted to convex compact sets. For learning and optimization, log-concave distributions correspond to convex losses. In this paper, we analyze the case of non-convex losses with compact convex constraint sets and IID external data variables. We term the resulting method the projected stochastic gradient Langevin algorithm (PSGLA). We show the algorithm achieves a deviation of $O(T^{-1/4}(\log T)^{1/2})$ from its target distribution in 1-Wasserstein distance. For optimization and learning, we show that the algorithm achieves $ε$-suboptimal solutions, on average, provided that it is run for a time that is polynomial in $ε^{-1}$ and slightly super-exponential in the problem dimension.

cs.LG

Causal Structure Identification from Corrupt Data-Streams

Complex networked systems can be modeled and represented as graphs, with nodes representing the agents and the links describing the dynamic coupling between them. The fundamental objective of network identification for dynamic systems is to identify causal influence pathways. However, dynamically related data-streams that originate from different sources are prone to corruption caused by asynchronous time stamps, packet drops, and noise. In this article, we show that identifying causal structure using corrupt measurements results in the inference of spurious links. A necessary and sufficient condition that delineates the effects of corruption on a set of nodes is obtained. Our theory applies to nonlinear systems, and systems with feedback loops. Our results are obtained by the analysis of conditional directed information in dynamic Bayesian networks. We provide consistency results for the conditional directed information estimator that we use by showing almost-sure convergence.

eess.SY

Network Structure Identification from Corrupt Data Streams

Complex networked systems can be modeled as graphs with nodes representing the agents and links describing the dynamic coupling between them. Previous work on network identification has shown that the network structure of linear time-invariant (LTI) systems can be reconstructed from the joint power spectrum of the data streams. These results assumed that data is perfectly measured. However, real-world data is subject to many corruptions, such as inaccurate time-stamps, noise, and data loss. We show that identifying the structure of linear time-invariant systems using corrupt measurements results in the inference of erroneous links. We provide an exact characterization and prove that such erroneous links are restricted to the neighborhood of the perturbed node. We extend the analysis of LTI systems to the case of Markov random fields with corrupt measurements. We show that data corruption in Markov random fields results in spurious probabilistic relationships in precisely the locations where spurious links arise in LTI systems.

eess.SY