SearcharxivSearch

arXiv subjects

Mohammad Boveiri

Publications and source records attributed to Mohammad Boveiri.

3 recordsLinked to original sources

Discrete-Time Adaptive Control in High Dimensions: Near Dimension-Free Performance via Mirror Descent

Motivated by the use of modern high-capacity models in real-time control problems, this paper studies the adaptive control of high-dimensional discrete-time nonlinear systems with an unknown matrix-valued parameter. We focus on regimes where the number of unknown parameter entries is large, but the parameter matrix possesses exploitable structure, such as entrywise sparsity, group sparsity, low rank, and row-stochastic or density-matrix structure. To quantify transient performance, we consider a regret criterion relative to a nominal controller with full knowledge of the true parameter. We show that standard Euclidean update schemes, including recursive least squares and gradient descent, are ill-suited to this setting: their transient performance deteriorates as the dimension increases, even when the true parameter has low intrinsic complexity. To address this limitation, we propose a novel class of mirror-descent-type adaptive laws equipped with a non-Euclidean Polyak-type step size that exploit the geometry induced by the parameter structure. For the proposed update laws, we establish asymptotic state convergence and derive regret bounds with at most logarithmic dependence on the ambient dimension. Numerical experiments demonstrate the effectiveness of the proposed schemes.

math.OC

Accelerating Adaptive Systems via Normalized Parameter Estimation Laws

In this paper, we propose a new class of parameter estimation laws for adaptive systems, called \emph{normalized parameter estimation laws}. A key feature of these estimation laws is that they accelerate the convergence of the system state, $\mathit{x(t)}$, to the origin. We quantify this improvement by showing that our estimation laws guarantee finite integrability of the $\mathit{r}$-th root of the squared norm of the system state, i.e., \( \mathit{\|x(t)\|}_2^{2/\mathit{r}} \in \mathcal{L}_1, \) where $\mathit{r} \geq 1$ is a pre-specified parameter that, for a broad class of systems, can be chosen arbitrarily large. In contrast, standard Lyapunov-based estimation laws only guarantee integrability of $\mathit{\|x(t)\|}_2^2$ (i.e., $\mathit{r} = 1$). We motivate our method by showing that, for large values of $r$, this guarantee serves as a sparsity-promoting mechanism in the time domain, meaning that it penalizes prolonged signal duration and slow decay, thereby promoting faster convergence of $\mathit{x(t)}$. The proposed estimation laws do not rely on time-varying or high adaptation gains and do not require persistent excitation. Moreover, they can be applied to systems with matched and unmatched uncertainties, regardless of their dynamic structure, as long as a control Lyapunov function (CLF) exists. Finally, they are compatible with any CLF-based certainty equivalence controllers. We further develop higher-order extensions of our estimation laws by incorporating momentum into the estimation dynamics. We illustrate the performance improvements achieved with the proposed scheme through various numerical experiments.

eess.SY

Variance-Reduced Cascade Q-learning: Algorithms and Sample Complexity

We study the problem of estimating the optimal Q-function of $γ$-discounted Markov decision processes (MDPs) under the synchronous setting, where independent samples for all state-action pairs are drawn from a generative model at each iteration. We introduce and analyze a novel model-free algorithm called Variance-Reduced Cascade Q-learning (VRCQ). VRCQ comprises two key building blocks: (i) the established direct variance reduction technique and (ii) our proposed variance reduction scheme, Cascade Q-learning. By leveraging these techniques, VRCQ provides superior guarantees in the $\ell_\infty$-norm compared with the existing model-free stochastic approximation-type algorithms. Specifically, we demonstrate that VRCQ is minimax optimal. Additionally, when the action set is a singleton (so that the Q-learning problem reduces to policy evaluation), it achieves non-asymptotic instance optimality while requiring the minimum number of samples theoretically possible. Our theoretical results and their practical implications are supported by numerical experiments.

stat.ML