SearcharxivSearch

arXiv subjects

Akshay Prasadan

Publications and source records attributed to Akshay Prasadan.

6 recordsLinked to original sources

Non-Parametric Model Calibration with Stochastic Control Parameters

We present a method for calibrating a computer model using non-parametric techniques where the inputs are stochastic but include calibration parameters whose distributions are unknown and control parameters whose distributions are specified. Our solution gives a distributional estimate over the input space that is consistent with observed field data, while also preserving the distribution of the known marginal of the control parameters. This property is desirable since stochastic inputs often include physical processes affecting the experimental conditions, and a scientifically plausible calibration estimate should preserve well-established distributional properties of these inputs. The method builds on recently developed non-parametric computer model calibration techniques based on the disintegration of measure and Bayesian inference.

stat.ME

Robust mean estimation under star-shaped constraints with heavy-tailed noise

We study the problem of robust mean estimation with adversarially contaminated data under star-shaped constraints in a heavy-tailed noise setting, where only a finite second moment $ \sigma ^2 $ is assumed. For a contamination level $ \varepsilon$ below some constant, we show that the minimax rate of the squared $ \ell_2 $ loss is $ \max( \delta ^{*2}, \varepsilon \sigma ^2) \wedge d^2 $ for a star-shaped set with diameter $ d $ (set $d = \infty$ if the set is unbounded), with $ \delta ^* $ determined via the local entropy $ \log M^\mathrm{ loc }(\delta ,c) $ as \begin{align*} \delta ^*:= \sup\bigg\{\delta \geq 0: N\frac{\delta ^2}{\sigma ^2}\leq \log M^\mathrm{ loc }(\delta ,c) \bigg\}, \end{align*} where $ c $ is a sufficiently large constant. Crucially, we require that the sample size satisfies $N \gtrsim \mathop{ \sup }\limits_{\delta \geq 0} \log M^\mathrm{ loc }(\delta ,c)$. We also show that the minimax rate is $ \max(\delta^{*2},\varepsilon ^2\sigma ^2) \wedge d^2 $ for known or sign-symmetric distributions, matching the rate achieved in the Gaussian case.

math.ST

Continuity of the Solution of a Non-Parametric Bayesian Statistical Calibration Procedure

Recent work has developed a non-parametric Bayesian approach to the calibration of a computer model, which abstractly amounts to the inversion of a pushforward of stochastic input parameters by a smooth map. The framework has been used in several complex scientific applications, motivating our investigation on the continuity of the solution operator with respect to the distribution on the input parameters. We demonstrate that the solution operator for this approach is uniformly continuous in the total variation metric and weakly continuous for a broad class of distributions.

stat.ME

Information theoretic limits of robust sub-Gaussian mean estimation under star-shaped constraints

We obtain the minimax rate for a mean location model with a bounded star-shaped set $K \subseteq \mathbb{R}^n$ constraint on the mean, in an adversarially corrupted data setting with Gaussian noise. We assume an unknown fraction $\epsilon \le 1/2-\kappa$ for some fixed $\kappa\in(0,1/2]$ of $N$ observations are arbitrarily corrupted. We obtain a minimax risk up to proportionality constants under the squared $\ell_2$ loss of $\max(\eta^{*2},\sigma^2\epsilon^2)\wedge d^2$ with \begin{align*} \eta^* = \sup \bigg\{\eta \ge 0 : \frac{N\eta^2}{\sigma^2} \leq \log \mathcal{M}_K^{\operatorname{loc}}(\eta,c)\bigg\}, \end{align*} where $\log \mathcal{M}_K^{\operatorname{loc}}(\eta,c)$ denotes the local entropy of the set $K$, $d$ is the diameter of $K$, $\sigma^2$ is the variance, and $c$ is some sufficiently large absolute constant. A variant of our algorithm achieves the same rate for settings with known or symmetric sub-Gaussian noise, with a smaller breakdown point, still of constant order. We further study the case of unknown sub-Gaussian noise and show that the rate is slightly slower: $\max(\eta^{*2},\sigma^2\epsilon^2\log(1/\epsilon))\wedge d^2$. We generalize our results to the case when $K$ is star-shaped but unbounded.

math.ST

Some facts about the optimality of the LSE in the Gaussian sequence model with convex constraint

We consider a convex constrained Gaussian sequence model and characterize necessary and sufficient conditions for the least squares estimator (LSE) to be minimax optimal. For a closed convex set $K\subset \mathbb{R}^n$ we observe $Y=\mu+\xi$ for $\xi\sim \mathcal{N}(0,\sigma^2\mathbb{I}_n)$ and $\mu\in K$ and aim to estimate $\mu$. We characterize the worst case risk of the LSE in multiple ways by analyzing the behavior of the local Gaussian width on $K$. We demonstrate that optimality is equivalent to a Lipschitz property of the local Gaussian width mapping. We also provide theoretical algorithms that search for the worst case risk. We then provide examples showing optimality or suboptimality of the LSE on various sets, including $\ell_p$ balls for $p\in[1,2]$, pyramids, solids of revolution, and multivariate isotonic regression, among others.

math.ST

Characterizing the minimax rate of nonparametric regression under bounded star-shaped constraints

We quantify the minimax rate for a nonparametric regression model over a star-shaped function class $\mathcal{F}$ with bounded diameter. We obtain a minimax rate of ${\varepsilon^{\ast}}^2\wedge\mathrm{diam}(\mathcal{F})^2$ where \[\varepsilon^{\ast} =\sup\{\varepsilon\ge 0:n\varepsilon^2 \le \log M_{\mathcal{F}}^{\operatorname{loc}}(\varepsilon,c)\},\] where $\log M_{\mathcal{F}}^{\operatorname{loc}}(\cdot, c)$ is the local metric entropy of $\mathcal{F}$, $c$ is some absolute constant scaling down the entropy radius, and our loss function is the squared population $L_2$ distance over our input space $\mathcal{X}$. In contrast to classical works on the topic [cf. Yang and Barron, 1999], our results do not require functions in $\mathcal{F}$ to be uniformly bounded in sup-norm. In fact, we propose a condition that simultaneously generalizes boundedness in sup-norm and the so-called $L$-sub-Gaussian assumption that appears in the prior literature. In addition, we prove that our estimator is adaptive to the true point in the convex-constrained case, and to the best of our knowledge this is the first such estimator in this general setting. This work builds on the Gaussian sequence framework of Neykov [2022] using a similar algorithmic scheme to achieve the minimax rate. Our algorithmic rate also applies with sub-Gaussian noise. We illustrate the utility of this theory with examples including multivariate monotone functions, linear functionals over ellipsoids, and Lipschitz classes.

math.ST