Searcharxiv⌕ Search

arXiv subjects

Arun Kumar Kuchibhotla

Publications and source records attributed to Arun Kumar Kuchibhotla.

At least 19 recordsLinked to original sources

Berry-Esseen bounds for multivariate martingale difference sequences in the Kolmogorov distance

We derive new Gaussian approximations for finite martingale difference sequences in $\mathbb{R}^d$ with respect to the Kolmogorov distance. Under appropriate conditions, our bounds exhibit a dependence of order $n^{-1/4}$ on the length of the sequence and of order $\mathrm{polylog}(d)$ on the dimension. As an application, we derive a high-dimensional Berry-Esseen bound over hyper-rectangles for martingale sequences generated from Markov chains.

math.PR↗

Inference for Median and a Generalization of HulC

It is well-documented in the literature that sample splitting offers significant methodological and theoretical advantages in statistical inference. This, for example, includes cross-fitting in double machine learning, universal inference for parametric inference, and split conformal prediction. The recently proposed inference method, HulC, also falls into this category. HulC operates by viewing the target functional of interest as an approximate median of an estimator and applying the classical distribution-free confidence intervals for median with minimal sample size. When the estimators are asymptotically normal, HulC intervals are shown to be 50\% wider than the Wald intervals asymptotically, on average, for $95\%$ coverage. Interestingly, this ratio of widths converges to a non-degenerate distribution. In this paper, we propose a generalization of HulC that are only 25\% wider than the Wald intervals for asymptotically normal estimators, irrespective of the nominal coverage. Furthermore, similarly to HulC, these generalized intervals remain valid for a significantly wider class of problems with non-normal limiting distributions. To better understand width properties under non-normal limiting distributions, we analyze distribution-free confidence intervals for the median when the Lebesgue density at the median is either zero or infinite. Surprisingly, we find that properly scaled, the interval width converges to a non-degenerate random variable.

math.ST↗

Generalized Asymptotic Limit Theory and Inference for Isotonic Regression

Monotonicity is a natural shape constraint in nonparametric regression problems, arising for instance when predicting factory yield as a monotone function of labor hours. The widely used isotonic least squares estimator (LSE) does not require any tuning parameters and its rate of convergence and pointwise limiting distribution are well studied, assuming a specific local shape for the true monotone function. We introduce a general condition on the local behavior of this true function, uncovering a far richer family of asymptotic distributions than previously known. Valid inference in the classical framework has remained challenging due to the need to estimate nuisance parameters, and no existing methods address inference in our broader setup. We resolve this by showing the symmetry of these new limiting distributions, which allows the HulC procedure of Kuchibhotla, Balakrishnan, and Wasserman (2024) to produce asymptotically valid confidence intervals. More generally, our framework enables inference that remains uniformly valid over a suitably regular class of true functions.

math.ST↗

On the L{é}vy concentration function of Gaussian quadratic forms with applications to second order U-statistics

We provide an upper-bound for the L{é}vy concentration function: $$ Q_{S}(\varepsilon):= \sup_{x \in\mathbb{R}}\mathbb{P} (x < S \leq x+\varepsilon) $$ where $S$ is a weighted sum of noncentral chi-square random variables: $$ S:= \sum_{k=1}^\infty λ_k (Z_k^2 - 1) + μ_kZ_k $$ Here, $\{Z_k\}_{k=1}^\infty$ is a sequence of independent standard Gaussian random variables and $\{λ_k\}_{k=1}^\infty, \{μ_k\}_{k=1}^\infty$ are real valued, square summable sequences. Random variables of this type often appear as limiting distributions of second order U-statistics. Our bound is adaptive, in that it recovers (up to constant factors) Gaussian type concentration function estimates if $\|λ\|_2$ is negligible compared to $\|μ\|_2$ and chi-square estimates if $\|μ\|_{2}$ is negligible compared to $\|λ\|_2$. Our bound generalizes existing bounds in various ways. In particular, we make no assumptions regarding the number of nonzero $|λ_k|$ or the size of the minimal $|λ_k|$, nor do we make any assumptions on the signs of $λ_k$. Finally, we apply our bound to some examples of interest, specifically quadratic forms that arise in limit theorems for second-order U-statistics.

math.PR↗

Doubly Robust and Efficient Calibration of Prediction Sets for Right-Censored Time-to-Event Outcomes

Our objective is to construct well-calibrated prediction sets for a time-to-event outcome subject to right-censoring with guaranteed coverage. Inspired by modern conformal inference, our approach avoids the need for a well-specified parametric or semiparametric survival model. Unlike existing conformal methods for survival data, which assume Type-I censoring with fully observed censoring times, we consider the more common right-censoring setting in which only the censoring time or only the event time is observed, whichever comes first. Under a standard conditional independence censoring condition, we propose and analyze several lower prediction bounds for the survival time of a future observation, including inverse-probability-of-censoring weighting, and its augmented version based on the semiparametric efficient influence function for the relevant marginal quantile of the outcome accounting for dependent censoring. We formally establish asymptotic coverage guarantees of the proposed methods, and demonstrate both theoretically and through empirical experiments, that the augmented approach substantially improves efficiency over all other proposed methods. Specifically, its coverage error bound is doubly robust, and therefore of second order, thus ensuring that it is asymptotically negligible relative to the coverage error of the other methods.

stat.ME↗

From Isotonic to Lipschitz Regression: A New Interpolative Perspective on Shape-restricted Estimation

This manuscript bridges nonparametric smoothness-based and shape-restricted estimation, which may appear as two disjoint paradigms in the field. The proposed approach is motivated by a conceptually simple observation: every Lipschitz function is a sum of a monotonic and a linear function. This principle is further generalized to the higher-order monotonicity and multivariate settings. A family of estimators is proposed based on a sample-splitting procedure, inheriting desirable methodological, theoretical, and computational properties of shape-restricted estimators. The theoretical analysis provides convergence guarantees of the estimator under heteroscedastic and heavy-tailed errors, as well as adaptivity to the unknown ``complexity" of the true regression function. The generality of the proposed decomposition framework is demonstrated through new approximation results and numerical studies.

stat.ME↗

Honest Inference for Stochastic Optimization

This manuscript studies a general approach to construct confidence sets for the solution of stochastic optimization, rendering empirical risk minimization as special cases. Statistical inference for stochastic optimization poses significant challenges due to the non-standard limiting behaviors of the corresponding estimator, which arise in settings with increasing dimension of parameters, non-smooth objectives, or constraints. We propose a simple and unified method that guarantees validity in both regular and irregular cases. We provide a unified treatment of validity, conservativeness, and the size of the resulting confidence sets. In particular, the presented width analysis demonstrates the adaptive behavior of the confidence set to the unknown degree of instance-specific regularity. We apply the proposed method to several high-dimensional and irregular statistical problems. Numerical results for all statistical applications are provided.

math.ST↗

On a Probability Inequality for Order Statistics with Applications to Bootstrap, Conformal Prediction, and more

``Behind every limit theorem, there is an inequality'' said Kolmogorov. We say ``for every inequality, there is an approximate inequality under approximate regularity conditions.'' Suppose $X, X'$ are independent and identically distributed random variables. Then $X \le X'$ with a probability of at least $1/2$, irrespective of the underlying (common) distribution. One can ask what happens to the probability if $X, X'$ are independent but not identically distributed. It should be approximately $1/2$ if the distributions are approximately equal. Similarly, what if the random variables are dependent? It should, again, be approximately $1/2$ if the random variables are approximately independent. We explore an extension of this probability inequality involving order statistics and develop approximate versions of such an inequality under violations of independence and identical distribution assumptions. We further show that this inequality can be used as a basis to prove asymptotic validity of bootstrap/subsampling, finite-sample validity of conformal prediction, permutation tests, and asymptotic validity of rank tests without group invariance. Specifically, in the context of resampling inference, our results can be seen as a finite-sample instantiation of some results by Peter Hall and yield an alternative ``cheap bootstrap'' that applies to high-dimensional data.

math.ST↗

Sharp Debiasing for Smooth Functional Estimation in Banach Spaces

This paper studies the estimation of smooth functionals $f(θ)$ of a mean parameter $θ= \mathbb{E}_P[W]$ for a distribution $P$ on a general Banach space. We propose a cross-fitted estimator based on a single sample splitting and establish non-asymptotic moment bounds and Berry--Esséen bounds for both $m$-smooth and infinitely smooth functionals under the finite moment assumptions. Our framework is applied to precision matrix estimation and the inference of projection parameters in high-dimensional regression. In these Euclidean settings, the proposed estimators achieve asymptotic normality under the dimension regime $d \log^2(en) = o(n)$ without requiring any structural assumptions (e.g., sparsity). We discuss computational relaxations that enables polynomial-time implementation for a range of matrix functionals.

math.ST↗

Time-uniform conformal and PAC prediction

Given that machine learning algorithms are increasingly being deployed to aid in high stakes decision-making, uncertainty quantification methods that wrap around these black box models such as conformal prediction have received much attention in recent years. In sequential settings, where data are observed/generated in a streaming fashion, traditional conformal methods do not provide any guarantee without fixing the sample size. More importantly, traditional conformal methods cannot cope with sequentially updated predictions. As such, we develop an extension of the conformal prediction and related probably approximately correct (PAC) prediction frameworks to sequential settings where the number of data points is not fixed in advance. The resulting prediction sets are anytime-valid in that their expected coverage is at the required level at any time chosen by the analyst even if this choice depends on the data. We present theoretical guarantees for our proposed methods and demonstrate their validity and utility on simulated and real datasets.

stat.ML↗

Statistical Properties of Rectified Flow

Rectified flow (Liu et al., 2022; Liu, 2022; Wu et al., 2023) is a method for defining a transport map between two distributions, and enjoys popularity in machine learning, although theoretical results supporting the validity of these methods are scant. The rectified flow can be regarded as an approximation to optimal transport, but in contrast to other transport methods that require optimization over a function space, computing the rectified flow only requires standard statistical tools such as regression or density estimation, which we leverage to develop empirical versions of transport maps. We study some structural properties of the rectified flow, including existence, uniqueness, and regularity, as well as the related statistical properties, such as rates of convergence and central limit theorems, for some selected estimators. To do so, we analyze the bounded and unbounded cases separately as each presents unique challenges. In both cases, we are able to establish convergence at faster rates than those for the usual nonparametric regression and density estimation.

math.ST↗

Inference for quantile-parametrized families via CDF confidence bands

Quantile-based distribution families are an important subclass of parametric families, capable of exhibiting a wide range of behaviors using very few parameters. These parametric models present significant challenges for classical methods, since the CDF and density do not have a closed-form expression. Furthermore, approximate maximum likelihood estimation and related procedures may yield non-$\sqrt{n}$ and non-normal asymptotics over regions of the parameter space, making bootstrap and resampling techniques unreliable. We develop a novel inference framework that constructs confidence sets by inverting distribution-free confidence bands for the empirical CDF through the known quantile function. Our proposed inference procedure provides a principled and assumption-lean alternative in this setting, requiring no distributional assumptions beyond the parametric model specification and avoiding the computational and theoretical difficulties associated with likelihood-based methods for these complex parametric families. We demonstrate our framework on Tukey Lambda and generalized Lambda distributions, evaluate its performance through simulation studies, and illustrate its practical utility with an application to both a small-sample dataset (Twin Study) and a large-sample dataset (Spanish household incomes).

stat.ME↗

Dual Induction CLT for High-dimensional m-dependent Data

We derive novel and sharp high-dimensional Berry--Esseen bounds for the sum of $m$-dependent random vectors over the class of hyper-rectangles exhibiting only a poly-logarithmic dependence in the dimension. Our results hold under minimal assumptions, such as non-degenerate covariances and finite third moments, and exhibit an optimal sample complexity of order $m^{(q-1)/(q-2)}/\sqrt{n}$. Aside from logarithmic terms, the resulting rates match the optimal rates established in the univariate case. When specialized to the sums of independent non-degenerate random vectors, our results produce sharp and, in some cases, optimal rates under the weakest possible conditions. We develop a novel inductive relationship between anti-concentration inequalities and Berry--Esseen bounds inspired by the classical Lindeberg swapping method and the concentration inequality approach for dependent data that may be of independent interest.

math.PR↗

Multiply Robust Conformal Risk Control with Coarsened Data

Conformal Prediction (CP) has recently received a tremendous amount of interest, leading to a wide range of new theoretical and methodological results for predictive inference with formal theoretical guarantees. However, the vast majority of CP methods assume that all units in the training data have fully observed data on both the outcome and covariates of primary interest, an assumption that rarely holds in practice. In reality, training data are often missing the outcome, a subset of covariates, or both on some units. In addition, time-to-event outcomes in the training set may be censored due to dropout or administrative end-of-follow-up. Accurately accounting for such coarsened data in the training sample while fulfilling the primary objective of well-calibrated conformal predictive inference, requires robustness and efficiency considerations. In this paper, we consider the general problem of obtaining distribution-free valid prediction regions for an outcome given coarsened training data. Leveraging modern semiparametric theory, we achieve our goal by deriving the efficient influence function of the quantile of the outcome we aim to predict, under a given semiparametric model for the coarsened data, carefully combined with a novel conformal risk control procedure. Our principled use of semiparametric theory has the key advantage of facilitating flexible machine learning methods such as random forests to learn the underlying nuisance functions of the semiparametric model. A straightforward application of the proposed general framework produces prediction intervals with stronger coverage properties under covariate shift, as well as the construction of multiply robust prediction sets in monotone missingness scenarios. We further illustrate the performance of our methods through various simulation studies.

math.ST↗

A Weighted Likelihood Approach Based on Statistical Data Depths

We propose a general approach to construct weighted likelihood estimating equations with the aim of obtaining robust parameter estimates. We modify the standard likelihood equations by incorporating a weight that reflects the statistical depth of each data point relative to the model, as opposed to the sample. An observation is considered regular when the corresponding difference of these two depths is close to zero. When this difference is large the observation score contribution is downweighted. We study the asymptotic properties of the proposed estimator, including consistency and asymptotic normality, for a broad class of weight functions. In particular, we establish asymptotic normality under the standard regularity conditions typically assumed for the maximum likelihood estimator (MLE). Our weighted likelihood estimator achieves the same asymptotic efficiency as the MLE in the absence of contamination, while maintaining a high degree of robustness in contaminated settings. In stark contrast to the traditional minimum divergence/disparity estimators, our results hold even if the dimension of the data diverges with the sample size, without requiring additional assumptions on the existence or smoothness of the underlying densities. We also derive the finite sample breakdown point of our estimator for both location and scatter matrix in the elliptically symmetric model. Detailed results and examples are presented for robust parameter estimation in the multivariate normal model. Robustness is further illustrated using two real data sets and a Monte Carlo simulation study.

math.ST↗

Assumption-Lean Honest Inference for $Z$-functionals

We develop a general assumption-lean framework for constructing uniformly valid confidence sets for functionals defined by moment equalities, referred to as $Z$-functionals. Our approach combines self-normalized statistics with a test inversion principle, enabling honest inference under mild regularity conditions and without explicit variance estimation. To enhance geometric tractability, we propose novel split-normalized and Gateaux-normalized statistics that yield computationally feasible and interpretable confidence sets. A central contribution of this work is a comprehensive non-asymptotic width analysis: we derive high-probability upper bounds on the diameter of the proposed confidence sets, and quantify their proximity to Wald intervals under minimal assumptions. Applications to high-dimensional non-sparse linear and generalized linear regression demonstrate that our procedures achieve valid coverage and near-optimal rate of convergence for the width/diameter, while the classical methods including Wald and bootstrap fail.

math.ST↗

Finite sample valid confidence sets of mode

Estimating the mode of a unimodal distribution is a classical problem in statistics. Although there are several approaches for point-estimation of mode in the literature, very little has been explored about the interval-estimation of mode. Our work proposes a collection of novel methods of obtaining finite sample valid confidence set of the mode of a unimodal distribution. We analyze the behaviour of the width of the proposed confidence sets under some regularity assumptions of the density about the mode and show that the width of these confidence sets shrink to zero near optimally. Simply put, we show that it is possible to build finite sample valid confidence sets for the mode that shrink to a singleton as sample size increases. We support the theoretical results by showing the performance of the proposed methods on some synthetic data-sets. We believe that our confidence sets can be improved both in construction and in terms of rate.

math.ST↗

The Berry-Esseen Bound for High-dimensional Self-normalized Sums

This manuscript studies the Gaussian approximation of the coordinate-wise maximum of self-normalized statistics in high-dimensional settings. We derive an explicit Berry-Esseen bound under weak assumptions on the absolute moments. When the third absolute moment is finite, our bound scales as $\log^{5/4}(d)/n^{1/8}$ where $n$ is the sample size and $d$ is the dimension. Hence, our bound tends to zero as long as $\log(d)=o(n^{1/10})$. Our results on self-normalized statistics represent substantial advancements, as such a bound has not been previously available in the high-dimensional central limit theorem (CLT) literature.

math.PR↗