SearcharxivSearch

arXiv subjects

Jon A. Wellner

Publications and source records attributed to Jon A. Wellner.

At least 19 recordsLinked to original sources

A New Approach to Tests and Confidence Bands for Distribution Functions

We introduce new goodness-of-fit tests and corresponding confidence bands for distribution functions. They are inspired by multi-scale methods of testing and based on refined laws of the iterated logarithm for the normalized uniform empirical process $\mathbb{U}_n (t)/\sqrt{t(1-t)}$ and its natural limiting process, the normalized Brownian bridge process $\mathbb{U}(t)/\sqrt{t(1-t)}$. The new tests and confidence bands refine the procedures of Berk and Jones (1979) and Owen (1995). Roughly speaking, the high power and accuracy of the latter methods in the tail regions of distributions are essentially preserved while gaining considerably in the central region. The goodness-of-fit tests perform well in signal detection problems involving sparsity, as in Ingster (1997), Donoho and Jin (2004) and Jager and Wellner (2007), but also under contiguous alternatives. Our analysis of the confidence bands sheds new light on the influence of the underlying $ϕ$-divergences.

math.ST

Hardy's Inequality and Its Descendants

We formulate and prove a generalization of Hardy's inequality (Hardy,1925) in terms of random variables and show that it contains the usual (or familiar) continuous and discrete forms of Hardy's inequality. Next we improve the recent version by Li and Mao of Hardy's inequality with weights for general Borel measures and mixed norms so that it implies the discrete version of Liao and the Hardy inequality with weights of Muckenhoupt as well as the mixed norm versions due to Hardy and Littlewood, Bliss, and Bradley. An equivalent formulation in terms of random variables is given as well. We also formulate a reverse version of Hardy's inequality, the closely related Copson inequality, a reverse Copson inequality and a Carleman-Pólya-Knopp inequality via random variables. Finally we connect our Copson inequality with counting process martingales and survival analysis, and briefly discuss other applications.

math.PR

Bi-$s^*$-Concave Distributions

We introduce new shape-constrained classes of distribution functions on R, the bi-$s^*$-concave classes. In parallel to results of Dümbgen, Kolesnyk, and Wilke (2017) for what they called the class of bi-log-concave distribution functions, we show that every $s$-concave density $f$ has a bi-$s^*$-concave distribution function $F$ for $s^*\leq s/(s+1)$. Confidence bands building on existing nonparametric bands, but accounting for the shape constraint of bi-$s^*$-concavity, are also considered. The new bands extend those developed by Dümbgen et al. (2017) for the constraint of bi-log-concavity. We also make connections between bi-$s^*$-concavity and finiteness of the Csörgő - Révész constant of $F$ which plays an important role in the theory of quantile processes.

math.ST

The density ratio of Poisson binomial versus Poisson distributions

Let $b(x)$ be the probability that a sum of independent Bernoulli random variables with parameters $p_1, p_2, p_3, \ldots \in [0,1)$ equals $x$, where $λ:= p_1 + p_2 + p_3 + \cdots$ is finite. We prove two inequalities for the maximal ratio $b(x)/π_λ(x)$, where $π_λ$ is the weight function of the Poisson distribution with parameter $λ$.

math.ST

Complex sampling designs: uniform limit theorems and applications

In this paper, we develop a general approach to proving global and local uniform limit theorems for the Horvitz-Thompson empirical process arising from complex sampling designs. Global theorems such as Glivenko-Cantelli and Donsker theorems, and local theorems such as local asymptotic modulus and related ratio-type limit theorems are proved for both the Horvitz-Thompson empirical process, and its calibrated version. Limit theorems of other variants and their conditional versions are also established. Our approach reveals an interesting feature: the problem of deriving uniform limit theorems for the Horvitz-Thompson empirical process is essentially no harder than the problem of establishing the corresponding finite-dimensional limit theorems. These global and local uniform limit theorems are then applied to important statistical problems including (i) $M$-estimation (ii) $Z$-estimation (iii) frequentist theory of Bayes procedures, all with weighted likelihood, to illustrate their wide applicability.

math.ST

Univariate log-concave density estimation with symmetry or modal constraints

We study nonparametric maximum likelihood estimation of a log-concave density function $f_0$ which is known to satisfy further constraints, where either (a) the mode $m$ of $f_0$ is known, or (b) $f_0$ is known to be symmetric about a fixed point $m$. We develop asymptotic theory for both constrained log-concave maximum likelihood estimators (MLE's), including consistency, global rates of convergence, and local limit distribution theory. In both cases, we find the MLE's pointwise limit distribution at $m$ (either the known mode or the known center of symmetry) and at a point $x_0 \ne m$. Software to compute the constrained estimators is available in the R package \verb+logcondens.mode+. The symmetry-constrained MLE is particularly useful in contexts of location estimation. The mode-constrained MLE is useful for mode-regression. The mode-constrained MLE can also be used to form a likelihood ratio test for the location of the mode of $f_0$. These problems are studied in separate papers. In particular, in a separate paper we show that, under a curvature assumption, the likelihood ratio statistic for the location of the mode can be used for hypothesis tests or confidence intervals that do not depend on either tuning parameters or nuisance parameters.

math.ST

Convergence rates of least squares regression estimators with heavy-tailed errors

We study the performance of the Least Squares Estimator (LSE) in a general nonparametric regression model, when the errors are independent of the covariates but may only have a $p$-th moment ($p\geq 1$). In such a heavy-tailed regression setting, we show that if the model satisfies a standard `entropy condition' with exponent $α\in (0,2)$, then the $L_2$ loss of the LSE converges at a rate \begin{align*} \mathcal{O}_{\mathbf{P}}\big(n^{-\frac{1}{2+α}} \vee n^{-\frac{1}{2}+\frac{1}{2p}}\big). \end{align*} Such a rate cannot be improved under the entropy condition alone. This rate quantifies both some positive and negative aspects of the LSE in a heavy-tailed regression setting. On the positive side, as long as the errors have $p\geq 1+2/α$ moments, the $L_2$ loss of the LSE converges at the same rate as if the errors are Gaussian. On the negative side, if $p<1+2/α$, there are (many) hard models at any entropy level $α$ for which the $L_2$ loss of the LSE converges at a strictly slower rate than other robust estimators. The validity of the above rate relies crucially on the independence of the covariates and the errors. In fact, the $L_2$ loss of the LSE can converge arbitrarily slowly when the independence fails. The key technical ingredient is a new multiplier inequality that gives sharp bounds for the `multiplier empirical process' associated with the LSE. We further give an application to the sparse linear regression model with heavy-tailed covariates and errors to demonstrate the scope of this new inequality.

math.ST

Inference for the mode of a log-concave density

We study a likelihood ratio test for the location of the mode of a log-concave density. Our test is based on comparison of the log-likelihoods corresponding to the unconstrained maximum likelihood estimator of a log-concave density and the constrained maximum likelihood estimator where the constraint is that the mode of the density is fixed, say at $m$. The constrained estimation problem is studied in detail in Doss and Wellner [2018]. Here the results of that paper are used to show that, under the null hypothesis (and strict curvature of $-\log f$ at the mode), the likelihood ratio statistic is asymptotically pivotal: that is, it converges in distribution to a limiting distribution which is free of nuisance parameters, thus playing the role of the $χ_1^2$ distribution in classical parametric statistical problems. By inverting this family of tests we obtain new (likelihood ratio based) confidence intervals for the mode of a log-concave density $f$. These new intervals do not depend on any smoothing parameters. We study the new confidence intervals via Monte Carlo methods and illustrate them with two real data sets. The new intervals seem to have several advantages over existing procedures. Software implementing the test and confidence intervals is available in the R package \verb+logcondens.mode+.

math.ST

Robustness of shape-restricted regression estimators: an envelope perspective

Classical least squares estimators are well-known to be robust with respect to moment assumptions concerning the error distribution in a wide variety of finite-dimensional statistical problems; generally only a second moment assumption is required for least squares estimators to maintain the same rate of convergence that they would satisfy if the errors were assumed to be Gaussian. In this paper, we give a geometric characterization of the robustness of shape-restricted least squares estimators (LSEs) to error distributions with an $L_{2,1}$ moment, in terms of the `localized envelopes' of the model. This envelope perspective gives a systematic approach to proving oracle inequalities for the LSEs in shape-restricted regression problems in the random design setting, under a minimal $L_{2,1}$ moment assumption on the errors. The canonical isotonic and convex regression models, and a more challenging additive regression model with shape constraints are studied in detail. Strikingly enough, in the additive model both the adaptation and robustness properties of the LSE can be preserved, up to error distributions with an $L_{2,1}$ moment, for estimating the shape-constrained proxy of the marginal $L_2$ projection of the true regression function. This holds essentially regardless of whether or not the additive model structure is correctly specified. The new envelope perspective goes beyond shape constrained models. Indeed, at a general level, the localized envelopes give a sharp characterization of the convergence rate of the $L_2$ loss of the LSE between the worst-case rate as suggested by the recent work of the authors [25], and the best possible parametric rate.

math.ST

On the isoperimetric constant, covariance inequalities and $L_p$-Poincaré inequalities in dimension one

Firstly, we derive in dimension one a new covariance inequality of $L_{1}-L_{\infty}$ type that characterizes the isoperimetric constant as the best constant achieving the inequality. Secondly, we generalize our result to $L_{p}-L_{q}$ bounds for the covariance. Consequently, we recover Cheeger's inequality without using the co-area formula. We also prove a generalized weighted Hardy type inequality that is needed to derive our covariance inequalities and that is of independent interest. Finally, we explore some consequences of our covariance inequalities for $L_{p}$-Poincaré inequalities and moment bounds. In particular, we obtain optimal constants in general $L_{p}$-Poincaré inequalities for measures with finite isoperimetric constant, thus generalizing in dimension one Cheeger's inequality, which is a $L_{p}$-Poincaré inequality for $p=2$, to any real $p\geq 1$.

math.PR

Efron's monotonicity property for measures on $\mathbb{R}^2$

First we prove some kernel representations for the covariance of two functions taken on the same random variable and deduce kernel representations for some functionals of a continuous one-dimensional measure. Then we apply these formulas to extend Efron's monotonicity property, given in Efron [1965] and valid for independent log-concave measures, to the case of general measures on $\mathbb{R}^2$. The new formulas are also used to derive some further quantitative estimates in Efron's monotonicity property.

math.ST

Estimation of mean residual life

Yang (1978) considered an empirical estimate of the mean residual life function on a fixed finite interval. She proved it to be strongly uniformly consistent and (when appropriately standardized) weakly convergent to a Gaussian process. These results are extended to the whole half line, and the variance of the the limiting process is studied. Also, nonparametric simultaneous confidence bands for the mean residual life function are obtained by transforming the limiting process to Brownian motion.

math.ST

Bi-$s^*$-concave distributions

We introduce a new shape-constrained class of distribution functions on R, the bi-$s^*$-concave class. In parallel to results of Dümbgen, Kolesnyk, and Wilke (2017) for what they called the class of bi-log-concave distribution functions, we show that every s-concave density f has a bi-$s^*$-concave distribution function $F$ and that every bi-$s^*$-concave distribution function satisfies $γ(F) \le 1/(1+s)$ where finiteness of $$ γ(F) \equiv \sup_{x} F(x) (1-F(x)) \frac{| f' (x)|}{f^2 (x)}, $$ the Csörgő - Révész constant of F, plays an important role in the theory of quantile processes on $R$.

math.ST

The Bennett-Orlicz norm

Lederer and van de Geer (2013) introduced a new Orlicz norm, the Bernstein-Orlicz norm, which is connected to Bernstein type inequalities. Here we introduce another Orlicz norm, the Bennett-Orlicz norm, which is connected to Bennett type inequalities. The new Bennett-Orlicz norm yields inequalities for expectations of maxima which are potentially somewhat tighter than those resulting from the Bernstein-Orlicz norm when they are both applicable. We discuss cross connections between these norms, exponential inequalities of the Bernstein, Bennett, and Prokhorov types, and make comparisons with results of Talagrand (1989, 1994), and Boucheron, Lugosi, and Massart (2013).

math.ST

Entropy of convex functions on $R^d$

Let $Ω$ be a bounded closed convex set in ${\mathbb R}^d$ with non-empty interior, and let ${\cal C}_r(Ω)$ be the class of convex functions on $Ω$ with $L^r$-norm bounded by $1$. We obtain sharp estimates of the $ε$-entropy of ${\cal C}_r(Ω)$ under $L^p(Ω)$ metrics, $1\le p \frac{dr}{d+(d-1)r}$ is attained by the closed unit ball. While a general convex body can be approximated by inscribed polytopes, the entropy rate does not carry over to the limiting body. Our results have applications to questions concerning rates of convergence of nonparametric estimators of high-dimensional shape-constrained functions.

math.ST

Finite sampling inequalities: an application to two-sample Kolmogorov-Smirnov statistics

We review a finite-sampling exponential bound due to Serfling and discuss related exponential bounds for the hypergeometric distribution. We then discuss how such bounds motivate some new results for two-sample empirical processes. Our development complements recent results by Wei and Dudley (2011) concerning exponential bounds for two-sided Kolmogorov - Smirnov statistics by giving corresponding results for one-sided statistics with emphasis on "adjusted" inequalities of the type proved originally by Dvoretzky, Kiefer, and Wolfowitz (1956) and by Massart (1990) for one-sample versions of these statistics.

math.ST

Exponential bounds for the hypergeometric distribution

We establish exponential bounds for the hypergeometric distribution which include a finite sampling correction factor, but are otherwise analogous to bounds for the binomial distribution due to León and Perron (2003) and Talagrand (1994). We also establish a convex ordering for sampling without replacement from populations of real numbers between zero and one: a population of all zeros or ones (and hence yielding a hypergeometric distribution in the upper bound) gives the extreme case.

math.ST

Multivariate convex regression: global risk bounds and adaptation

We study the problem of estimating a multivariate convex function defined on a convex body in a regression setting with random design. We are interested in optimal rates of convergence under a squared global continuous $l_2$ loss in the multivariate setting $(d\geq 2)$. One crucial fact is that the minimax risks depend heavily on the shape of the support of the regression function. It is shown that the global minimax risk is on the order of $n^{-2/(d+1)}$ when the support is sufficiently smooth, but that the rate $n^{-4/(d+4)}$ is when the support is a polytope. Such differences in rates are due to difficulties in estimating the regression function near the boundary of smooth regions. We then study the natural bounded least squares estimators (BLSE): we show that the BLSE nearly attains the optimal rates of convergence in low dimensions, while suffering rate-inefficiency in high dimensions. We show that the BLSE adapts nearly parametrically to polyhedral functions when the support is polyhedral in low dimensions by a local entropy method. We also show that the boundedness constraint cannot be dropped when risk is assessed via continuous $l_2$ loss. Given rate sub-optimality of the BLSE in higher dimensions, we further study rate-efficient adaptive estimation procedures. Two general model selection methods are developed to provide sieved adaptive estimators (SAE) that achieve nearly optimal rates of convergence for particular "regular" classes of convex functions, while maintaining nearly parametric rate-adaptivity to polyhedral functions in arbitrary dimensions. Interestingly, the uniform boundedness constraint is unnecessary when risks are measured in discrete $l_2$ norms.

math.ST