SearcharxivSearch

arXiv subjects

Lutz Duembgen

Publications and source records attributed to Lutz Duembgen.

At least 19 recordsLinked to original sources

A New Algorithm for Totally Positive Approximations

We revisit the problem of approximating a bivariate distribution with finite support by another such distribution which is totally positive of order two (TP2). Approximation is meant in a maximum likelihood sense.

stat.CO

On the tails of log-concave density estimators

It is shown that the nonparametric maximum likelihood estimator of a univariate log-concave probability density satisfies desirable consistency properties in the tail regions. Specifically, let $P$ and $f$ denote the true underlying distribution and density, respectively. If $\hat{f}_n$ is the estimated log-concave density, and $\hatφ_n = \log \hat{f}_n$, then we specify sequences $(b_n)_{n\in \mathbb{N}}$ such that $P([b_n,\infty)) \to 0$ at a specific speed, ensuring that the absolute errors or absolute relative errors of $\hat{f}_n, \ \hatφ_n$ and $\hatφ_n'$ converge to zero uniformly on sets $[a, b_n]$. The main tools, besides characterizations of $\hat{f}_n$, are exponential and maximal inequalities for truncated moments of log-concave distributions, which are of independent interest.

math.ST

Nonparametric Smoothing of Directional and Axial Data

We discuss generalized linear models for directional data where the conditional distribution of the response is a von Mises-Fisher distribution in arbitrary dimension or a Bingham distribution on the unit circle. To do this properly, we parametrize von Mises-Fisher distributions by Euclidean parameters and investigate computational aspects of this parametrization. Then we modify this approach for local polynomial regression as a means of nonparametric smoothing of distributional data. The methods are illustrated with simulated data and a data set from planetary sciences involving covariate vectors on a sphere with axial response.

stat.ME

Connecting model-based and model-free approaches to linear least squares regression

In a regression setting with a response vector and given regressor vectors, a typical question is to what extent the response is related to these regressors, specifically, how well it can be approximated by a linear combination of the latter. Classical methods for this question are based on statistical models for the conditional distribution of the response, given the regressors. In the present paper it is shown that various p-values resulting from this model-based approach have also a purely data-analytic, model-free interpretation. This finding is derived in a rather general context. In addition, we introduce equivalence regions, a reinterpretation of confidence regions in the model-free context.

math.ST

Estimation of a Likelihood Ratio Ordered Family of Distributions

Consider bivariate observations $(X_1,Y_1), \ldots, (X_n,Y_n) \in \mathbb{R}\times \mathbb{R}$ with unknown conditional distributions $Q_x$ of $Y$, given that $X = x$. The goal is to estimate these distributions under the sole assumption that $Q_x$ is isotonic in $x$ with respect to likelihood ratio order. If the observations are identically distributed, a related goal is to estimate the joint distribution $\mathcal{L}(X,Y)$ under the sole assumption that it is totally positive of order two in a certain sense. An algorithm is developed which estimates the unknown family of distributions $(Q_x)_x$ via empirical likelihood. The benefit of the stronger regularization imposed by likelihood ratio order over the usual stochastic order is evaluated in terms of estimation and predictive performances on simulated as well as real data.

math.ST

Various New Inequalities for Beta Distributions

This note provides some new inequalities and approximations for beta distributions, including tail inequalities, exponential inequalities of Hoeffding and Bernstein type, Gaussian inequalities and approximations.

math.ST

Approximating Symmetrized Estimators of Scatter via Balanced Incomplete U-Statistics

We derive limiting distributions of symmetrized estimators of scatter, where instead of all $n(n-1)/2$ pairs of the $n$ observations we only consider $nd$ suitably chosen pairs, $1 \le d < \lfloor n/2\rfloor$. It turns out that the resulting estimators are asymptotically equivalent to the original one whenever $d = d(n) \to \infty$ at arbitrarily slow speed. We also investigate the asymptotic properties for arbitrary fixed $d$. These considerations and numerical examples indicate that for practical purposes, moderate fixed values of $d$ between,say, $10$ and $20$ yield already estimators which are computationally feasible and rather close to the original ones.

math.ST

On Stochastic Orders and Total Positivity

The usual stochastic order and the likelihood ratio order between probability distributions on the real line are reviewed in full generality. In addition, for the distribution of a random pair $(X,Y)$, it is shown that the conditional distributions of $Y$, given $X = x$, are increasing in $x$ with respect to the likelihood ratio order if and only if the joint distribution of $(X,Y)$ is totally positive of order two (TP2) in a certain sense. It is also shown that these three types of constraints are stable under weak convergence, and that weak convergence of TP2 distributions implies convergence of the conditional distributions just mentioned.

math.ST

Honest calibration assessment for binary outcome predictions

Probability predictions from binary regressions or machine learning methods ought to be calibrated: If an event is predicted to occur with probability $x$, it should materialize with approximately that frequency, which means that the so-called calibration curve $p(\cdot)$ should equal the identity, $p(x) = x$ for all $x$ in the unit interval. We propose honest calibration assessment based on novel confidence bands for the calibration curve, which are valid only subject to the natural assumption of isotonicity. Besides testing the classical goodness-of-fit null hypothesis of perfect calibration, our bands facilitate inverted goodness-of-fit tests whose rejection allows for the sought-after conclusion of a sufficiently well specified model. We show that our bands have a finite sample coverage guarantee, are narrower than existing approaches, and adapt to the local smoothness of the calibration curve $p$ and the local variance of the binary observations. In an application to model predictions of an infant having a low birth weight, the bounds give informative insights on model calibration.

math.ST

A New Approach to Tests and Confidence Bands for Distribution Functions

We introduce new goodness-of-fit tests and corresponding confidence bands for distribution functions. They are inspired by multi-scale methods of testing and based on refined laws of the iterated logarithm for the normalized uniform empirical process $\mathbb{U}_n (t)/\sqrt{t(1-t)}$ and its natural limiting process, the normalized Brownian bridge process $\mathbb{U}(t)/\sqrt{t(1-t)}$. The new tests and confidence bands refine the procedures of Berk and Jones (1979) and Owen (1995). Roughly speaking, the high power and accuracy of the latter methods in the tail regions of distributions are essentially preserved while gaining considerably in the central region. The goodness-of-fit tests perform well in signal detection problems involving sparsity, as in Ingster (1997), Donoho and Jin (2004) and Jager and Wellner (2007), but also under contiguous alternatives. Our analysis of the confidence bands sheds new light on the influence of the underlying $ϕ$-divergences.

math.ST

Bounding distributional errors via density ratios

We present some new and explicit error bounds for the approximation of distributions. The approximation error is quantified by the maximal density ratio of the distribution $Q$ to be approximated and its proxy $P$. This non-symmetric measure is more informative than and implies bounds for the total variation distance. Explicit approximation problems include, among others, hypergeometric by binomial distributions, binomial by Poisson distributions, and beta by gamma distributions. In many cases we provide both upper and (matching) lower bounds.

math.ST

Honest Confidence Bands for Isotonic Quantile Curves

We provide confidence bands for isotonic quantile curves in nonparametric univariate regression with guaranteed given coverage probability. The method is an adaptation of the confidence bands of Duembgen and Johns (2004) for isotonic median curves.

math.ST

Refining Invariant Coordinate Selection via Local Projection Pursuit

Independent component selection (ICS), introduced by Tyler et al. (2009, JRSS B), is a powerful tool to find potentially interesting projections of multivariate data. In some cases, some of the projections proposed by ICS come close to really interesting ones, but little deviations can result in a blurred view which does not reveal the feature (e.g. a clustering) which would otherwise be clearly visible. To remedy this problem, we propose an automated and localized version of projection pursuit (PP), cf. Huber (1985, Ann. Statist.}. Precisely, our local search is based on gradient descent applied to estimated differential entropy as a function of the projection matrix.

stat.CO

Exact Confidence Bounds in Discrete Models -- Algorithmic Aspects of Sterne's Method

In this manuscript we review two methods to construct exact confidence bounds for an unknown real parameter in a general class of discrete statistical models. These models include the binomial family, the Poisson family as well as distributions connected to odds ratios in two-by-two tables. In particular, we discuss Sterne's (1954) method in our general framework and present an explicit algorithm for the computation of the resulting confidence bounds. The methods are illustrated with various examples.

stat.CO

Active set algorithms for estimating shape-constrained density ratios

In many instances, imposing a constraint on the shape of a density is a reasonable and flexible assumption. It offers an alternative to parametric models which can be too rigid and to other nonparametric methods requiring the choice of tuning parameters. This paper treats the nonparametric estimation of log-concave or log-convex density ratios by means of active set algorithms in a unified framework. In the setting of log-concave densities, the new algorithm is similar to but substantially faster than previously considered active set methods. Log-convexity is a less common shape constraint which is described by some authors as "tail inflation". The active set method proposed here is novel in this context. As a by-product, new goodness-of-fit tests of single hypotheses are formulated and are shown to be more powerful than higher criticism tests in a simulation study.

stat.CO

Local Estimation of a Multivariate Density and its Derivatives

We analyze four different approaches to estimate a multivariate probability density (or the log-density) and its first and second order derivatives. Two methods, local log-likelihood and local Hyvärinen score estimation, are in terms of weighted scoring rules with local quadratic models. The other two approaches are matching of local moments and kernel density estimation. All estimators depend on a general kernel, and we use the Gaussian kernel to provide explicit examples. Asymptotic properties of the estimators are derived and compared. In terms of rates of convergence, a refined local moment matching estimator is the best.

math.ST

The density ratio of Poisson binomial versus Poisson distributions

Let $b(x)$ be the probability that a sum of independent Bernoulli random variables with parameters $p_1, p_2, p_3, \ldots \in [0,1)$ equals $x$, where $λ:= p_1 + p_2 + p_3 + \cdots$ is finite. We prove two inequalities for the maximal ratio $b(x)/π_λ(x)$, where $π_λ$ is the weight function of the Poisson distribution with parameter $λ$.

math.ST