SearcharxivSearch

arXiv subjects

Katharina Proksch

Publications and source records attributed to Katharina Proksch.

15 recordsLinked to original sources

Extending Characterizations of Multivariate Laws via Distance Distributions

We extend a theorem of Maa, Pearl, and Bartoszynski, which links equality of interpoint distance distributions to equality of underlying multivariate distributions, beyond the restrictive class of homogeneous, translation-invariant distance functions. Our approach replaces geometric assumptions on the distance with analytic conditions: volume-regularity of distance-induced balls, Lebesgue differentiability with respect to the distance, and bounded centered oscillations of densities. Under these conditions, equality of interpoint distance distributions continues to imply equality of the generating laws. The result persists under monotone continuous transformations of homogeneous, translation-invariant distances, recovering the original statement, and it extends to compact Riemannian manifolds equipped with the geodesic metric. We further develop a quantitative version of the theorem, i.e., inequalities that connect discrepancies of interpoint distance distributions to the $L^2$-distance between densities, and obtain explicit rates under Ahlfors $α$-regularity of the distance function and $β$-Hölder continuity of densities, capturing dependence on dimensionality. Several representative examples illustrate the applicability of the generalization to domain-specific distances used in modern statistics. The examples include non-homogeneous non-translation invariant distances such as Canberra, entropic distances, and the Bray--Curtis dissimilarity.

math.ST

Fatigue detection via sequential testing of biomechanical data using martingale statistic

Injuries to the knee joint are very common for long-distance and frequent runners, an issue which is often attributed to fatigue. We address the problem of fatigue detection from biomechanical data from different sources, consisting of lower extremity joint angles and ground reaction forces from running athletes with the goal of better understanding the impact of fatigue on the biomechanics of runners in general and on an individual level. This is done by sequentially testing for change in a datastream using a simple martingale test statistic. Time-uniform probabilistic martingale bounds are provided which are used as thresholds for the test statistic. Sharp bounds can be developed by a hybrid of a piece-wise linear- and a law of iterated logarithm- bound over all time regimes, where the probability of an early detection is controlled in a uniform way. If the underlying distribution of the data gradually changes over the course of a run, then a timely upcrossing of the martingale over these bounds is expected. The methods are developed for a setting when change sets in gradually in an incoming stream of data. Parameter selection for the bounds are based on simulations and methodological comparison is done with respect to existing advances. The algorithms presented here can be easily adapted to an online change-detection setting. Finally, we provide a detailed data analysis based on extensive measurements of several athletes and benchmark the fatigue detection results with the runners' individual feedback over the course of the data collection. Qualitative conclusions on the biomechanical profiles of the athletes can be made based on the shape of the martingale trajectories even in the absence of an upcrossing of the threshold.

stat.ME

A two-sample test based on averaged Wilcoxon rank sums over interpoint distances

An important class of two-sample multivariate homogeneity tests is based on identifying differences between the distributions of interpoint distances. While generating distances from point clouds offers a straightforward and intuitive way for dimensionality reduction, it also introduces dependencies to the resulting distance samples. We propose a simple test based on Wilcoxon's rank sum statistic for which we prove asymptotic normality under the null hypothesis and fixed alternatives under mild conditions on the underlying distributions of the point clouds. Furthermore, we show consistency of the test and derive a variance approximation that allows to construct a computationally feasible, distribution-free test with good finite sample performance. The power and robustness of the test for high-dimensional data and low sample sizes is demonstrated by numerical simulations. Finally, we apply the proposed test to case-control testing on microarray data in genetic studies, which is considered a notorious case for a high number of variables and low sample sizes.

stat.ME

From Small Scales to Large Scales: Distance-to-Measure Density based Geometric Analysis of Complex Data

How can we tell complex point clouds with different small scale characteristics apart, while disregarding global features? Can we find a suitable transformation of such data in a way that allows to discriminate between differences in this sense with statistical guarantees? In this paper, we consider the analysis and classification of complex point clouds as they are obtained, e.g., via single molecule localization microscopy. We focus on the task of identifying differences between noisy point clouds based on small scale characteristics, while disregarding large scale information such as overall size. We propose an approach based on a transformation of the data via the so-called Distance-to-Measure (DTM) function, a transformation which is based on the average of nearest neighbor distances. For each data set, we estimate the probability density of average local distances of all data points and use the estimated densities for classification. While the applicability is immediate and the practical performance of the proposed methodology is very good, the theoretical study of the density estimators is quite challenging, as they are based on i.i.d. observations that have been obtained via a complicated transformation. In fact, the transformed data are stochastically dependent in a non-local way that is not captured by commonly considered dependence measures. Nonetheless, we show that the asymptotic behaviour of the density estimator is driven by a kernel density estimator of certain i.i.d. random variables by using theoretical properties of U-statistics, which allows to handle the dependencies via a Hoeffding decomposition. We show via a numerical study and in an application to simulated single molecule localization microscopy data of chromatin fibers that unsupervised classification tasks based on estimated DTM-densities achieve excellent separation results.

stat.ME

Towards quantitative super-resolution microscopy: Molecular maps with statistical guarantees

Quantifying the number of molecules from fluorescence microscopy measurements is an important topic in cell biology and medical research. In this work, we present a consecutive algorithm for super-resolution (STED) scanning microscopy that provides molecule counts in automatically generated image segments and offers statistical guarantees in form of asymptotic confidence intervals. To this end, we first apply a multiscale scanning procedure on STED microscopy measurements of the sample to obtain a system of significant regions, each of which contains at least one molecule with prescribed uniform probability. This system of regions will typically be highly redundant and consists of rectangular building blocks. To choose an informative but non-redundant subset of more naturally shaped regions, we hybridize our system with the result of a generic segmentation algorithm. The diameter of the segments can be of the order of the resolution of the microscope. Using multiple photon coincidence measurements of the same sample in confocal mode, we are then able to estimate the brightness and number of the molecules and give uniform confidence intervals on the molecule counts for each previously constructed segment. In other words, we establish a so-called molecular map with uniform error control. The performance of the algorithm is investigated on simulated and real data.

stat.AP

Simultaneous inference for Berkson errors-in-variables regression under fixed design

In various applications of regression analysis, in addition to errors in the dependent observations also errors in the predictor variables play a substantial role and need to be incorporated in the statistical modeling process. In this paper we consider a nonparametric measurement error model of Berkson type with fixed design regressors and centered random errors, which is in contrast to much existing work in which the predictors are taken as random observations with random noise. Based on an estimator that takes the error in the predictor into account and on a suitable Gaussian approximation, we derive %uniform confidence statements for the function of interest. In particular, we provide finite sample bounds on the coverage error of uniform confidence bands, where we circumvent the use of extreme-value theory and rather rely on recent results on anti-concentration of Gaussian processes. In a simulation study we investigate the performance of the uniform confidence sets for finite samples.

math.ST

Gromov-Wasserstein Distance based Object Matching: Asymptotic Inference

In this paper, we aim to provide a statistical theory for object matching based on the Gromov-Wasserstein distance. To this end, we model general objects as metric measure spaces. Based on this, we propose a simple and efficiently computable asymptotic statistical test for pose invariant object discrimination. This is based on an empirical version of a $β$-trimmed lower bound of the Gromov-Wasserstein distance. We derive for $β\in[0,1/2)$ distributional limits of this test statistic. To this end, we introduce a novel $U$-type process indexed in $β$ and show its weak convergence. Finally, the theory developed is investigated in Monte Carlo simulations and applied to structural protein comparisons.

math.ST

Tests for qualitative features in the random coefficients model

The random coefficients model is an extension of the linear regression model that allows for unobserved heterogeneity in the population by modeling the regression coefficients as random variables. Given data from this model, the statistical challenge is to recover information about the joint density of the random coefficients which is a multivariate and ill-posed problem. Because of the curse of dimensionality and the ill-posedness, pointwise nonparametric estimation of the joint density is difficult and suffers from slow convergence rates. Larger features, such as an increase of the density along some direction or a well-accentuated mode can, however, be much easier detected from data by means of statistical tests. In this article, we follow this strategy and construct tests and confidence statements for qualitative features of the joint density, such as increases, decreases and modes. We propose a multiple testing approach based on aggregating single tests which are designed to extract shape information on fixed scales and directions. Using recent tools for Gaussian approximations of multivariate empirical processes, we derive expressions for the critical value. We apply our method to simulated and real data.

stat.ME

Flaring of Blazars from an Analytical, Time-dependent Model for Combined Synchrotron and Synchrotron Self-Compton Radiative Losses of Multiple Ultrarelativistic Electron Populations

We present a fully analytical, time-dependent leptonic one-zone model that describes a simplified radiation process of multiple interacting ultrarelativistic electron populations, accounting for the flaring of GeV blazars. In this model, several mono-energetic, ultrarelativistic electron populations are successively and instantaneously injected into the emission region, i.e., a magnetized plasmoid propagating along the blazar jet, and subjected to linear, time-independent synchrotron radiative losses, which are caused by a constant magnetic field, and nonlinear, time-dependent synchrotron self-Compton radiative losses in the Thomson limit. Considering a general multiple-injection scenario is, from a physical point of view, more realistic than the usual single-injection scenario invoked in common blazar models, as blazar jets may extend over tens of kiloparsecs and, thus, most likely pick up several particle populations from intermediate clouds. We analytically compute the electron number density by solving a kinetic equation using Laplace transformations and the method of matched asymptotic expansions. Moreover, we explicitly calculate the optically thin synchrotron intensity, the synchrotron self-Compton intensity in the Thomson limit, as well as the associated total fluences. In order to mimic injections of finite duration times and radiative transport, we model flares by sequences of these instantaneous injections, suitably distributed over the entire emission region. Finally, we present a parameter study for the total synchrotron and synchrotron self-Compton fluence spectral energy distributions for a generic three-injection scenario, varying the magnetic field strength, the Doppler factor, and the initial electron energy of the first injection in realistic parameter domains, demonstrating that our model can reproduce the typical broad-band behavior seen in observational data.

astro-ph.HE

Risk Estimators for Choosing Regularization Parameters in Ill-Posed Problems - Properties and Limitations

This paper discusses the properties of certain risk estimators recently proposed to choose regularization parameters in ill-posed problems. A simple approach is Stein's unbiased risk estimator (SURE), which estimates the risk in the data space, while a recent modification (GSURE) estimates the risk in the space of the unknown variable. It seems intuitive that the latter is more appropriate for ill-posed problems, since the properties in the data space do not tell much about the quality of the reconstruction. We provide theoretical studies of both estimators for linear Tikhonov regularization in a finite dimensional setting and estimate the quality of the risk estimators, which also leads to asymptotic convergence results as the dimension of the problem tends to infinity. Unlike previous papers, who studied image processing problems with a very low degree of ill-posedness, we are interested in the behavior of the risk estimators for increasing ill-posedness. Interestingly, our theoretical results indicate that the quality of the GSURE risk can deteriorate asymptotically for ill-posed problems, which is confirmed by a detailed numerical study. The latter shows that in many cases the GSURE estimator leads to extremely small regularization parameters, which obviously cannot stabilize the reconstruction. Similar but less severe issues with respect to robustness also appear for the SURE estimator, which in comparison to the rather conservative discrepancy principle leads to the conclusion that regularization parameter choice based on unbiased risk estimation is not a reliable procedure for ill-posed problems. A similar numerical study for sparsity regularization demonstrates that the same issue appears in nonlinear variational regularization approaches.

math.ST

Multiscale scanning in inverse problems

In this paper we propose a multiscale scanning method to determine active components of a quantity $f$ w.r.t. a dictionary $\mathcal{U}$ from observations $Y$ in an inverse regression model $Y=Tf+\xi$ with linear operator $T$ and general random error $\xi$. To this end, we provide uniform confidence statements for the coefficients $\langle \varphi, f\rangle$, $\varphi \in \mathcal U$, under the assumption that $(T^*)^{-1} \left(\mathcal U\right)$ is of wavelet-type. Based on this we obtain a multiple test that allows to identify the active components of $\mathcal{U}$, i.e. $\left\langle f, \varphi\right\rangle \neq 0$, $\varphi \in \mathcal U$, at controlled, family-wise error rate. Our results rely on a Gaussian approximation of the underlying multiscale statistic with a novel scale penalty adapted to the ill-posedness of the problem. The scale penalty furthermore ensures weak convergence of the statistic's distribution towards a Gumbel limit under reasonable assumptions. The important special cases of tomography and deconvolution are discussed in detail. Further, the regression case, when $T = \text{id}$ and the dictionary consists of moving windows of various sizes (scales), is included, generalizing previous results for this setting. We show that our method obeys an oracle optimality, i.e. it attains the same asymptotic power as a single-scale testing procedure at the correct scale. Simulations support our theory and we illustrate the potential of the method as an inferential tool for imaging. As a particular application we discuss super-resolution microscopy and analyze experimental STED data to locate single DNA origami.

stat.ME

Uncertainty Limits on Solutions of Inverse Problems over Multiple Orders of Magnitude using Bootstrap Methods: An Astroparticle Physics Example

Astroparticle experiments such as IceCube or MAGIC require a deconvolution of their measured data with respect to the response function of the detector to provide the distributions of interest, e.g. energy spectra. In this paper, appropriate uncertainty limits that also allow to draw conclusions on the geometric shape of the underlying distribution are determined using bootstrap methods, which are frequently applied in statistical applications. Bootstrap is a collective term for resampling methods that can be employed to approximate unknown probability distributions or features thereof. A clear advantage of bootstrap methods is their wide range of applicability. For instance, they yield reliable results, even if the usual normality assumption is violated. The use, meaning and construction of uncertainty limits to any user-specific confidence level in the form of confidence intervals and levels are discussed. The precise algorithms for the implementation of these methods, applicable for any deconvolution algorithm, are given. The proposed methods are applied to Monte Carlo simulations to show their feasibility and their precision in comparison to the statistical uncertainties calculated with the deconvolution software TRUEE.

astro-ph.IM

Multiscale inference for a multivariate density with applications to X-ray astronomy

In this paper we propose methods for inference of the geometric features of a multivariate density. Our approach uses multiscale tests for the monotonicity of the density at arbitrary points in arbitrary directions. In particular, a significance test for a mode at a specific point is constructed. Moreover, we develop multiscale methods for identifying regions of monotonicity and a general procedure for detecting the modes of a multivariate density. It is is shown that the latter method localizes the modes with an effectively optimal rate. The theoretical results are illustrated by means of a simulation study and a data example. The new method is applied to and motivated by the determination and verification of the position of high-energy sources from X-ray observations by the Swift satellite which is important for a multiwavelength analysis of objects such as Active Galactic Nuclei.

math.ST

Confidence bands for multivariate and time dependent inverse regression models

Uniform asymptotic confidence bands for a multivariate regression function in an inverse regression model with a convolution-type operator are constructed. The results are derived using strong approximation methods and a limit theorem for the supremum of a stationary Gaussian field over an increasing system of sets. As a particular application, asymptotic confidence bands for a time dependent regression function $f_t(x)$ ($x\in \mathbb {R}^d,t\in \mathbb {R}$) in a convolution-type inverse regression model are obtained. Finally, we demonstrate the practical feasibility of our proposed methods in a simulation study and an application to the estimation of the luminosity profile of the elliptical galaxy NGC5017. To the best knowledge of the authors, the results presented in this paper are the first which provide uniform confidence bands for multivariate nonparametric function estimation in inverse problems.

math.ST

Confidence Corridors for Multivariate Generalized Quantile Regression

We focus on the construction of confidence corridors for multivariate nonparametric generalized quantile regression functions. This construction is based on asymptotic results for the maximal deviation between a suitable nonparametric estimator and the true function of interest which follow after a series of approximation steps including a Bahadur representation, a new strong approximation theorem and exponential tail inequalities for Gaussian random fields. As a byproduct we also obtain confidence corridors for the regression function in the classical mean regression. In order to deal with the problem of slowly decreasing error in coverage probability of the asymptotic confidence corridors, which results in meager coverage for small sample sizes, a simple bootstrap procedure is designed based on the leading term of the Bahadur representation. The finite sample properties of both procedures are investigated by means of a simulation study and it is demonstrated that the bootstrap procedure considerably outperforms the asymptotic bands in terms of coverage accuracy. Finally, the bootstrap confidence corridors are used to study the efficacy of the National Supported Work Demonstration, which is a randomized employment enhancement program launched in the 1970s. This article has supplementary materials.

math.ST