SearcharxivSearch

arXiv subjects

Geurt Jongbloed

Publications and source records attributed to Geurt Jongbloed.

At least 19 recordsLinked to original sources

Inverting Poisson-Laguerre tessellations

While it is well-known how to compute the cells of a Laguerre tessellation for a given set of weighted generator points, it is not obvious how to invert a Laguerre tessellation. That is, given that one observes a Laguerre tessellation, how can one retrieve the weighted generators corresponding to the observed cells. In this paper, we consider inversion of a class of random Laguerre tessellations known as Poisson-Laguerre tessellations. The weighted generators of observed cells of a Poisson-Laguerre tessellation are of interest because knowledge of these weighted generators is useful for statistical inference of Poisson-Laguerre tessellations. For general Laguerre tessellations we provide a characterization of all configurations of weighted generator points which yield the same Laguerre tessellation. For Poisson-Laguerre tessellations we propose a method for consistent inversion, meaning that as one observes the tessellation through increasing observation windows, a closer approximation of the original weighted generators can be obtained. In a simulation study we examine both performance of the inversion procedure, as well as the use of the obtained approximated weighted generators for nonparametrically estimating the weight distribution function corresponding to a Poisson-Laguerre tessellation.

stat.ME

Model quality in football: Quantifying the quality of an Expected Threat model

The recent growth in data availability in football has increased the risk of incorrect use of data-driven models, making guidelines on their validation and application necessary. The Expected Threat (xT) model is an accessible option for football organisations that start building in-house methods, yet little is known about how to assess its quality. The aim of this study is twofold: to examine how the model quality depends on the number of game states and the number of training points, and to translate these results into guidelines for constructing and applying the model. Using the Markov chain underlying the model, we perform theoretical analyses and simulations to study the estimation error. These show that the estimation error is approximately lognormal for a given number of training points and game states. Additionally, we combine the simulations with expert consultation to establish the estimation error beyond which player evaluations based on the Expected Threat model become unreliable for scouting applications. From this, we derive rules of thumb to ensure the quality of an Expected Threat model before application, and we illustrate through an example how a validated model can be applied in practice. Because the approach generalises to Expected Possession Value models, this paper illustrates a framework to systematically quantify model quality, despite the ground truth being unobservable in football analytics.

stat.AP

Semiparametric Uncertainty Quantification via Isotonized Posterior for Deconvolutions

We address the problem of uncertainty quantification for the deconvolution model \(Z = X + Y\), where \(X\) and \(Y\) are nonnegative random variables and the goal is to estimate the signal's distribution of \(X \sim F_0\) supported on~\([0,\infty)\), from observations where the noise distribution is known. Existing frequentist methods often produce confidence intervals for $F_0(x)$ that depend on unknown nuisance parameters, such as the density of \(X\) and its derivative, which are difficult to estimate in practice. This paper introduces a novel and computationally efficient nonparametric Bayesian approach, based on projecting the posterior, to overcome this limitation. Our method leverages the solution \(p\) to a specific Volterra integral equation as in \cite{74}, which relates the cumulative distribution function (CDF) of the signal, \(F_0\), to the distribution of the observables. We place a Dirichlet Process prior directly on the distribution of the observed data $Z$, yielding a simple, conjugate posterior. To ensure the resulting estimates for \(F_0\) are valid CDFs, we isotonize posterior draws taking the Greatest Convex Majorant of the primitive of the posterior draws and defining what we term the Isotonic Inverse Posterior. We show that this framework yields posterior credible sets for \(F_0\) that are not only computationally fast to generate but also possess asymptotically correct frequentist coverage after a straightforward recalibration technique for the so-called Bayes Chernoff distribution introduced in \cite{54}. Our approach thus does not require the estimation of nuisance parameters to deliver uncertainty quantification for the parameter of interest $F_0(x)$. The practical effectiveness and robustness of the method are demonstrated through a simulation study with various noise distributions for $Y$.

stat.ME

The trade-off between model flexibility and accuracy of the Expected Threat model in football

With an average football (soccer) match recording over 3,000 on-ball events, effective use of this event data is essential for practitioners at football clubs to obtain meaningful insights. Models can extract more information from this data, and explainable methods can make them more accessible to practitioners. The Expected Threat model has been praised for its explainability and offers an accessible option. However, selecting the grid size is a challenging key design choice that has to be made when applying the Expected Threat model. Using a finer grid leads to a more flexible model that can better distinguish between different situations, but the accuracy of the estimates deteriorates with a more flexible model. Consequently, practitioners face challenges in balancing the trade-off between model flexibility and model accuracy. In this study, the Expected Threat model is analyzed from a theoretical perspective and simulations are performed based on the Markov chain of the model to examine its behavior in practice. Our theoretical results establish an upper bound on the error of the Expected Threat model for different flexibilities. Based on the simulations, a more accurate characterization of the model's error is provided, improving over the theoretical bound. Finally, these insights are converted into a practical rule of thumb to help practitioners choose the right balance between the model flexibility and the desired accuracy of the Expected Threat model.

stat.AP

Nonparametric Estimation in Uniform Deconvolution and Interval Censoring

In the uniform deconvolution problem one is interested in estimating the distribution function $F_0$ of a nonnegative random variable, based on a sample with additive uniform noise. A peculiar and not well understood phenomenon of the nonparametric maximum likelihood estimator in this setting is the dichotomy between the situations where $F_0(1)=1$ and $F_0(1)<1$. If $F_0(1)=1$, the MLE can be computed in a straightforward way and its asymptotic pointwise behavior can be derived using the connection to the so-called current status problem. However, if $F_0(1)<1$, one needs an iterative procedure to compute it and the asymptotic pointwise behavior of the nonparametric maximum likelihood estimator is not known. In this paper we describe the problem, connect it to interval censoring problems and a more general model studied in Groeneboom (2024) to state two competing naturally occurring conjectures for the case $F_0(1)<1$. Asymptotic arguments related to smooth functional theory and extensive simulations lead us to to bet on one of these two conjectures.

math.ST

Semiparametric Bernstein-von Mises Phenomenon via Isotonized Posterior in Wicksell's problem

In this paper, we propose a novel Bayesian approach for nonparametric estimation in Wicksell's problem. This has important applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in materials science to determine the 3D microstructure of a material, using its 2D cross sections. We deviate from the classical Bayesian nonparametric approach, which would place a Dirichlet Process (DP) prior on the distribution function of the unobservables, by directly placing a DP prior on the distribution function of the observables. Our method offers computational simplicity due to the conjugacy of the posterior and allows for asymptotically efficient estimation by projecting the posterior onto the \( \mathbb{L}_2 \) subspace of increasing, right-continuous functions. Indeed, the resulting Isotonized Inverse Posterior (IIP) satisfies a Bernstein--von Mises (BvM) phenomenon with minimax asymptotic variance \( g_0(x)/2\gamma \), where \( \gamma > 1/2 \) reflects the degree of H\"older continuity of the true cdf at \( x \). Since the IIP gives automatic uncertainty quantification, it eliminates the need to estimate \( \gamma \). Our results provide the first semiparametric Bernstein--von Mises theorem for projection-based posteriors with a DP prior in inverse problems.

math.ST

Nonparametric inference for Poisson-Laguerre tessellations

In this paper, we consider statistical inference for Poisson-Laguerre tessellations in $\mathbb{R}^d$. The object of interest is a distribution function $F$ which uniquely determines the intensity measure of the underlying Poisson process. Two nonparametric estimators for $F$ are introduced which depend only on the points of the Poisson process which generate non-empty cells and the actual cells corresponding to these points. The proposed estimators are proven to be strongly consistent, as the observation window expands unboundedly to the whole space. We also consider a stereological setting, where one is interested in estimating the distribution function associated with the Poisson process of a higher dimensional Poisson-Laguerre tessellation, given that a corresponding sectional Poisson-Laguerre tessellation is observed.

math.ST

Asymptotically efficient estimation under local constraint in Wicksell's problem

We consider nonparametric estimation of the distribution function $F$ of squared sphere radii in the classical Wicksell problem. Under smoothness conditions on $F$ in a neighborhood of $x$, in \cite{21} it is shown that the Isotonic Inverse Estimator (IIE) is asymptotically efficient and attains rate of convergence $\sqrt{n / \log n}$. If $F$ is constant on an interval containing $x$, the optimal rate of convergence increases to $\sqrt{n}$ and the IIE attains this rate adaptively, i.e.\ without explicitly using the knowledge of local constancy. However, in this case, the asymptotic distribution is not normal. In this paper, we introduce three \textit{informed} projection-type estimators of $F$, which use knowledge on the interval of constancy and show these are all asymptotically equivalent and normal. Furthermore, we establish a local asymptotic minimax lower bound in this setting, proving that the three \textit{informed} estimators are asymptotically efficient and a convolution result showing that the IIE is not efficient. We also derive the asymptotic distribution of the difference of the IIE with the efficient estimators, demonstrating that the IIE is \textit{not} asymptotically equivalent to the \textit{informed} estimators. Through a simulation study, we provide evidence that the performance of the IIE closely resembles that of its competitors.

math.ST

Existence and approximation of densities of chord length- and cross section area distributions

In various stereological problems an $n$-dimensional convex body is intersected with an $(n-1)$-dimensional Isotropic Uniformly Random (IUR) hyperplane. In this paper the cumulative distribution function associated with the $(n-1)$-dimensional volume of such a random section is studied. This distribution is also known as chord length distribution and cross section area distribution in the planar and spatial case respectively. For various classes of convex bodies it is shown that these distribution functions are absolutely continuous with respect to Lebesgue measure. A Monte Carlo simulation scheme is proposed for approximating the corresponding probability density functions.

stat.AP

Testing for no effect in regression problems: a permutation approach

Often the question arises whether $Y$ can be predicted based on $X$ using a certain model. Especially for highly flexible models such as neural networks one may ask whether a seemingly good prediction is actually better than fitting pure noise or whether it has to be attributed to the flexibility of the model. This paper proposes a rigorous permutation test to assess whether the prediction is better than the prediction of pure noise. The test avoids any sample splitting and is based instead on generating new pairings of $(X_i,Y_j)$. It introduces a new formulation of the null hypothesis and rigorous justification for the test, which distinguishes it from previous literature. The theoretical findings are applied both to simulated data and to sensor data of tennis serves in an experimental context. The simulation study underscores how the available information affects the test. It shows that the less informative the predictors, the lower the probability of rejecting the null hypothesis of fitting pure noise and emphasizes that detecting weaker dependence between variables requires a sufficient sample size.

stat.ME

Adaptive and Efficient Isotonic Estimation in Wicksell's Problem

We consider nonparametric estimation in Wicksell's problem which has relevant applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in material sciences to determine the 3D microstructure of a material, using its 2D cross sections. In the classical setting, we study the isotonized version of the plug-in estimator (IIE) for the underlying cdf $F$ of the spheres' squared radii. This estimator is fully automatic, in the sense that it does not rely on tuning parameters, and we show it is adaptive to local smoothness properties of the distribution function $F$ to be estimated. Moreover, we prove a local asymptotic minimax lower bound in this non-standard setting, with $\sqrt{\log{n}/n}$-asymptotics and where the functional $F$ to be estimated is not regular. Combined, our results prove that the isotonic estimator (IIE) is an adaptive, easy-to-compute, and efficient estimator for estimating the underlying distribution function $F$.

math.ST

Credible intervals and bootstrap confidence intervals in monotone regression

In the recent paper [5], a Bayesian approach for constructing confidence intervals in monotone regression problems is proposed, based on credible intervals. We view this method from a frequentist point of view, and show that it corresponds to a percentile bootstrap method of which we give two versions. It is shown that a (non-percentile) smoothed bootstrap method has better behavior and does not need correction for over- or undercoverage. The proofs use martingale methods.

math.ST

Confidence intervals in monotone regression

We construct bootstrap confidence intervals for a monotone regression function. It has been shown that the ordinary nonparametric bootstrap, based on the nonparametric least squares estimator (LSE) $\hat f_n$ is inconsistent in this situation. We show, however, that a consistent bootstrap can be based on the smoothed $\hat f_n$, to be called the SLSE (Smoothed Least Squares Estimator). The asymptotic pointwise distribution of the SLSE is derived. The confidence intervals, based on the smoothed bootstrap, are compared to intervals based on the (not necessarily monotone) Nadaraya Watson estimator and the effect of Studentization is investigated. We also give a method for automatic bandwidth choice, correcting work in Sen and Xu (2015). The procedure is illustrated using a well known dataset related to climate change.

math.ST

Stereological determination of particle size distributions for similar convex bodies

Consider an opaque medium which contains 3D particles. All particles are convex bodies of the same shape, but they vary in size. The particles are randomly positioned and oriented within the medium and cannot be observed directly. Taking a planar section of the medium we obtain a sample of observed 2D section profile areas of the intersected particles. In this paper the distribution of interest is the underlying 3D particle size distribution for which an identifiability result is obtained. Moreover, a nonparametric estimator is proposed for this size distribution. The estimator is proven to be consistent and its performance is assessed in a simulation study.

stat.AP

Improving state estimation through projection post-processing for activity recognition with application to football

The past decade has seen an increased interest in human activity recognition based on sensor data. Most often, the sensor data come unannotated, creating the need for fast labelling methods. For assessing the quality of the labelling, an appropriate performance measure has to be chosen. Our main contribution is a novel post-processing method for activity recognition. It improves the accuracy of the classification methods by correcting for unrealistic short activities in the estimate. We also propose a new performance measure, the Locally Time-Shifted Measure (LTS measure), which addresses uncertainty in the times of state changes. The effectiveness of the post-processing method is evaluated, using the novel LTS measure, on the basis of a simulated dataset and a real application on sensor data from football. The simulation study is also used to discuss the choice of the parameters of the post-processing method and the LTS measure.

cs.CV

Statistical Integration of Heterogeneous Data with PO2PLS

The availability of multi-omics data has revolutionized the life sciences by creating avenues for integrated system-level approaches. Data integration links the information across datasets to better understand the underlying biological processes. However, high-dimensionality, correlations and heterogeneity pose statistical and computational challenges. We propose a general framework, probabilistic two-way partial least squares (PO2PLS), which addresses these challenges. PO2PLS models the relationship between two datasets using joint and data-specific latent variables. For maximum likelihood estimation of the parameters, we implement a fast EM algorithm and show that the estimator is asymptotically normally distributed. A global test for testing the relationship between two datasets is proposed, and its asymptotic distribution is derived. Notably, several existing omics integration methods are special cases of PO2PLS. Via extensive simulations, we show that PO2PLS performs better than alternatives in feature selection and prediction performance. In addition, the asymptotic distribution appears to hold when the sample size is sufficiently large. We illustrate PO2PLS with two examples from commonly used study designs: a large population cohort and a small case-control study. Besides recovering known relationships, PO2PLS also identified novel findings. The methods are implemented in our R-package PO2PLS. Supplementary materials for this article are available online.

stat.ME

Bayesian nonparametric estimation in the current status continuous mark model

In this paper we consider the current status continuous mark model where, if the event takes place before an inspection time $T$ a "continuous mark" variable is observed as well. A Bayesian nonparametric method is introduced for estimating the distribution function of the joint distribution of the event time ($X$) and mark ($Y$). We consider a prior that is obtained by assigning a distribution on heights of cells, where cells are obtained from a partition of the support of the density of $(X, Y)$. As distribution on cell heights, we consider both a Dirichlet prior and a prior based on the graph-Laplacian on the specified partition. Our main result shows that under appropriate conditions, the posterior distribution function contracts pointwisely at rate $\left(n/\log n\right)^{-\fracρ{3(ρ+2)}}$, where $ρ$ is the Hölder smoothness of the true density. In addition to our theoretical results, we provide computational methods for drawing from the posterior using probabilistic programming. The performance of our computational methods is illustrated in two examples.

math.ST

Nonparametric Bayesian estimation of a concave distribution function with mixed interval censored data

Assume we observe a finite number of inspection times together with information on whether a specific event has occurred before each of these times. Suppose replicated measurements are available on multiple event times. The set of inspection times, including the number of inspections, may be different for each event. This is known as mixed case interval censored data. We consider Bayesian estimation of the distribution function of the event time while assuming it is concave. We provide sufficient conditions on the prior such that the resulting procedure is consistent from the Bayesian point of view. We also provide computational methods for drawing from the posterior and illustrate the performance of the Bayesian method in both a simulation study and two real datasets.

math.ST