SearcharxivSearch

arXiv subjects

Eliza O'Reilly

Publications and source records attributed to Eliza O'Reilly.

17 recordsLinked to original sources

Wasserstein stability of the zero cell of a Poisson hyperplane tessellation under directional perturbations

A stationary Poisson hyperplane process in $\mathbb{R}^d$ is characterized by an intensity parameter and an even probability measure on the unit sphere called the directional distribution. In this work, we investigate the stability of the zero cell, i.e., the random convex polytope of the induced hyperplane tessellation containing the origin, under perturbations of the directional distribution. Our results provide quantitative bounds establishing local Hölder continuity of the distribution of the zero cell with respect to Wasserstein metrics on the space of random convex bodies and probability distributions on the unit sphere. As an application, we establish stability bounds for density estimators constructed from Poisson hyperplane tessellations. In particular, we bound the expected total variation distance between the zero-cell-based estimated probability measures in terms of the Wasserstein distance between their directional distributions.

math.PR

Statistical Advantages of Oblique Randomized Decision Trees and Forests

This work studies the statistical implications of using features comprised of general linear combinations of covariates to partition the data in randomized decision tree and forest regression algorithms. Using random tessellation theory in stochastic geometry, we provide a theoretical analysis of a class of efficiently generated random tree and forest estimators that allow for oblique splits along such features. We call these estimators oblique Mondrian trees and forests, as the trees are generated by first selecting a set of features from linear combinations of the covariates and then running a Mondrian process that hierarchically partitions the data along these features. Quadratic risk bounds and convergence rates are obtained for the flexible function class of multi-index models for dimension reduction, where the output is assumed to depend on a low-dimensional relevant feature subspace of the input domain. The results highlight how the risk of these estimators depends on the choice of features and quantify how robust the risk is with respect to error between the selected features along which the data is split and the true relevant feature subspace. The asymptotic analysis also provides conditions on the convergence rate a set of estimated relevant features must satisfy for oblique Mondrian estimators to obtain minimax optimal rates of convergence with respect to the dimension of the relevant feature subspace. Additionally, a lower bound on the risk of axis-aligned Mondrian trees (where features are restricted to the set of covariates) is obtained, proving that these estimators are suboptimal for general ridge functions, no matter how the distribution over the covariates used to divide the data at each tree node is weighted.

math.ST

TrIM: Transformed Iterative Mondrian Forests for Gradient-based Dimension Reduction and High-Dimensional Regression

We propose a computationally efficient algorithm for gradient-based linear dimension reduction and high-dimensional regression. The algorithm initially computes a Mondrian forest and uses this estimator to identify a relevant feature subspace of the inputs from an estimate of the expected gradient outer product (EGOP) of the regression function. In addition, we introduce an iterative approach known as Transformed Iterative Mondrian (TrIM) forest to improve the Mondrian forest estimator by using the EGOP estimate to update the set of features and weights used by the Mondrian partitioning mechanism. We obtain consistency guarantees and convergence rates for estimating the EGOP matrix and the random forest estimator obtained from one iteration of the TrIM algorithm. Lastly, we demonstrate the effectiveness of our proposed algorithm for learning the relevant feature subspace across various settings with both simulated and real data.

math.ST

Operatopes, Operanoids, and Noncommutative Zonoids

We study a class of convex bodies called operatopes that are obtained by taking Minkowski sums of affine images of an operator norm ball. This notion generalizes that of zonotopes which are Minkowksi sums of line segments. Taking the limit of the number of line segments to infinity yields the class of convex bodies called zonoids, which can also be viewed as the expectation of a random line segment. Expanding on this interpretation, we analogously define operanoids as the expectation of a random affine image of an operator norm ball. In studying the properties of operanoids when the dimension of the operator norm ball grows, we arrive at a new asymptotic regime for limits of convex bodies. This leads to the more general class of convex bodies called noncommutative zonoids, and we use the framework of free probability theory to illustrate basic properties and examples. Finally, we discuss applications of operanoids and noncommutative zonoids in statistics and stochastic processes.

math.MG

Optimal Regularization Under Uncertainty: Distributional Robustness and Convexity Constraints

Regularization is a central tool for addressing ill-posedness in inverse problems and statistical estimation, with the choice of a suitable penalty often determining the reliability and interpretability of downstream solutions. While recent work has characterized optimal regularizers for well-specified data distributions, practical deployments are often complicated by distributional uncertainty and the need to enforce structural constraints such as convexity. In this paper, we introduce a framework for distributionally robust optimal regularization, which identifies regularizers that remain effective under perturbations of the data distribution. Our approach leverages convex duality to reformulate the underlying distributionally robust optimization problem, eliminating the inner maximization and yielding formulations that are amenable to numerical computation. We show how the resulting robust regularizers interpolate between memorization of the training distribution and uniform priors, providing insights into their behavior as robustness parameters vary. For example, we show how certain ambiguity sets, such as those based on the Wasserstein-1 distance, naturally induce regularity in the optimal regularizer by promoting regularizers with smaller Lipschitz constants. We further investigate the setting where regularizers are required to be convex, formulating a convex program for their computation and illustrating their stability with respect to distributional shifts. Taken together, our results provide both theoretical and computational foundations for designing regularizers that are reliable under model uncertainty and structurally constrained for robust deployment.

math.OC

The Uniformly Rotated Mondrian Kernel

Random feature maps are used to decrease the computational cost of kernel machines in large-scale problems. The Mondrian kernel is one such example of a fast random feature approximation of the Laplace kernel, generated by a computationally efficient hierarchical random partition of the input space known as the Mondrian process. In this work, we study a variation of this random feature map by applying a uniform random rotation to the input space before running the Mondrian process to approximate a kernel that is invariant under rotations. We obtain a closed-form expression for the isotropic kernel that is approximated, as well as a uniform convergence rate of the uniformly rotated Mondrian kernel to this limit. To this end, we utilize techniques from the theory of stationary random tessellations in stochastic geometry and prove a new result on the geometry of the typical cell of the superposition of uniformly rotated Mondrian tessellations. Finally, we test the empirical performance of this random feature map on both synthetic and real-world datasets, demonstrating its improved performance over the Mondrian kernel on a dataset that is debiased from the standard coordinate axes.

cs.LG

The Star Geometry of Critic-Based Regularizer Learning

Variational regularization is a classical technique to solve statistical inference tasks and inverse problems, with modern data-driven approaches parameterizing regularizers via deep neural networks showcasing impressive empirical performance. Recent works along these lines learn task-dependent regularizers. This is done by integrating information about the measurements and ground-truth data in an unsupervised, critic-based loss function, where the regularizer attributes low values to likely data and high values to unlikely data. However, there is little theory about the structure of regularizers learned via this process and how it relates to the two data distributions. To make progress on this challenge, we initiate a study of optimizing critic-based loss functions to learn regularizers over a particular family of regularizers: gauges (or Minkowski functionals) of star-shaped bodies. This family contains regularizers that are commonly employed in practice and shares properties with regularizers parameterized by deep neural networks. We specifically investigate critic-based losses derived from variational representations of statistical distances between probability measures. By leveraging tools from star geometry and dual Brunn-Minkowski theory, we illustrate how these losses can be interpreted as dual mixed volumes that depend on the data distribution. This allows us to derive exact expressions for the optimal regularizer in certain cases. Finally, we identify which neural network architectures give rise to such star body gauges and when do such regularizers have favorable properties for optimization. More broadly, this work highlights how the tools of star geometry can aid in understanding the geometry of unsupervised regularizer learning.

cs.LG

Optimal Regularization for a Data Source

In optimization-based approaches to inverse problems and to statistical estimation, it is common to augment criteria that enforce data fidelity with a regularizer that promotes desired structural properties in the solution. The choice of a suitable regularizer is typically driven by a combination of prior domain information and computational considerations. Convex regularizers are attractive computationally but they are limited in the types of structure they can promote. On the other hand, nonconvex regularizers are more flexible in the forms of structure they can promote and they have showcased strong empirical performance in some applications, but they come with the computational challenge of solving the associated optimization problems. In this paper, we seek a systematic understanding of the power and the limitations of convex regularization by investigating the following questions: Given a distribution, what is the optimal regularizer for data drawn from the distribution? What properties of a data source govern whether the optimal regularizer is convex? We address these questions for the class of regularizers specified by functionals that are continuous, positively homogeneous, and positive away from the origin. We say that a regularizer is optimal for a data distribution if the Gibbs density with energy given by the regularizer maximizes the population likelihood (or equivalently, minimizes cross-entropy loss) over all regularizer-induced Gibbs densities. As the regularizers we consider are in one-to-one correspondence with star bodies, we leverage dual Brunn-Minkowski theory to show that a radial function derived from a data distribution is akin to a ``computational sufficient statistic'' as it is the key quantity for identifying optimal regularizers and for assessing the amenability of a data source to convex regularization.

math.OC

Minimax Rates for High-Dimensional Random Tessellation Forests

Random forests are a popular class of algorithms used for regression and classification. The algorithm introduced by Breiman in 2001 and many of its variants are ensembles of randomized decision trees built from axis-aligned partitions of the feature space. One such variant, called Mondrian forests, was proposed to handle the online setting and is the first class of random forests for which minimax rates were obtained in arbitrary dimension. However, the restriction to axis-aligned splits fails to capture dependencies between features, and random forests that use oblique splits have shown improved empirical performance for many tasks. In this work, we show that a large class of random forests with general split directions also achieve minimax optimal convergence rates in arbitrary dimension. This class includes STIT forests, a generalization of Mondrian forests to arbitrary split directions, as well as random forests derived from Poisson hyperplane tessellations. These are the first results showing that random forest variants with oblique splits can obtain minimax optimality in arbitrary dimension. Our proof technique relies on the novel application of the theory of stationary random tessellations in stochastic geometry to statistical learning theory.

math.ST

Spectrahedral Regression

Convex regression is the problem of fitting a convex function to a data set consisting of input-output pairs. We present a new approach to this problem called spectrahedral regression, in which we fit a spectrahedral function to the data, i.e. a function that is the maximum eigenvalue of an affine matrix expression of the input. This method represents a significant generalization of polyhedral (also called max-affine) regression, in which a polyhedral function (a maximum of a fixed number of affine functions) is fit to the data. We prove bounds on how well spectrahedral functions can approximate arbitrary convex functions via statistical risk analysis. We also analyze an alternating minimization algorithm for the non-convex optimization problem of fitting the best spectrahedral function to a given data set. We show that this algorithm converges geometrically with high probability to a small ball around the optimal parameter given a good initialization. Finally, we demonstrate the utility of our approach with experiments on synthetic data sets as well as real data arising in applications such as economics and engineering design.

math.OC

Stochastic geometry to generalize the Mondrian Process

The stable under iterated tessellation (STIT) process is a stochastic process that produces a recursive partition of space with cut directions drawn independently from a distribution over the sphere. The case of random axis-aligned cuts is known as the Mondrian process. Random forests and Laplace kernel approximations built from the Mondrian process have led to efficient online learning methods and Bayesian optimization. In this work, we utilize tools from stochastic geometry to resolve some fundamental questions concerning STIT processes in machine learning. First, we show that a STIT process with cut directions drawn from a discrete distribution can be efficiently simulated by lifting to a higher dimensional axis-aligned Mondrian process. Second, we characterize all possible kernels that stationary STIT processes and their mixtures can approximate. We also give a uniform convergence rate for the approximation error of the STIT kernels to the targeted kernels, generalizing the work of [3] for the Mondrian case. Third, we obtain consistency results for STIT forests in density estimation and regression. Finally, we give a formula for the density estimator arising from an infinite STIT random forest. This allows for precise comparisons between the Mondrian forest, the Mondrian kernel and the Laplace kernel in density estimation. Our paper calls for further developments at the novel intersection of stochastic geometry and machine learning.

stat.ML

Couplings for determinantal point processes and their reduced Palm distributions with a view to quantifying repulsiveness

For a determinantal point process $X$ with a kernel $K$ whose spectrum is strictly less than one, Andr{é} Goldman has established a coupling to its reduced Palm process $X^u$ at a point $u$ with $K(u,u)>0$ so that almost surely $X^u$ is obtained by removing a finite number of points from $X$. We sharpen this result, assuming weaker conditions and establishing that $X^u$ can be obtained by removing at most one point from $X$, where we specify the distribution of the difference $ξ_u:=X\setminus X^u$. This is used for discussing the degree of repulsiveness in DPPs in terms of $ξ_u$, including Ginibre point processes and other specific parametric models for DPPs.

math.PR

Thin-shell concentration for zero cells of stationary Poisson mosaics

We study the concentration of the norm of a random vector $Y$ uniformly sampled in the centered zero cell of two types of stationary and isotropic random mosaics in $\mathbb{R}^n$ for large dimensions $n$. For a stationary and isotropic Poisson-Voronoi mosaic, $Y$ has a radial and log-concave distribution, implying that ${|Y|}/{\mathbb{E}(|Y|^2)^{\frac{1}{2}}}$ approaches one for large $n$. Assuming the cell intensity of the random mosaic scales like $e^{n ρ_n}$, where $\lim_{n \to \infty} ρ_n = ρ$, $|Y|$ is on the order of $\sqrt{n}$ for large $n$. For the Poisson-Voronoi mosaic, we show that $|Y|/\sqrt{n}$ concentrates to $e^{-ρ}(2πe)^{-\frac{1}{2}}$ as $n$ increases, and for a stationary and isotropic Poisson hyperplane mosaic, we show there is a range $(R_{\ell}, R_u)$ such that ${|Y|}/{\sqrt{n}}$ will be within this range with high probability for large $n$. The rates of convergence are also computed in both cases.

math.PR

Facets of spherical random polytopes

Facets of the convex hull of $n$ independent random vectors chosen uniformly at random from the unit sphere in $\mathbb{R}^d$ are studied. A particular focus is given on the height of the facets as well as the expected number of facets as the dimension increases. Regimes for $n$ and $d$ with different asymptotic behavior of these quantities are identified and asymptotic formulas in each case are established. Extensions of some known results in fixed dimension to the case where dimension tends to infinity are described.

math.PR

The stochastic geometry of unconstrained one-bit data compression

A stationary stochastic geometric model is proposed for analyzing the data compression method used in one-bit compressed sensing. The data set is an unconstrained stationary set, for instance all of $\mathbb{R}^n$ or a stationary Poisson point process in $\mathbb{R}^n$. It is compressed using a stationary and isotropic Poisson hyperplane tessellation, assumed independent of the data. That is, each data point is compressed using one bit with respect to each hyperplane, which is the side of the hyperplane it lies on. This model allows one to determine how the intensity of the hyperplanes must scale with the dimension $n$ to ensure sufficient separation of different data by the hyperplanes as well as sufficient proximity of the data compressed together. The results have direct implications in compressive sensing and in source coding.

math.PR

Reach of Repulsion for Determinantal Point Processes in High Dimensions

Goldman [7] proved that the distribution of a stationary determinantal point process (DPP) $Φ$ can be coupled with its reduced Palm version $Φ^{0,!}$ such that there exists a point process $η$ where $Φ= Φ^{0,!} \cup η$ in distribution and $Φ^{0,!} \cap η= \emptyset$. The points of $η$ characterize the repulsive nature of a typical point of $Φ$. In this paper, the first moment measure of $η$ is used to study the repulsive behavior of DPPs in high dimensions. It is shown that many families of DPPs have the property that the total number of points in $η$ converges in probability to zero as the space dimension $n$ goes to infinity. It is also proved that for some DPPs there exists an $R^*$ such that the decay of the first moment measure of $η$ is slowest in a small annulus around the sphere of radius $\sqrt{n}R^*$. This $R^*$ can be interpreted as the asymptotic reach of repulsion of the DPP. Examples of classes of DPP models exhibiting this behavior are presented and an application to high dimensional Boolean models is given.

math.PR

End-to-End Optimization of High Throughput DNA Sequencing

At the core of high throughput DNA sequencing platforms lies a bio-physical surface process that results in a random geometry of clusters of homogenous short DNA fragments typically hundreds of base pairs long - bridge amplification. The statistical properties of this random process and length of the fragments are critical as they affect the information that can be subsequently extracted, i.e., density of successfully inferred DNA fragment reads. The ensemble of overlapping DNA fragment reads are then used to computationally reconstruct the much longer target genome sequence, e.g, ranging from hundreds of thousands to billions of base pairs. The success of the reconstruction in turn depends on having a sufficiently large ensemble of DNA fragments that are sufficiently long. In this paper using stochastic geometry we model and optimize the end-to-end process linking and partially controlling the statistics of the physical processes to the success of the computational step. This provides, for the first time, a framework capturing salient features of such sequencing platforms that can be used to study cost, performance or sensitivity of the sequencing process.

q-bio.GN