SearcharxivSearch

arXiv subjects

Ricardo Fraiman

Publications and source records attributed to Ricardo Fraiman.

At least 19 recordsLinked to original sources

Imbalanced Classification under Capacity Constraints

Detecting observations from a minority class under severe class imbalance is a central challenge in applications such as fraud detection, medical screening, and industrial quality control. In these settings, each positive prediction triggers a costly follow-up action, an MRI scan, a transaction audit, whose execution is subject to real operational constraints. This paper proposes a formal classification framework under capacity constraints: given a user-defined bound limit $b$ on the proportion of observations that can be labeled as belonging to the minority class, the goal is to find the classifier that maximizes sensitivity on that class. We characterize the optimal classifier under this constraint and establish its equivalence with the classical Bayes classifier under a reweighting of the prior probabilities. We also introduce a capacity-adjusted performance metric $M$ that accounts for the effective detection rate when the capacity constraint is binding. The framework is implemented on top of standard learning methods, k-NN, SVM, random forests, and neural networks, and statistical consistency is established for each. We further show that these methods reduce to post-hoc thresholding when no hyperparameters are oriented toward the capacity-constrained objective, and introduce a capacity-aware support vector machine that exploits the constraint during training and achieves the strongest empirical performance. Experiments on the Taiwanese credit card default dataset confirm that capacity-constrained classifiers substantially outperform both classical approaches and SMOTE under high imbalance regimes. The framework extends naturally to multiclass settings and online environments.

stat.ML

A bootstrap approach for testing invariance under parameterized group actions: orthogonal reflections and axial symmetry

Testing whether a multivariate distribution is invariant under an orthogonal transformation is a classical problem when the transformation is fixed in advance. We address a harder situation: the transformation is unknown and must be inferred from the data. We frame this as invariance under a group action whose representation is indexed by an unknown parameter. We work in $\mathbb{R}^d$ under a simple-spectrum assumption on the covariance matrix $\Sigma$. Under this assumption, any orthogonal transformation that preserves the distribution must commute with $\Sigma$, and therefore must be a {reflection through a subspace spanned by a subset of the principal directions. This reduces the search over the orthogonal group to a finite family of candidate reflections, one for each subset of principal axes.} For each candidate, we construct a Kolmogorov--Smirnov-type statistic based on projected data and sample splitting. We derive its asymptotic distribution in a triangular-array framework and establish bootstrap validity under suitable regularity conditions. Axial symmetry about an unspecified direction (that is, invariance under reflection across an unknown one-dimensional subspace) and hyperplane (Householder) symmetry about an unspecified normal direction are the two leading particular cases: we treat them in detail, the former driving the simulation study and the latter the real-data application, to show how a concrete problem involving an unknown group representation can be successfully addressed.

math.ST

Estimating axial symmetry using random projections

This paper studies the problem of identifying directions of axial symmetry in multivariate distributions. Theoretical results are derived on how the measure or cardinality of the set of symmetry directions relates to spherical symmetry. The problem is framed using random projections, leading to a proof that in \(\RR^2\), agreement on two random projections is enough to identify the true axes of symmetry. A corresponding result for higher dimensions is conjectured. An estimator for the symmetry directions is proposed and proved to be consistent in the plane.

math.ST

A Cramér-Wold theorem for mixtures

We show how a Cramér-Wold theorem for a family of multivariate probability distributions can be used to generate a similar theorem for mixtures (convex combinations) of distributions drawn from the same family. Using this abstract result, we establish a Cramér-Wold theorem for mixtures of multivariate Gaussian distributions. According to this theorem, two such mixtures can be distinguished by projecting them onto a certain predetermined finite set of lines, the number of lines depending only on the total number Gaussian distributions involved and on the ambient dimension. A similar result is also obtained for mixtures of multivariate $t$-distributions.

math.PR

Depth based trimmed means

Robust estimation of location is a fundamental problem in statistics, particularly in scenarios where data contamination by outliers or model misspecification is a concern. In univariate settings, methods such as the sample median and trimmed means balance robustness and efficiency by mitigating the influence of extreme observations. This paper extends these robust techniques to the multivariate context through the use of data depth functions, which provide a natural means to order and rank multidimensional data. We review several depth measures and discuss their role in generalizing trimmed mean estimators beyond one dimension. Our main contributions are twofold: first, we prove the almost sure consistency of the multivariate trimmed mean estimator under mixing conditions; second, we establish a general limit distribution theorem for a broad family of depth-based estimators, encompassing popular examples such as Tukey's and projection depth. These theoretical advancements not only enhance the understanding of robust location estimation in high-dimensional settings but also offer practical guidelines for applications in areas such as machine learning, economic analysis, and financial risk assessment. A small example with simulated data is performed, varying the depth measure used and the percentage of trimmed data.

math.ST

Two statistical problems for multivariate mixture distributions

We address two important statistical problems: that of estimating mixtures of multivariate normal distributions and mixtures of $t$-distributions based on univariate projections, and that of quantifying a discrepancy between mixture distributions induced by two model-based clusterings. In the second problem, rather than introducing a direct metric on partitions, we propose a model-based distributional discrepancy between the fitted mixture distributions associated with two clusterings. The results are based on an earlier work of the authors, where it was shown that mixtures of multivariate Gaussian or $t$-distributions can be distinguished by projecting them onto a certain predetermined finite set of lines, the number of lines depending only on the total number of distributions involved and on the ambient dimension. We also compare our proposal with robust versions of the expectation-maximization method EM. In each case, we present algorithms for effecting the task, and compare them with existing methods by carrying out some simulations.

math.ST

Home range estimation under a restricted sampling scheme

The analysis of animal movement has gained attention recently. New continuous-time models and statistical methods have been developed to estimate some sets related to their movements, such as the home-range and the core-area among others, when the information of the trajectory is provided by a GPS. Because data transfer costs and GPS battery life are practical constraints in ecological studies, the experimental designer must make critical sampling decisions in order to maximize information. To capture fine-scale motion, long-term behavior must be sacrificed, and vice versa. To overcome this limitation, we introduce the on--off sampling scheme, where the GPS is alternately on and off. This scheme is already used in practice but with insufficient statistical theoretical support. We prove the consistency of home-range estimators with an underlying reflected diffusion model under this sampling method (in terms of the Hausdorff distance). The same rate of convergence is achieved as in the case where the GPS is always on for the whole experiment. This is illustrated by a simulation study and real data. We also provide estimators of the stationary distribution, its level sets (which give estimators of the core area), and the drift function.

math.ST

A quantitative Heppes Theorem and multivariate Bernoulli distributions

Using some extensions of a theorem of Heppes on finitely supported discrete probability measures, we address the problems of classification and testing based on projections. In particular, when the support of the distributions is known in advance (as for instance for multivariate Bernoulli distributions), a single suitably chosen projection determines the distribution. Several applications of these results are considered.

math.PR

A Cramér-Wold theorem for elliptical distributions

According to a well-known theorem of Cramér and Wold, if $P$ and $Q$ are two Borel probability measures on $\mathbb{R}^d$ whose projections $P_L,Q_L$ onto each line $L$ in $\mathbb{R}^d$ satisfy $P_L=Q_L$, then $P=Q$. Our main result is that, if $P$ and $Q$ are both elliptical distributions, then, to show that $P=Q$, it suffices merely to check that $P_L=Q_L$ for a certain set of $(d^2+d)/2$ lines $L$. Moreover $(d^2+d)/2$ is optimal. The class of elliptical distributions contains the Gaussian distributions as well as many other multivariate distributions of interest. Our theorem contrasts with other variants of the Cramér-Wold theorem, in that no assumption is made about the finiteness of moments of $P$ and $Q$. We use our results to derive a statistical test for equality of elliptical distributions, and carry out a small simulation study of the test, comparing it with other tests from the literature. We also give an application to learning (binary classification), again illustrated with a small simulation

math.PR

Universally consistent estimation of the reach

The reach of a set $M \subset \mathbb R^d$, also known as condition number when $M$ is a manifold, was introduced by Federer in 1959. The reach is a central concept in geometric measure theory, set estimation, manifold learning, among others areas. We introduce a universally consistent estimate of the reach, just assuming that the reach is positive. Under an additional assumption we provide rates of convergence. We also show that it is not possible to determine, based on a finite sample, if the reach of the support of a density is zero or not. We provide a small simulation study and a bias correction method for the case when $M$ is a manifold.

math.ST

Statistical analysis of measures of non-convexity

Several measures of non-convexity (departures from convexity) have been introduced in the literature, both for sets and functions. Some of them are of geometric nature, while others are more of topological nature. We address the statistical analysis of some of these measures of non-convexity of a set $S$, by dealing with their estimation based on a sample of points in $S$. We introduce also a new measure of non-convexity. We discuss briefly about these different notions of non-convexity, prove consistency and find the asymptotic distribution for the proposed estimators. We also consider the practical implementation of these estimators and illustrate their applicability to a real data example.

math.ST

Estimation of surface area

We study the problem of estimating the surface area of the boundary $\partial S$ of a sufficiently smooth set $S\subset\mathbb{R}^d$ when the available information is only a finite subset $\X\subset S$. We propose two estimators. The first makes use of the Devroye--Wise support estimator and is based on Crofton's formula, which, roughly speaking, states that the $(d-1)$-dimensional surface area of a smooth enough set is the mean number of intersections of randomly chosen lines. For that purpose, we propose an estimator of the number of intersections of such lines with support based on the Devroye--Wise support estimators. The second surface area estimator makes use of the $α$-convex hull of $\X$, which is denoted by $C_α(\X)$. More precisely, it is the $(d-1)$-dimensional surface area of $C_α(\X)$, as denoted by $|C_α(\X)|_{d-1}$, which is proven to converge to the $(d-1)$-dimensional surface area of $\partial S$. Moreover, $|C_α(\X)|_{d-1}$ can be computed using Crofton's formula. Our results depend on the Hausdorff distance between $S$ and $\X$ for the Devroye--Wise estimator, and the Hausdorff distance between $\partial S$ and $\partial C_α(\X)$ for the second estimator.

math.ST

Application of the Cramér-Wold theorem to testing for invariance under group actions

We address the problem of testing for the invariance of a probability measure under the action of a group of linear transformations. We propose a procedure based on consideration of one-dimensional projections, justified using a variant of the Cramér-Wold theorem. Our test procedure is powerful, computationally efficient, and dimension-independent, extending even to the case of infinite-dimensional spaces (multivariate functional data). It includes, as special cases, tests for exchangeability and sign-invariant exchangeability. We compare our procedure with some previous proposals in these cases, in a small simulation study. The paper concludes with two real-data examples.

math.ST

Level set and density estimation on manifolds

We tackle the problem of the estimation of the level sets L_f(λ) of the density f of a random vector X supported on a smooth manifold M\subsetR^d , from an iid sample of X. To do that we introduce a kernel-based estimator f^n,h , which is a slightly modified version of the one proposed in [45], and proves its a.s. uniform convergence to f . Then, we propose two estimators of L f (λ), the first one is a plug-in: L f^n,h (λ), which is proven to be a.s. consistent in Hausdorff distance and distance in measure, if L f(λ) does not meet the boundary of M . While the second one assumes that L f(λ) is r-convex, and is estimated by means of the r-convex hull of L f^n,h(λ). The performance of our proposal is illustrated through some simulated examples. In a real data example we analyze the intensity and direction of strong and moderate winds.

math.ST

Level sets of depth measures in abstract spaces

The lens depth of a point has been recently extended to general metric spaces, which is not the case for most depths. It is defined as the probability of being included in the intersection of two random balls centred at two random points X and Y, with the same radius d(X, Y). We study the consistency in Hausdorff and measure distance, of the level sets of the empirical lens depth, based on an iid sample on a general metric space. We also prove that the boundary of the empirical level sets are consistent estimators of their population counterparts, and analyze two real-life examples

math.ST

Retrieving the structure of probabilistic sequences of auditory stimuli from EEG data

Using a new probabilistic approach we model the relationship between sequences of auditory stimuli generated by stochastic chains and the electroencephalographic (EEG) data acquired while 19 participants were exposed to those stimuli. The structure of the chains generating the stimuli are characterized by rooted and labeled trees whose leaves, henceforth called contexts, represent the sequences of past stimuli governing the choice of the next stimulus. A classical conjecture claims that the brain assigns probabilistic models to samples of stimuli. If this is true, then the context tree generating the sequence of stimuli should be encoded in the brain activity. Using an innovative statistical procedure we show that this context tree can effectively be extracted from the EEG data, thus giving support to the classical conjecture.

q-bio.NC

Weighted lens depth: Some applications to supervised classification

Starting with Tukey's pioneering work in the 1970's, the notion of depth in statistics has been widely extended especially in the last decade. These extensions include high dimensional data, functional data, and manifold-valued data. In particular, in the learning paradigm, the depth-depth method has become a useful technique. In this paper we extend the notion of lens depth to the case of data in metric spaces, and prove its main properties, with particular emphasis on the case of Riemannian manifolds, where we extend the concept of lens depth in such a way that it takes into account non-convex structures on the data distribution. Next we illustrate our results with some simulation results and also in some interesting real datasets, including pattern recognition in phylogenetic trees using the depth--depth approach.

math.ST

On semi-supervised learning

Semi-supervised learning deals with the problem of how, if possible, to take advantage of a huge amount of unclassified data, to perform a classification in situations when, typically, there is little labeled data. Even though this is not always possible (it depends on how useful, for inferring the labels, it would be to know the distribution of the unlabeled data), several algorithm have been proposed recently. %but in general they are not proved to outperform A new algorithm is proposed, that under almost necessary conditions, %and it is proved that it attains asymptotically the performance of the best theoretical rule as the amount of unlabeled data tends to infinity. The set of necessary assumptions, although reasonable, show that semi-supervised classification only works for very well conditioned problems. The focus is on understanding when and why semi-supervised learning works when the size of the initial training sample remains fixed and the asymptotic is on the size of the unlabeled data. The performance of the algorithm is assessed in the well known "Isolet" real-data of phonemes, where a strong dependence on the choice of the initial training sample is shown.

stat.ML