Searcharxiv⌕ Search

arXiv subjects

Amparo Baíllo

Publications and source records attributed to Amparo Baíllo.

8 recordsLinked to original sources

Bootstrap tests for almost goodness-of-fit

We introduce the \textit{almost goodness-of-fit} test, a procedure to assess whether a (parametric) model provides a good representation of the probability distribution generating the observed sample. Specifically, given a distribution function $F$ and a parametric family $\mathcal{G}=\{ G(\boldsymbolθ) : \boldsymbolθ \in Θ\}$, we consider the testing problem \[ H_0: \| F - G(\boldsymbolθ_F) \|_p \geq ε\quad \text{vs} \quad H_1: \| F - G(\boldsymbolθ_F) \|_p < ε, \] where $ε>0$ is a margin of error and $G(\boldsymbolθ_F)$ denotes a representative of $F$ within the parametric class. The approximate model is determined via an M-estimator of the parameters. %The objective is the approximate validation of a distribution or an entire parametric family up to a pre-specified threshold value. The methodology also quantifies the percentage improvement of the proposed model relative to a non-informative (constant) benchmark. The test statistic is the $\mathrm{L}^p$-distance between the empirical distribution function and that of the estimated model. We present two consistent, easy-to-implement, and flexible bootstrap schemes to carry out the test. The performance of the proposal is illustrated through simulation studies and analysis and real-data applications.

stat.ME↗

Kernel K-means clustering of distributional data

We consider the problem of clustering a sample of probability distributions from a random distribution on $\mathbb R^p$. Our proposed partitioning method makes use of a symmetric, positive-definite kernel $k$ and its associated reproducing kernel Hilbert space (RKHS) $\mathcal H$. By mapping each distribution to its corresponding kernel mean embedding in $\mathcal H$, we obtain a sample in this RKHS where we carry out the $K$-means clustering procedure, which provides an unsupervised classification of the original sample. The procedure is simple and computationally feasible even for dimension $p>1$. The simulation studies provide insight into the choice of the kernel and its tuning parameter. The performance of the proposed clustering procedure is illustrated on a collection of Synthetic Aperture Radar (SAR) images.

stat.ML↗

Tests for almost stochastic dominance

We introduce a 2-dimensional stochastic dominance (2DSD) index to characterize both strict and almost stochastic dominance. Based on this index, we derive an estimator for the minimum violation ratio (MVR), also known as the critical parameter, of the almost stochastic ordering condition between two variables. We determine the asymptotic properties of the empirical 2DSD index and MVR for the most frequently used stochastic orders. We also provide conditions under which the bootstrap estimators of these quantities are strongly consistent. As an application, we develop consistent bootstrap testing procedures for almost stochastic dominance. The performance of the tests is checked via simulations and the analysis of real data.

econ.EM↗

Extremal points of Lorenz curves and applications to inequality analysis

We find the set of extremal points of Lorenz curves with fixed Gini index and compute the maximal $L^1$-distance between Lorenz curves with given values of their Gini coefficients. As an application we introduce a bidimensional index that simultaneously measures relative inequality and dissimilarity between two populations. This proposal employs the Gini indices of the variables and an $L^1$-distance between their Lorenz curves. The index takes values in a right-angled triangle, two of whose sides characterize perfect relative inequality-expressed by the Lorenz ordering between the underlying distributions. Further, the hypotenuse represents maximal distance between the two distributions. As a consequence, we construct a chart to, graphically, either see the evolution of (relative) inequality and distance between two income distributions over time or to compare the distribution of income of a specific population between a fixed time point and a range of years. We prove the mathematical results behind the above claims and provide a full description of the asymptotic properties of the plug-in estimator of this index. Finally, we apply the proposed bidimensional index to several real EU-SILC income datasets to illustrate its performance in practice.

econ.EM↗

A survey and a new selection criterion for statistical home range estimation

The home range of a specific animal describes the geographic area where this individual spends most of the time while carrying out its usual activities (eating, resting, reproduction, ...). Although a well-established definition of this concept is lacking, there is a variety of home range estimators. The first objective of this work is to review and categorize the statistical methodologies proposed in the literature to approximate the home range of an animal, based on a sample of observed locations. The second aim is to address the open question of choosing the "best" home range from a collection of them based on the same sample. We introduce a numerical index, based on a penalization criterion, to rank the estimated home ranges. The key idea is to balance the excess area covered by the estimator (with respect to the original sample) and a shape descriptor measuring the over-adjustment of the home range to the data. To our knowledge, apart from computing the home range area, our ranking procedure is the first one which is both applicable to real data and to any type of home range estimator. Further, the optimization of the selection index provides in fact a way to select the smoothing parameter for the kernel home range estimator. For clarity of exposition, we have applied all the estimation procedures and our selection proposal to a set of real locations of a Mongolian wolf using R as the statistical software. As a byproduct, this review contains a thorough revision of the implementation of home range estimators in the R language.

q-bio.QM↗

New distance measures for classifying X-ray astronomy data into stellar classes

The classification of the X-ray sources into classes (such as extragalactic sources, background stars, ...) is an essential task in astronomy. Typically, one of the classes corresponds to extragalactic radiation, whose photon emission behaviour is well characterized by a homogeneous Poisson process. We propose to use normalized versions of the Wasserstein and Zolotarev distances to quantify the deviation of the distribution of photon interarrival times from the exponential class. Our main motivation is the analysis of a massive dataset from X-ray astronomy obtained by the Chandra Orion Ultradeep Project (COUP). This project yielded a large catalog of 1616 X-ray cosmic sources in the Orion Nebula region, with their series of photon arrival times and associated energies. We consider the plug-in estimators of these metrics, determine their asymptotic distributions, and illustrate their finite-sample performance with a Monte Carlo study. We estimate these metrics for each COUP source from three different classes. We conclude that our proposal provides a striking amount of information on the nature of the photon emitting sources. Further, these variables have the ability to identify X-ray sources wrongly catalogued before. As an appealing conclusion, we show that some sources, previously classified as extragalactic emissions, have a much higher probability of being young stars in Orion Nebula.

math.ST↗

Homogeneity tests for Michaelis-Menten curves with application to fluorescence resonance energy transfer data

Resonance energy transfer methods are in wide use for evaluating protein-protein interactions and protein conformational changes in living cells. Fluorescence resonance energy transfer (FRET) measures energy transfer as a function of the acceptor:donor ratio, generating FRET saturation curves. Modeling these curves by Michaelis-Menten kinetics allows characterization by two parameters, which serve to evaluate apparent affinity between two proteins and to compare this affinity in different experimental conditions. To reduce the effect of sampling variability, several statistical samples of the saturation curve are generated in the same biological conditions. Here we study three procedures to determine whether statistical samples in a collection are homogeneous, in the sense that they are extracted from the same regression model. From the hypothesis testing viewpoint, we considered an F test and a procedure based on bootstrap resampling. The third method analyzed the problem from the model selection viewpoint, and used the Akaike information criterion (AIC). Although we only considered the Michaelis-Menten model, all statistical procedures would be applicable to any other nonlinear regression model. We compared the performance of the homogeneity testing methods in a Monte Carlo study and through analysis in living cells of FRET saturation curves for dimeric complexes of CXCR4, a seven-transmembrane receptor of the G protein-coupled receptor family. We show that the F test, the bootstrap procedure and the model selection method lead in general to similar conclusions, although AIC gave the best results when sample sizes were small, whereas the F test and the bootstrap method were more appropriate for large samples. In practice, all three methods are easy to use simultaneously and show consistency, facilitating conclusions on sample homogeneity.

q-bio.BM↗

Supervised classification for a family of Gaussian functional models

In the framework of supervised classification (discrimination) for functional data, it is shown that the optimal classification rule can be explicitly obtained for a class of Gaussian processes with "triangular" covariance functions. This explicit knowledge has two practical consequences. First, the consistency of the well-known nearest neighbors classifier (which is not guaranteed in the problems with functional data) is established for the indicated class of processes. Second, and more important, parametric and nonparametric plug-in classifiers can be obtained by estimating the unknown elements in the optimal rule. The performance of these new plug-in classifiers is checked, with positive results, through a simulation study and a real data example.

stat.ML↗