SearcharxivSearch

arXiv subjects

Raúl Gouet

Publications and source records attributed to Raúl Gouet.

8 recordsLinked to original sources

Estimating the tail index of Pareto-type distributions from geometric records

In this paper, we develop a novel inferential approach based on geometric records for estimating the tail index of heavy-tailed distributions. We construct a maximum likelihood estimator for the Pareto model and establish strong consistency and asymptotic normality, providing also an explicit expression for the asymptotic variance. These results are then extended to a broad class of Pareto-type distributions. The performance of the estimator is assessed via Monte Carlo simulation and compared with classical estimators from the literature. The proposed method is particularly well suited for settings where data arrive sequentially, as it yields smooth estimation trajectories. It is also especially advantageous in applications such as destructive testing, where measuring each item is costly. In this context, the estimator achieves a comparable level of estimation accuracy to Hill's estimator, but with a considerably lower number of fully measured items. An application to the analysis of the distribution of fluctuations of the Dow Jones Industrial Average (DJI) is also presented.

math.ST

Characterisation of distributions via record-like observations

We characterise probability distributions via a martingale property associated with a natural generalisation of record values, known as $δ$-records. For an independent and identically distributed sequence $(X_n)$ with running maximum $M_n$, let $N_n$ be the number of $δ$-records (those $X_k$ with $X_k>M_{k-1}+δ$). We determine distributions for which $N_n-cM_n$ is a martingale, and show that this property uniquely determines the underlying distribution within broad classes. We show that the problem can be reformulated in terms of a delay-integrated Cauchy functional equation. A distinctive feature of this equation is that it is required to hold on a set that depends on the unknown distribution itself, which both complicates the analysis and allows for a rich variety of solutions. A complete characterisation is obtained when $δ<0$. For $δ>0$, all solutions with bounded support are identified. In the case of $δ>0$ and unbounded support, we consider both continuous and lattice distributions. In the continuous case, the characterisation reduces to a delay differential equation, which admits classical exponential-type solutions as well as broader families, including mixtures of exponential and gamma distributions. An analogous discrete analysis leads to difference equations whose solutions include mixtures of geometric and negative binomial distributions. In particular, this yields a new characterisation of the geometric distribution based on weak records.

math.PR

Stochastic ordering, attractiveness and couplings in non-conservative particle systems

We analyse the stochastic comparison of interacting particle systems allowing for multiple arrivals, departures and non-conservative jumps of individuals between sites. That is, if $k$ individuals leave site $x$ for site $y$, a possibly different number $l$ arrive at destination. This setting includes new models, when compared to the conservative case, such as metapopulation models with deaths during migrations. It implies a sharp increase of technical complexity, given the numerous changes to consider. Known results are significantly generalised, even in the conservative case, as no particular form of the transition rates is assumed. We obtain necessary and sufficient conditions on the rates for the stochastic comparison of the processes and prove their equivalence with the existence of an order-preserving Markovian coupling. As a corollary, we get necessary and sufficient conditions for the attractiveness of the processes. A salient feature of our approach lies in the presentation of the coupling in terms of solutions to network flow problems. We illustrate the applicability of our results to a flexible family of population models described as interacting particle systems, with a range of parameters controlling births, deaths, catastrophes or migrations. We provide explicit conditions on the parameters for the stochastic comparison and attractiveness of the models, showing their usefulness in studying their limit behaviour. Additionally, we give three examples of constructing the coupling.

math.PR

Exact and asymptotic properties of $δ$-records in the linear drift model

The study of records in the Linear Drift Model (LDM) has attracted much attention recently due to applications in several fields. In the present paper we study $δ$-records in the LDM, defined as observations which are greater than all previous observations, plus a fixed real quantity $δ$. We give analytical properties of the probability of $δ$-records and study the correlation between $δ$-record events. We also analyse the asymptotic behaviour of the number of $δ$-records among the first $n$ observations and give conditions for convergence to the Gaussian distribution. As a consequence of our results, we solve a conjecture posed in J. Stat. Mech. 2010, P10013, regarding the total number of records in a LDM with negative drift. Examples of application to particular distributions, such as Gumbel or Pareto are also provided. We illustrate our results with a real data set of summer temperatures in Spain, where the LDM is consistent with the global-warming phenomenon.

math.ST

Minimax convergence rate for estimating the Wasserstein barycenter of random measures on the real line

This paper is focused on the statistical analysis of probability measures $ν_{1},\ldots,ν_{n}$ on $\mathbb{R}$ that can be viewed as independent realizations of an underlying stochastic process. We consider the situation of practical importance where the random measures $ν_{i}$ are absolutely continuous with densities $f_{i}$ that are not directly observable. In this case, instead of the densities, we have access to datasets of real random variables $(X_{i,j})_{1 \leq i \leq n; \; 1 \leq j \leq p_{i} }$ organized in the form of $n$ experimental units, such that $X_{i,1},\ldots,X_{i,p_{i}}$ are iid observations sampled from a random measure $ν_{i}$ for each $1 \leq i \leq n$. In this setting, we focus on first-order statistics methods for estimating, from such data, a meaningful structural mean measure. For the purpose of taking into account phase and amplitude variations in the observations, we argue that the notion of Wasserstein barycenter is a relevant tool. The main contribution of this paper is to characterize the rate of convergence of a (possibly smoothed) empirical Wasserstein barycenter towards its population counterpart in the asymptotic setting where both $n$ and $\min_{1 \leq i \leq n} p_{i}$ may go to infinity. The optimality of this procedure is discussed from the minimax point of view with respect to the Wasserstein metric. We also highlight the connection between our approach and the curve registration problem in statistics. Some numerical experiments are used to illustrate the results of the paper on the convergence rate of empirical Wasserstein barycenters.

math.ST

Geodesic PCA in the Wasserstein space

We introduce the method of Geodesic Principal Component Analysis (GPCA) on the space of probability measures on the line, with finite second moment, endowed with the Wasserstein metric. We discuss the advantages of this approach, over a standard functional PCA of probability densities in the Hilbert space of square-integrable functions. We establish the consistency of the method by showing that the empirical GPCA converges to its population counterpart, as the sample size tends to infinity. A key property in the study of GPCA is the isometry between the Wasserstein space and a closed convex subset of the space of square-integrable functions, with respect to an appropriate measure. Therefore, we consider the general problem of PCA in a closed convex subset of a separable Hilbert space, which serves as basis for the analysis of GPCA and also has interest in its own right. We provide illustrative examples on simple statistical models, to show the benefits of this approach for data analysis. The method is also applied to a real dataset of population pyramids.

stat.ME

Extrapolation of Urn Models via Poissonization: Accurate Measurements of the Microbial Unknown

The availability of high-throughput parallel methods for sequencing microbial communities is increasing our knowledge of the microbial world at an unprecedented rate. Though most attention has focused on determining lower-bounds on the alpha-diversity i.e. the total number of different species present in the environment, tight bounds on this quantity may be highly uncertain because a small fraction of the environment could be composed of a vast number of different species. To better assess what remains unknown, we propose instead to predict the fraction of the environment that belongs to unsampled classes. Modeling samples as draws with replacement of colored balls from an urn with an unknown composition, and under the sole assumption that there are still undiscovered species, we show that conditionally unbiased predictors and exact prediction intervals (of constant length in logarithmic scale) are possible for the fraction of the environment that belongs to unsampled classes. Our predictions are based on a Poissonization argument, which we have implemented in what we call the Embedding algorithm. In fixed i.e. non-randomized sample sizes, the algorithm leads to very accurate predictions on a sub-sample of the original sample. We quantify the effect of fixed sample sizes on our prediction intervals and test our methods and others found in the literature against simulated environments, which we devise taking into account datasets from a human-gut and -hand microbiota. Our methodology applies to any dataset that can be conceptualized as a sample with replacement from an urn. In particular, it could be applied, for example, to quantify the proportion of all the unseen solutions to a binding site problem in a random RNA pool, or to reassess the surveillance of a certain terrorist group, predicting the conditional probability that it deploys a new tactic in a next attack.

stat.ME

Asymptotic normality for the counting process of weak records and δ-records in discrete models

Let $\{X_n,n\ge1\}$ be a sequence of independent and identically distributed random variables, taking non-negative integer values, and call $X_n$ a $δ$-record if $X_n>\max\{X_1,...,X_{n-1}\}+δ$, where $δ$ is an integer constant. We use martingale arguments to show that the counting process of $δ$-records among the first $n$ observations, suitably centered and scaled, is asymptotically normally distributed for $δ\ne0$. In particular, taking $δ=-1$ we obtain a central limit theorem for the number of weak records.

math.PR