SearcharxivSearch

arXiv subjects

Gane Samb Lo

Publications and source records attributed to Gane Samb Lo.

At least 19 recordsLinked to original sources

Independence of the indicator functions of record values for Multivariate independent data

We consider a sequence of random vectors on \(\mathbb{R}^d, \ d\geq 1\). We consider the record values based on the simultaneous strict inequality of the coordinates. The indicator record variable (irv) of the j-th observation is the function that assigns the value 1 (one) if that observation is a record value and the null value otherwise. Here, we give a detailed and a thorough proof that the indicator functions are independent, whenever the data are themselves independent, not necessarily iid , in \(\mathbb{R}^d, \ d\geq 1\). We compare that proof with available proofs in dimension one. Indeed, in seminal works on records, in particular in Ahsanullah(2024), Nevzorov(2001), Resnick (1987), Ahsanullah and Nevzorov (2015), etc., the independence of record indicator functions is usually validated based on logical reasoning, and so, is not rigorously proved. This allows us to undertake a detailed and thorough proof for independent data, not necessity iid , in \(\mathbb{R}^d, \ d\geq 1\). The proof includes leads for not even independent data.

math.PR

Asymptotic theory and statistical inference for the samples problems with heavy-tailed data using the functional empirical process

This paper introduces the Trimmed Functional Empirical Process (TFEP) as a robust framework for statistical inference when dealing with heavy-tailed or skewed distributions, where classical moments such as the mean or variance may be infinite or undefined. Standard approaches including the classical Functional Empirical Process (FEP), break down under such conditions, especially for distributions like Pareto, Cauchy, low degree of freedom Student-t, due to their reliance on finite-variance assumptions to guarantee asymptotic convergence. The TFEP approach addresses these limitations by trimming a controlled proportion of extreme order statistics, thereby stabilizing the empirical process and restoring asymptotic Gaussian behavior. We establish the weak convergence of the TFEP under mild regularity conditions and derive new asymptotic distributions for one-sample and twosample problems. These theoretical developments lead to robust confidence intervals for truncated means, variances, and their differences or ratios. The efficiency and reliability of the TFEP are supported by extensive Monte Carlo experiments and an empirical application to Senegalese income data. In all scenarios, the TFEP provides accurate inference where both Gaussian-based methods and the classical FEP break down. The methodology thus offers a powerful and flexible tool for statistical analysis in heavy-tailed and non-standard environments.

stat.ME

Asymptotic Statistical Theory for the Samples Problems using the Functional Empirical Process, revisited I

In this paper we study the asymptotic theory for samples problem based on the functional empirical process (fep), this new method is called general samples problem. We suggest this method to develop the full theory of estimation of means, variances, ratios of variances and difference of means for independent samples. We compare the results of our new method to the Gaussian method using simulated and real data. The obtained results are almost equivalent to those in the Gaussian case for samples's size equal to $10$. It has been prove that the estimation of the means difference is very precise regardless of the equality or inequality of variances for greater sizes of sample. This method is recommended when the sizes of samples is around or greater that $15$ and it requires the finiteness of the fourth order moment.

stat.ME

General asymptotic representations of indexes based on the functional empirical process and the residual functional empirical process and applications

The objective of this paper is to establish a general asymptotic representation (\textit{GAR}) for a wide range of statistics, employing two fundamental processes: the functional empirical process (\textit{fep}) and the residual functional empirical process introduced by Lo and Sall (2010a, 2010b), denoted as \textit{lrfep}. The functional empirical process (\textit{fep}) is defined as follows: $$ \mathbb{G}_n(h)=\frac{1}{\sqrt{n}} \sum_{j=1}^{n} \{h(X_j)-\mathbb{E}h(X_j)\}, $$ \Bin [where $X$, $X_1$, $\cdots$, $X_n$ is a sample from a random $d$-vectors $X$ of size $(n+1)$ with $n\geq 1$ and $h$ is a measurable function defined on $\mathbb{R}^d$ such that $\mathbb{E}h(X)^2<+\infty$]. It is a powerful tool for deriving asymptotic laws. An earlier and simpler version of this paper focused on the application of the (\textit{fep}) to statistics $J_n$ that can be turned into an asymptotic algebraic expression of empirical functions of the form $$ J_n=\mathbb{E}h(X) + n^{-1/2} \mathbb{G}_n(h) + o_{\mathbb{P}}(n^{-1/2}). \ \ \ \textit{SGAR} $$ \Bin However, not all statistics, in particular welfare indexes, conform to this form. In many scenarios, functions of the order statistics $X_{1,n}\leq$, $\cdots$, $\leq X_{n,n}$ are involved, resulting in $L$-statistics. In such cases, the (\textit{fep}) can still be utilized, but in combination with the related residual functional empirical process introduced by Lo and Sall (2010a, 2010b). This combination leads to general asymptotic representations (GAR) for a wide range of statistical indexes $$ J_n=\mathbb{E}h(X) + n^{-1/2} \biggr(\mathbb{G}_n(h) + \int_{0}^{1} \mathbb{G}_n(\tilde{f}_s) \ell(s) \ ds + o_{\mathbb{P}}(1)\biggr), \ \ \textit{FGAR} $$

math.ST

A Jarque-Bera test for skew normal data

The skew normal law has been introduced in Azzalin (1985) as an alternative to adjusting asymmetric data that share important patterns with the normal law. It has been extensively studied. However, there is so much to do in order to catch the diversity and the richness of the investigation of its normal counterpart. The General Jarque-Berra Test (GJBT) has been devised by Lo et al. (2015), Da et al. (2023) for arbitrary laws with at least finite first eight moments, as a generalization of the Jarque-Bera (1987) test that was specially set up for normal data. Here, we particularize it to skew normal data. When particularized in the skew normal law, this test is proven to be extremely powerful in detecting the true model for any $α\neq 0$ and rejected the normal law ($α=0$) whatever be the size of the data. We introduced the use of the samples duplication method to reach a high level of efficiency for the test.

stat.ME

The Real-Valued Bochner integral and the Modern Real-Valued Measurable function on $\mathbb{R}$

The like-Lebesgue integral of real-valued measurable functions (abbreviated as \textit{RVM-MI})is the most complete and appropriate integration Theory. Integrals are also defined in abstract spaces since Pettis (1938). In particular, Bochner integrals received much interest with very recent researches. It is very commode to use the \textit{RVM-MI} in constructing Bochner integral in Banach or in locally convex spaces. In this simple not, we prove that the Bochner integral and the \textit{RVM-MI} with respect to a finite measure $m$ are the same on $\mathbb{R}$. Applications of that equality may be useful in weak limits on Banach space.

math.FA

Second order Expansions for Extreme Quantiles of Burr Distributions and Asymptotic Theory of Record Values

In this paper we investigate the Burr distributions family which contains twelve members. Second order expansions of quantiles of the Burr's distributions are provided on which may be based statistical methods, in particular in extreme value theory. Beyond the proper interest of these expansions, we apply them to characterize the asymptotic laws of their records of Burr's distributions, lead to new statistical tests.

math.ST

Elements of Randoms Analysis about the Gamma Generalized Hyperbolic Distribution Levy Stochastic Process

In this paper, we study some aspects on random analysis on the Léevy stochastic processes with margins following generalized hyperbolic distributions generated by gamma laws. In particular we study the boundedness of its total variations and the quadratic variations. Next we give an empirical construction that enables the graphical representation of the paths of such stochastic processes. Comparisons with the Brownian motions are considered.

math.PR

An introduction to a general records theory both for dependent data and high dimensions: revisited 2

The probabilistic investigation on record values and record times of a sequence of random variables defined on the same probability space has received much attention from 1952 to now. A great deal of such theory focused on \textit{iid} or independent real-valued random variables. There exists a few results for real-valued dependent random variables. Some papers deal also with multivariate random variables. But a large theory regarding vectors and dependent data has yet to be done. In preparation of that, the probability laws of records are investigated here, without any assumption on the dependence structure. The results are extended sequences with values in partially ordered spaces whose order is compatible with measurability. The general characterizations are checked in known cases mostly for \textit{iid} sequences. The frame is ready for undertaking a vast study of records theory in high dimensions and for types of dependence.

math.PR

The exact probability law for the approximated similarity from the Minhashing method

We propose a probabilistic setting in which we study the probability law of the Rajaraman and Ullman \textit{RU} algorithm and a modified version of it denoted by \textit{RUM}. These algorithms aim at estimating the similarity index between huge texts in the context of the web. We give a foundation of this method by showing, in the ideal case of carefully chosen probability laws, the exact similarity is the mathematical expectation of the random similarity provided by the algorithm. Some extensions are given. \noindent \textbf{Résumé.} Nous proposons un cadre probabilistique dans lequel nous étudions la loi de probabilité de l'algorithme de Rajaraman et Ullman \textit{RU} ainsi qu'une version modifiée de cet algorithme notée \textit{RUM}. Ces alogrithmes visent à estimer l'indice de la similarité entre des textes de grandes tailles dans le contexte du Web. Nous donnons une base de validité de cette méthode en montrant que pour des lois de probabilités minutieusement choisies, la similarité exacte est l'espérance mathématique de la similarité aléatoire donnée par l'algorithme \textit{RUM}. Des généralisations sont abordées.

math.PR

Applying of the Extreme Value Theory for determining extreme claims in the automobile insurance sector: Case of a China car insurance

According to the Chinese Health Statistics Yearbook, in 2005, the number of traffic accidents was 187781 with total direct property losses of 103691.7 (10000 Yuan). This research aims to fill the gap in the literature by investigating the extreme claim sizes not only for the entire portfolio. This empirical study investigates the behavior of the upper tail of the claim size by class of policyholders.

stat.AP

$\mathbf{G}$-Central limit theorems and $\mathbf{G}$-invariance principles for associated random variables

The investigation asymptotic limits on associated data mainly focused on limit theorems of summands of associated data and on the related invariance principles. In a series of papers, we are going to set the general frame of the theory by considering an arbitrary infinitely decomposable (divisible) limit law for summands and study the associated functional laws converging to Lévy processes. The asymptotic frame of Newman (1980) is still used as a main tool. Detailed results are given when $G$ is a Gaussian law (as confirmation of known results) and when $G$ is a Poisson law. In the later case, classical results for independent and identically distributed data are extended to stationary and non-stationary associated data.

math.PR

$\ell^{\infty}$ Poisson invariance principles from two classical Poisson limit theorems and extension to non-stationary independent sequences

The simple Lévy Poisson process and scaled forms are explicitly constructed from partial sums of independent and identically distributed random variables and from sums of non-stationary independent random variables. For the latter, the weak limits are scaled Poisson processes. The method proposed here prepares generalizations to dependent data, to associated data in the first place.

math.PR

Extensions of two classical Poisson limit laws to non-stationary independent data

In earlier stages in the introduction to asymptotic methods in probability theory, the weak convergence of sequences $(X_n)_{n\geq 1}$ of Binomial of random variables (\textit{rv}'s) to a Poisson law is classical and easy-to prove. A version of such a result concerning sequences $(Y_n)_{n\geq 1}$ of negative binomial \textit{rv}'s also exists. In both cases, $X_n$ and $Y_n-n$ are by-row sums $S_n[X]$ and $S_n[Y]$ of arrays of Bernoulli \textit{rv}'s and corrected geometric \textit{rv}'s respectively. When considered in the general frame of asymptotic theorems of by-row sums of \textit{rv}'s of arrays, these two simple results in the independent and identically distributed scheme can be generalized to non-stationary data and beyond to non-stationary and dependent data. Further generalizations give interesting results that would not be found by direct methods. In this paper, we focus on generalizations to the non-stationary independent data. Extensions to dependent data will addressed later.

math.PR

The Pseudo-Lindley Alpha Power transformed distribution, mathematical characterizations and asymptotic properties

We introduce a new generalization of the Pseudo-Lindley distribution by applying alpha power transformation. The obtained distribution is referred as the Pseudo-Lindley alpha power transformed distribution (\textit{PL-APT}). Some tractable mathematical properties of the \textit{PL-APT} distribution as reliability, hazard rate, order statistics and entropies are provided. The maximum likelihood method is used to obtain the parameters' estimation of the \textit{PL-APT} distribution. The asymptotic properties of the proposed distribution are discussed. Also, a simulation study is performed to compare the modeling capability and flexibility of \textit{PL-APT} with Lindley and Pseudo-Lindley distributions. The \textit{PL-APT} provides a good fit as the Lindley and the Pseudo-Lindley distribution. The extremal domain of attraction of \textit{PL-APT} is found and its quantile and extremal quantile functions studied. Finally, the extremal value index is estimated by the double-indexed Hill's estimator (Ngom and Lo, 2016) and related asymptotic statistical tests are provided and characterized.

math.ST

Moments estimators and omnibus chi-square tests for some usual probability laws

For many probability laws, in parametric models, the estimation of the parameters can be done in the frame of the maximum likelihood method, or in the frame of moment estimation methods, or by using the plug-in method, etc. Usually, for estimating more than one parameter, the same frame is used. We focus on the moment estimation method in this paper. We use the instrumental tool of the functional empirical process (fep) in Lo (2016) to show how it is practical to derive, almost algebraically, the joint distribution Gaussian law and to derive omnibus chi-square asymptotic laws from it. We choose four distributions to illustrate the method (Gamma law, beta law, Uniform law and Fisher law) and completely describe the asymptotic laws of the moment estimators whenever possible. Simulations studies are performed to investigate for each case the smallest sizes for which the obtained statistical tests are recommendable. Generally, the omnibus chi-square test proposed here work fine with sample sizes around fifty

stat.ME

Simultaneous Joint Lower and Upper record values Probability Laws for Absolutely Continuous or Discrete Data

This paper investigates the probability density function ($pdf$) of the $(2n-1)$-vector $(n\geq 1)$ of both lower and upper record values for a sequence of independent random variables with common $pdf f$ defined on the same probability space, provided that the lower and upper record times are finite up to $n$. A lot is known about the lower or the upper record values when they are studied separately. When put together, the challenges are a far bigger complicated. The rare results in the literature still present important flaws. This paper begins a new and complete investigation with a few number of records: 2 and 3. Lessons from these simple cases will allow addressing the general formulation of simultaneous joint lower-upper records.

math.PR

A Course on Elementary Probability Theory

This book introduces to the theory of probabilities from the beginning. Assuming that the reader possesses the normal mathematical level acquired at the end of the secondary school, we aim to equip him with a solid basis in probability theory. The theory is preceded by a general chapter on counting methods. Then, the theory of probabilities is presented in a discrete framework. Two objectives are sought. The first is to give the reader the ability to solve a large number of problems related to probability theory, including application problems in a variety of disciplines. The second was to prepare the reader before he approached the manual on the mathematical foundations of probability theory. In this book, the reader will concentrate more on mathematical concepts, while in the present text, experimental frameworks are mostly found. If both objectives are met, the reader will have already acquired a definitive experience in problem-solving ability with the tools of probability theory and at the same time he is ready to move on to a theoretical course on probability theory based on the theory of measurement and integration. The book ends with a chapter that allows the reader to begin an intermediate course in mathematical statistics.

math.HO