SearcharxivSearch

arXiv subjects

Soumaila Dembele

Publications and source records attributed to Soumaila Dembele.

5 recordsLinked to original sources

Second order Expansions for Extreme Quantiles of Burr Distributions and Asymptotic Theory of Record Values

In this paper we investigate the Burr distributions family which contains twelve members. Second order expansions of quantiles of the Burr's distributions are provided on which may be based statistical methods, in particular in extreme value theory. Beyond the proper interest of these expansions, we apply them to characterize the asymptotic laws of their records of Burr's distributions, lead to new statistical tests.

math.ST

Elements of Randoms Analysis about the Gamma Generalized Hyperbolic Distribution Levy Stochastic Process

In this paper, we study some aspects on random analysis on the Léevy stochastic processes with margins following generalized hyperbolic distributions generated by gamma laws. In particular we study the boundedness of its total variations and the quadratic variations. Next we give an empirical construction that enables the graphical representation of the paths of such stochastic processes. Comparisons with the Brownian motions are considered.

math.PR

The exact probability law for the approximated similarity from the Minhashing method

We propose a probabilistic setting in which we study the probability law of the Rajaraman and Ullman \textit{RU} algorithm and a modified version of it denoted by \textit{RUM}. These algorithms aim at estimating the similarity index between huge texts in the context of the web. We give a foundation of this method by showing, in the ideal case of carefully chosen probability laws, the exact similarity is the mathematical expectation of the random similarity provided by the algorithm. Some extensions are given. \noindent \textbf{Résumé.} Nous proposons un cadre probabilistique dans lequel nous étudions la loi de probabilité de l'algorithme de Rajaraman et Ullman \textit{RU} ainsi qu'une version modifiée de cet algorithme notée \textit{RUM}. Ces alogrithmes visent à estimer l'indice de la similarité entre des textes de grandes tailles dans le contexte du Web. Nous donnons une base de validité de cette méthode en montrant que pour des lois de probabilités minutieusement choisies, la similarité exacte est l'espérance mathématique de la similarité aléatoire donnée par l'algorithme \textit{RUM}. Des généralisations sont abordées.

math.PR

Applying of the Extreme Value Theory for determining extreme claims in the automobile insurance sector: Case of a China car insurance

According to the Chinese Health Statistics Yearbook, in 2005, the number of traffic accidents was 187781 with total direct property losses of 103691.7 (10000 Yuan). This research aims to fill the gap in the literature by investigating the extreme claim sizes not only for the entire portfolio. This empirical study investigates the behavior of the upper tail of the claim size by class of policyholders.

stat.AP

Probabilistic, statistical and algorithmic aspects of the similarity of texts and application to Gospels comparison

The fundamental problem of similarity studies, in the frame of data-mining, is to examine and detect similar items in articles, papers, books, with huge sizes. In this paper, we are interested in the probabilistic, and the statistical and the algorithmic aspects in studies of texts. We will be using the approach of $k$\textit{-shinglings}, a $k$\textit{-shingling} being defined as a sequence of $k$ consecutive characters that are extracted from a text ($k\geq 1$ ). The main stake in this field is to find accurate and quick algorithms to compute the similarity in short times. This will be achieved in using approximation methods. The first approximation method is statistical and, is based on the theorem of Glivenko-Cantelli. The second is the banding technique. And the third concerns a modification of the algorithm proposed by Rajaraman and al (% \cite{AnandJeffrey}), denoted here as (RUM). The Jaccard index is the one used in this paper. We finally illustrate these results of the paper on the four Gospels. The results are very conclusive.

stat.ME