SearcharxivSearch

arXiv subjects

Bernard Bercu

Publications and source records attributed to Bernard Bercu.

At least 19 recordsLinked to original sources

A Central Limit Theorem for the Ewens-Pitman random partition in the large-$\theta$ regime via a martingale approach

The Ewens-Pitman model defines a distribution on random partitions of $\{1,\ldots,n\}$, with parameters $\alpha \in [0,1)$ and $\theta > -\alpha$; the case $\alpha=0$ reduces to the classical Ewens model from population genetics. We investigate the large-$n$ asymptotic behaviour of the Ewens-Pitman random partition in the nonstandard regime $\theta=\lambda n$ with $\lambda>0$, establishing joint fluctuation results for the total number of blocks $K_n^{\{n\}}$ and the counts $K_{r,n}^{\{n\}}$ of blocks of sizes $r=1,\dots,d$, for fixed $d\in\mathbb{N}$. In particular, for $\alpha\in[0,1)$ and $\theta=\lambda n$, our main result provides a strong law of large numbers and a central limit theorem for the $(d+1)$-dimensional vector $\mathbf{K}_{d,n}^{\{n\}} = \bigl(K_n^{\{n\}}, K_{1,n}^{\{n\}}, \dots, K_{d,n}^{\{n\}}\bigr)^T$ as $n \to \infty$. The proof exploits the Chinese restaurant sequential construction under $\theta=\lambda n$ and a central limit theorem for triangular arrays of martingales, extending techniques previously developed for the classical regime with fixed $\theta$. As corollaries of our results, we recover known asymptotics for $K_n^{\{n\}}$ and derive new strong laws and central limit theorems for each fixed $K_{r,n}^{\{n\}}$, thereby completing earlier weak-law results and providing a comprehensive asymptotic description of the Ewens-Pitman partition structure in the large-$\theta$ setting.

math.PR

A Gaussian process limit for the self-normalized Ewens-Pitman process

For an integer $n\geq1$, consider a random partition $\Pi_{n}$ of $\{1,\ldots,n\}$ into $K_{n}$ partition sets with $K_{r,n}$ partition subsets of size $r=1,\ldots,n$, and assume $\Pi_{n}$ distributed according to the Ewens-Pitman model with parameters $\alpha\in]0,1[$ and $\theta>-\alpha$. Although the large-$n$ asymptotic behaviors of $K_{n}$ and $K_{r,n}$ are well understood in terms of almost sure convergence and Gaussian fluctuations, much less is known about the asymptotic behavior of $P_{r,n}=K_{r,n}/K_n$ and of the self-normalized Ewens-Pitman process $(P_{1,n},P_{2,n},\dots)$. Motivated by the almost sure convergence of $(P_{1,n},P_{2,n},\dots)$ to the Sibuya distribution $p_{\alpha}=(p_{\alpha}(1),p_{\alpha}(2),\ldots)$, where $p_{\alpha}(r)$ is the probability mass at $r=1,2,\ldots$, we establish the $\ell^{2}$ distributional convergence \begin{displaymath} \sqrt{K_{n}}((P_{1,n},\,P_{2,n},\ldots)-p_{\alpha})\underset{n\rightarrow+\infty}{\overset{\cL}{\longrightarrow}}\mathcal{G}(\Gamma_\alpha), \end{displaymath} where $\mathcal{G}(\Gamma_\alpha)$ stands for a centered Gaussian process with covariance matrix $\Gamma_\alpha=diag(p_{\alpha}) - p_{\alpha} p_{\alpha}^T$. We apply our result to the estimation of the parameter

math.PR

An hybrid stochastic Newton algorithm for logistic regression

In this paper, we investigate a second-order stochastic algorithm for solving large-scale binary classification problems. We propose to make use of a new hybrid stochastic Newton algorithm that includes two weighted components in the Hessian matrix estimation: the first one coming from the natural Hessian estimate and the second associated with the stochastic gradient information. Our motivation comes from the fact that both parts evaluated at the true parameter of logistic regression, are equal to the Hessian matrix. This new formulation has several advantages and it enables us to prove the almost sure convergence of our stochastic algorithm to the true parameter. Moreover, we significantly improve the almost sure rate of convergence to the Hessian matrix. Furthermore, we establish the central limit theorem for our hybrid stochastic Newton algorithm. Finally, we show a surprising result on the almost sure convergence of the cumulative excess risk.

stat.CO

A new look on large deviations and concentration inequalities for the Ewens-Pitman model

The Ewens-Pitman model is a probability distribution for random partitions of the set $[n]=\{1,\ldots,n\}$, parameterized by $\alpha\in[0,1)$ and $\theta>-\alpha$, with $\alpha=0$ corresponding to the Ewens model in population genetics. The goal of this paper is to provide an alternative and concise proof of the Feng-Hoppe large deviation principle for the number $K_{n}$ of partition sets in the Ewens-Pitman model with $\alpha\in(0,1)$ and $\theta>-\alpha$. Our approach leverages an integral representation of the moment-generating function of $K_{n}$ in terms of the (one-parameter) Mittag-Leffler function, along with a sharp asymptotic expansion of it. This approach significantly simplifies the original proof of Feng-Hoppe large deviation principle, as it avoids all the technical difficulties arising from a continuity argument with respect to rational and non-rational values of $\alpha$. Beyond large deviations for $K_{n}$, our approach allows to establish a sharp concentration inequality for $K_n$ involving the rate function of the large deviation principle.

math.PR

On the multidimensional elephant random walk with stops

The goal of this paper is to investigate the asymptotic behavior of the multidimensional elephant random walk with stops (MERWS). In contrast with the standard elephant random walk, the elephant is allowed to stay on his own position. We prove that the Gram matrix associated with the MERWS, properly normalized, converges almost surely to the product of a deterministic matrix, related to the axes on which the MERWS moves uniformly, and a Mittag-Leffler distribution. It allows us to extend all the results previously established for the one-dimensional elephant random walk with stops. More precisely, in the diffusive and critical regimes, we prove the almost sure convergence of the MERWS. In the superdiffusive regime, we establish the almost sure convergence of the MERWS, properly normalized, to a nondegenerate random vector. We also study the self-normalized asymptotic normality of the MERWS.

math.PR

On the SAGA algorithm with decreasing step

Stochastic optimization naturally appear in many application areas, including machine learning. Our goal is to go further in the analysis of the Stochastic Average Gradient Accelerated (SAGA) algorithm. To achieve this, we introduce a new $\lambda$-SAGA algorithm which interpolates between the Stochastic Gradient Descent ($\lambda=0$) and the SAGA algorithm ($\lambda=1$). Firstly, we investigate the almost sure convergence of this new algorithm with decreasing step which allows us to avoid the restrictive strong convexity and Lipschitz gradient hypotheses associated to the objective function. Secondly, we establish a central limit theorem for the $\lambda$-SAGA algorithm. Finally, we provide the non-asymptotic $\mathbb{L}^p$ rates of convergence.

math.OC

Regularized estimation of Monge-Kantorovich quantiles for spherical data

Tools from optimal transport (OT) theory have recently been used to define a notion of quantile function for directional data. In practice, regularization is mandatory for applications that require out-of-sample estimates. To this end, we introduce a regularized estimator built from entropic optimal transport, by extending the definition of the entropic map to the spherical setting. We propose a stochastic algorithm to directly solve a continuous OT problem between the uniform distribution and a target distribution, by expanding Kantorovich potentials in the basis of spherical harmonics. In addition, we define the directional Monge-Kantorovich depth, a companion concept for OT-based quantiles. We show that it benefits from desirable properties related to Liu-Zuo-Serfling axioms for the statistical analysis of directional data. Building on our regularized estimators, we illustrate the benefits of our methodology for data analysis.

stat.ME

Sharp analysis on the joint distribution of the number of descents and inverse descents in a random permutation

Chatteerjee and Diaconis have recently shown the asymptotic normality for the joint distribution of the number of descents and inverse descents in a random permutation. A noteworthy point of their results is that the asymptotic variance of the normal distribution is diagonal, which means that the number of descents and inverse descents are asymptotically uncorrelated.The goal of this paper is to go further in this analysis by proving a large deviation principlefor the joint distribution. We shall show that the rate function of the joint distributionis the sum of the rate functions of the marginal distributions, which also means that the number of descents and inverse descents are asymptotically independent at the large deviation level. However,we are going to prove that they are finely dependent at the sharp large deviation level.

math.CO

A martingale approach to Gaussian fluctuations and laws of iterated logarithm for Ewens-Pitman model

The Ewens-Pitman model refers to a distribution for random partitions of $[n]=\{1,\ldots,n\}$, which is indexed by a pair of parameters $\alpha \in [0,1)$ and $\theta>-\alpha$, with $\alpha=0$ corresponding to the Ewens model in population genetics. The large $n$ asymptotic properties of the Ewens-Pitman model have been the subject of numerous studies, with the focus being on the number $K_{n}$ of partition sets and the number $K_{r,n}$ of partition subsets of size $r$, for $r=1,\ldots,n$. While for $\alpha=0$ asymptotic results have been obtained in terms of almost-sure convergence and Gaussian fluctuations, for $\alpha\in(0,1)$ only almost-sure convergences are available, with the proof for $K_{r,n}$ being given only as a sketch. In this paper, we make use of martingales to develop a unified and comprehensive treatment of the large $n$ asymptotic behaviours of $K_{n}$ and $K_{r,n}$ for $\alpha\in(0,1)$, providing alternative, and rigorous, proofs of the almost-sure convergences of $K_{n}$ and $K_{r,n}$, and covering the gap of Gaussian fluctuations. We also obtain new laws of the iterated logarithm for $K_{n}$ and $K_{r,n}$.

math.PR

Monge-Kantorovich superquantiles and expected shortfalls with applications to multivariate risk measurements

We propose center-outward superquantile and expected shortfall functions, with applications to multivariate risk measurements, extending the standard notion of value at risk and conditional value at risk from the real line to $\mathbb{R}^d$. Our new concepts are built upon the recent definition of Monge-Kantorovich quantiles based on the theory of optimal transport, and they provide a natural way to characterize multivariate tail probabilities and central areas of point clouds. They preserve the univariate interpretation of a typical observation that lies beyond or ahead a quantile, but in a meaningful multivariate way. We show that they characterize random vectors and their convergence in distribution, which underlines their importance. Our new concepts are illustrated on both simulated and real datasets.

math.ST

Stochastic optimal transport in Banach Spaces for regularized estimation of multivariate quantiles

We introduce a new stochastic algorithm for solving entropic optimal transport (EOT) between two absolutely continuous probability measures $\mu$ and $\nu$. Our work is motivated by the specific setting of Monge-Kantorovich quantiles where the source measure $\mu$ is either the uniform distribution on the unit hypercube or the spherical uniform distribution. Using the knowledge of the source measure, we propose to parametrize a Kantorovich dual potential by its Fourier coefficients. In this way, each iteration of our stochastic algorithm reduces to two Fourier transforms that enables us to make use of the Fast Fourier Transform (FFT) in order to implement a fast numerical method to solve EOT. We study the almost sure convergence of our stochastic algorithm that takes its values in an infinite-dimensional Banach space. Then, using numerical experiments, we illustrate the performances of our approach on the computation of regularized Monge-Kantorovich quantiles. In particular, we investigate the potential benefits of entropic regularization for the smooth estimation of multivariate quantiles using data sampled from the target measure $\nu$.

math.PR

Sharp large deviations and concentration inequalities for the number of descents in a random permutation

The goal of this paper is to go further in the analysis of the behavior of the number of descents in a random permutation. Via two different approaches relying on a suitable martingale decomposition or on the Irwin-Hall distribution, we prove that the number of descents satisfies a sharp large deviation principle. A very precise concentration inequality involving the rate function in the large deviation principle is also provided.

math.PR

On the elephant random walk with stops playing hide and seek with the Mittag-Leffler distribution

The aim of this paper is to investigate the asymptotic behavior of the so-called elephant random walk with stops (ERWS). In contrast with the standard elephant random walk, the elephant is allowed to be lazy by staying on his own position. We prove that the number of ones of the ERWS, properly normalized, converges almost surely to a Mittag-Leffler distribution. It allows us to carry out a sharp analysis on the asymptotic behavior of the ERWS. In the diffusive and critical regimes, we establish the almost sure convergence of the ERWS. We also show that it is necessary to self-normalized the position of the ERWS by the random number of ones in order to prove the asymptotic normality. In the superdiffusive regime, we establish the almost sure convergence of the ERWS, properly normalized, to a nondegenerate random variable. Moreover, we also show that the fluctuation of the ERWS around its limiting random variable is still Gaussian.

math.PR

A stochastic Gauss-Newton algorithm for regularized semi-discrete optimal transport

We introduce a new second order stochastic algorithm to estimate the entropically regularized optimal transport cost between two probability measures. The source measure can be arbitrary chosen, either absolutely continuous or discrete, while the target measure is assumed to be discrete. To solve the semi-dual formulation of such a regularized and semi-discrete optimal transportation problem, we propose to consider a stochastic Gauss-Newton algorithm that uses a sequence of data sampled from the source measure. This algorithm is shown to be adaptive to the geometry of the underlying convex optimization problem with no important hyperparameter to be accurately tuned. We establish the almost sure convergence and the asymptotic normality of various estimators of interest that are constructed from this stochastic Gauss-Newton algorithm. We also analyze their non-asymptotic rates of convergence for the expected quadratic risk in the absence of strong convexity of the underlying objective function. The results of numerical experiments from simulated data are also reported to illustrate the finite sample properties of this Gauss-Newton algorithm for stochastic regularized optimal transport, and to show its advantages over the use of the stochastic gradient descent, stochastic Newton and ADAM algorithms.

math.ST

How to estimate the memory of the Elephant Random Walk

We introduce an original way to estimate the memory parameter of the elephant random walk, a fascinating discrete time random walk on integers having a complete memory of its entire history. Our estimator is nothing more than a quasi-maximum likelihood estimator, based on a second order Taylor approximation of the log-likelihood function. We show the almost sure convergence of our estimate in the diffusive, critical and superdiffusive regimes. The local asymptotic normality of our statistical procedure is established in the diffusive regime, while the local asymptotic mixed normality is proven in the superdiffusive regime. Asymptotic and exact confidence intervals as well as statistical tests are also provided. All our analysis relies on asymptotic results for martingales and the quadratic variations associated.

math.PR

New insights on the minimal random walk

The aim of this paper is to deepen the analysis of the asymptotic behavior of the so-called minimal random walk (MRW) using a new martingale approach. The MRW is a discrete-time random walk with infinite memory that has three regimes depending on the location of its two parameters. In the diffusive and critical regimes, we establish new results on the almost sure asymptotic behavior of the MRW, such as the quadratic strong law and the law of the iterated logarithm. In the superdiffusive regime, we prove the almost sure convergence of the MRW, properly normalized, to a nondegenerate random variable. Moreover, we show that the fluctuation of the MRW around its limiting random variable is still Gaussian.

math.PR

Asymptotic analysis of random walks on ice and graphite

The purpose of this paper is to investigate the asymptotic behavior of random walks on three-dimensional crystal structures. We focus our attention on the 1h structure of the ice and the 2h structure of graphite. We establish the strong law of large numbers and the asymptotic normality for both random walks on ice and graphite. All our analysis relies on asymptotic results for multi-dimensional martingales.

math-ph