SearcharxivSearch

arXiv subjects

Florencia Leonardi

Publications and source records attributed to Florencia Leonardi.

At least 19 recordsLinked to original sources

Optimal recovery by maximum and integrated conditional likelihood in the general Stochastic Block Model

In this paper, we obtain new results on the weak and strong consistency of the maximum and integrated conditional likelihood estimators for the community detection problem in the Stochastic Block Model with $k$ communities and unknown parameters. In particular, we show that maximum conditional likelihood achieves the optimal known threshold for exact recovery in the logarithmic degree regime. For the integrated conditional likelihood, we obtain a sub-optimal constant, but still obtain strong consistency in the logarithmic degree regime. Both methods are shown to be weakly consistent in the divergent degree regime. These results fill in the gap in the theory of community detection with maximum likelihood and integrated conditional likelihood, solving open problems in the literature.

math.ST

Threshold detection under a semiparametric regression model

Linear regression models have been extensively considered in the literature. However, in some practical applications they may not be appropriate all over the range of the covariate. In this paper, a more flexible model is introduced by considering a regression model $Y=r(X)+\varepsilon$ where the regression function $r(\cdot)$ is assumed to be linear for large values in the domain of the predictor variable $X$. More precisely, we assume that $r(x)=α_0+β_0 x$ for $x> u_0$, where the value $u_0$ is identified as the smallest value satisfying such a property. A penalized procedure is introduced to estimate the threshold $u_0$. The considered proposal focusses on a semiparametric approach since no parametric model is assumed for the regression function for values smaller than $u_0$. Consistency properties of both the threshold estimator and the estimators of $(α_0,β_0)$ are derived, under mild assumptions. Through a numerical study, the small sample properties of the proposed procedure and the importance of introducing a penalization are investigated. The analysis of a real data set allows us to demonstrate the usefulness of the penalized estimators.

math.ST

Model selection for Markov random fields on graphs under a mixing condition

In this work, we propose a global model selection criterion to estimate the graph of conditional dependencies of a random vector based on a finite sample. By global criterion, we mean optimizing a function over the entire set of possible graphs, eliminating the need to estimate the individual neighborhoods and subsequently combine them to estimate the graph. We prove the almost sure convergence of the graph estimator. This convergence holds provided the data is a realization of a multivariate stochastic process that satisfies a mixing condition. To the best of our knowledge, these are the first results to show the consistency of a model selection criterion for Markov random fields on graphs under non-independent data.

math.ST

Consistent model selection for the Degree Corrected Stochastic Blockmodel

The Degree Corrected Stochastic Block Model (DCSBM) was introduced by \cite{karrer2011stochastic} as a generalization of the stochastic block model in which vertices of the same community are allowed to have distinct degree distributions. On the modelling side, this variability makes the DCSBM more suitable for real life complex networks. On the statistical side, it is more challenging due to the large number of parameters when dealing with community detection. In this paper we prove that the penalized marginal likelihood estimator is strongly consistent for the estimation of the number of communities. We consider \emph{dense} or \emph{semi-sparse} random networks, and our estimator is \emph{unbounded}, in the sense that the number of communities $k$ considered can be as big as $n$, the number of nodes in the network.

math.ST

Structure recovery for partially observed discrete Markov random fields on graphs under not necessarily positive distributions

We propose a penalized pseudo-likelihood criterion to estimate the graph of conditional dependencies in a discrete Markov random field that can be partially observed. We prove the convergence of the estimator in the case of a finite or countable infinite set of nodes. In the finite case, the underlying graph can be recovered with probability one, while in the countable infinite case, we can recover any finite sub-graph with probability one by allowing the candidate neighborhoods to grow as a function o(log n), with n the sample size. Our method requires minimal assumptions on the probability distribution, and contrary to other approaches in the literature, the usual positivity condition is not needed. We evaluate the performance of the estimator on simulated data, and we apply the methodology to a real dataset of stock index markets in different countries.

stat.ME

Population based change-point detection for the identification of homozygosity islands

In this paper, we propose a new method for offline change-point detection on some parameters of the distribution of a random vector. We introduce a penalized maximum likelihood approach that can be efficiently computed by a dynamic programming algorithm or approximated by a fast greedy binary splitting algorithm. We prove both algorithms converge almost surely to the set of change-points under very general assumptions on the distribution and independent sampling of the random vector. In particular, we show the assumptions leading to the consistency of the algorithms are satisfied by categorical and Gaussian random variables. This new approach is motivated by the problem of identifying homozygosity islands on the genome of individuals in a population. Our method directly tackles the issue of identification of the homozygosity islands at the population level, without the need of analyzing single individuals and then combining the results, as is made nowadays in state-of-the-art approaches.

stat.ME

Strong consistency of Krichevsky-Trofimov estimator for the number of communities in the Stochastic Block Model

In this paper we introduce an estimator for the number of communities in the Stochastic Block Model (SBM), based on the maximization of a penalized version of the so-called Krichevsky-Trofimov mixture distribution. We prove its eventual almost sure convergence to the underlying number of communities, without assuming a known upper bound on that quantity. Our results apply to both the dense and the sparse regimes. To our knowledge this is the first consistency result for the estimation of the number of communities in the SBM in the unbounded case, that is when the number of communities is allowed to grow with the same size.

math.ST

A note on perfect simulation for exponential random graph models

In this paper we propose a perfect simulation algorithm for the Exponential Random Graph Model, based on the Coupling From The Past method of Propp & Wilson (1996). We use a Glauber dynamics to construct the Markov Chain and we prove the monotonicity of the ERGM for a subset of the parametric space. We also obtain an upper bound on the running time of the algorithm that depends on the mixing time of the Markov chain.

stat.CO

Computationally efficient change point detection for high-dimensional regression

Large-scale sequential data is often exposed to some degree of inhomogeneity in the form of sudden changes in the parameters of the data-generating process. We consider the problem of detecting such structural changes in a high-dimensional regression setting. We propose a joint estimator of the number and the locations of the change points and of the parameters in the corresponding segments. The estimator can be computed using dynamic programming or, as we emphasize here, it can be approximated using a binary search algorithm with $O(n \log(n) \mathrm{Lasso}(n))$ computational operations while still enjoying essentially the same theoretical properties; here $\mathrm{Lasso}(n)$ denotes the computational cost of computing the Lasso for sample size $n$. We establish oracle inequalities for the estimator as well as for its binary search approximation, covering also the case with a large (asymptotically growing) number of change points. We evaluate the performance of the proposed estimation algorithms on simulated data and apply the methodology to real data.

stat.ME

Nonparametric statistical inference for the context tree of a stationary ergodic process

We consider the problem of estimating the context tree of a stationary ergodic process with finite alphabet without imposing additional conditions on the process. As a starting point we introduce a Hamming metric in the space of irreducible context trees and we use the properties of the weak topology in the space of ergodic stationary processes to prove that if the Hamming metric is unbounded, there exist no consistent estimators for the context tree. Even in the bounded case we show that there exist no two-sided confidence bounds. However we prove that one-sided inference is possible in this general setting and we construct a consistent estimator that is a lower bound for the context tree of the process with an explicit formula for the coverage probability. We develop an efficient algorithm to compute the lower bound and we apply the method to test a linguistic hypothesis about the context tree of codified written texts in European Portuguese.

math.ST

A test of hypotheses for random graph distributions built from EEG data

The theory of random graphs is being applied in recent years to model neural interactions in the brain. While the probabilistic properties of random graphs has been extensively studied in the literature, the development of statistical inference methods for this class of objects has received less attention. In this work we propose a non-parametric test of hypotheses to test if two samples of random graphs were originated from the same probability distribution. We show how to compute efficiently the test statistic and we study its performance on simulated data. We apply the test to compare graphs of brain functional network interactions built from electroencephalographic (EEG) data collected during the visualization of point light displays depicting human locomotion.

stat.AP

A model selection approach for multiple sequence segmentation and dimensionality reduction

In this paper we consider the problem of segmenting $n$ aligned random sequences of equal length $m$, into a finite number of independent blocks. We propose to use a penalized maximum likelihood criterion to infer simultaneously the number of points of independence as well as the position of each one of these points. We show how to compute the estimator efficiently by means of a dynamic programming algorithm with time complexity $O(m^2n)$. We also propose another algorithm, called hierarchical algorithm, that provides an approximation to the estimator when the sample size increases and runs in time $O(mn)$. Our main theoretical result is the proof of almost sure consistency of the estimator and the convergence of the hierarchical algorithm when the sample size $n$ grows to infinity. We illustrate the convergence of these algorithms through some simulation examples and we apply the method to a real protein sequence alignment of Ebola Virus.

stat.ME

Loss of memory of hidden Markov models and Lyapunov exponents

In this paper we prove that the asymptotic rate of exponential loss of memory of a finite state hidden Markov model is bounded above by the difference of the first two Lyapunov exponents of a certain product of matrices. We also show that this bound is in fact realized, namely for almost all realizations of the observed process we can find symbols where the asymptotic exponential rate of loss of memory attains the difference of the first two Lyapunov exponents. These results are derived in particular for the observed process and for the filter; that is, for the distribution of the hidden state conditioned on the observed sequence. We also prove similar results in total variation.

math.PR

Finding the basic neighborhood in variable range Markov random fields: application in SNP association studies

The SNPs (Single Nucleotide Polymorphisms) genotyping platforms are of great value for gene mapping of complex diseases. Nowadays, the high-density of these molecular markers enables studies of dependence patterns between loci over the genome, allowing a simultaneous inference of dependence structure and disease association. In this paper we propose a method based on the theory of variable range Markov random fields to estimate the extent of dependence among SNPs allowing variable windows along the genome. The advantage of this method is that it allows the simultaneous prediction of dependence and independence regions among SNPs, without restricting a priori the range of dependence. We introduce an estimator based on the idea of penalized maximum likelihood to find the conditional dependence neighborhood of each SNP in the sample and we prove its consistency. We apply our method to autosomal SNPs genotypic data with unknown phase in the context of case-control association studies. By examining rheumatoid arthritis data from the Genetic Analysis Workshop 16 (GAW16), we show the utility of the Markov model under variable range dependence.

stat.ME

Context tree selection and linguistic rhythm retrieval from written texts

The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet they are conjectured, on linguistic grounds, to implement different rhythms. We show that this linguistic question can be formulated as a problem of model selection in the class of variable length Markov chains. To carry on this approach, we compare texts from European and Brazilian Portuguese. These texts are previously encoded according to some basic rhythmic features of the sentences which can be automatically retrieved. This is an entirely new approach from the linguistic point of view. Our statistical contribution is the introduction of the smallest maximizer criterion which is a constant free procedure for model selection. As a by-product, this provides a solution for the problem of optimal choice of the penalty constant when using the BIC to select a variable length Markov chain. Besides proving the consistency of the smallest maximizer criterion when the sample size diverges, we also make a simulation study comparing our approach with both the standard BIC selection and the Peres-Shields order estimation. Applied to the linguistic sample constituted for our case study, the smallest maximizer criterion assigns different context-tree models to the two dialects of Portuguese. The features of the selected models are compatible with current conjectures discussed in the linguistic literature.

stat.ML

Context Tree Selection: A Unifying View

The present paper investigates non-asymptotic properties of two popular procedures of context tree (or Variable Length Markov Chains) estimation: Rissanen's algorithm Context and the Penalized Maximum Likelihood criterion. First showing how they are related, we prove finite horizon bounds for the probability of over- and under-estimation. Concerning overestimation, no boundedness or loss-of-memory conditions are required: the proof relies on new deviation inequalities for empirical probabilities of independent interest. The underestimation properties rely on loss-of-memory and separation conditions of the process. These results improve and generalize the bounds obtained previously. Context tree models have been introduced by Rissanen as a parsimonious generalization of Markov models. Since then, they have been widely used in applied probability and statistics.

math.ST

Testing statistical hypothesis on random trees and applications to the protein classification problem

Efficient automatic protein classification is of central importance in genomic annotation. As an independent way to check the reliability of the classification, we propose a statistical approach to test if two sets of protein domain sequences coming from two families of the Pfam database are significantly different. We model protein sequences as realizations of Variable Length Markov Chains (VLMC) and we use the context trees as a signature of each protein family. Our approach is based on a Kolmogorov--Smirnov-type goodness-of-fit test proposed by Balding et al. [Limit theorems for sequences of random trees (2008), DOI: 10.1007/s11749-008-0092-z]. The test statistic is a supremum over the space of trees of a function of the two samples; its computation grows, in principle, exponentially fast with the maximal number of nodes of the potential trees. We show how to transform this problem into a max-flow over a related graph which can be solved using a Ford--Fulkerson algorithm in polynomial time on that number. We apply the test to 10 randomly chosen protein domain families from the seed of Pfam-A database (high quality, manually curated families). The test shows that the distributions of context trees coming from different families are significantly different. We emphasize that this is a novel mathematical approach to validate the automatic clustering of sequences in any context. We also study the performance of the test via simulations on Galton--Watson related processes.

math.ST

Some upper bounds for the rate of convergence of penalized likelihood context tree estimators

We find upper bounds for the probability of underestimation and overestimation errors in penalized likelihood context tree estimation. The bounds are explicit and applies to processes of not necessarily finite memory. We allow for general penalizing terms and we give conditions over the maximal depth of the estimated trees in order to get strongly consistent estimates. This generalizes previous results obtained in the case of estimation of the order of a Markov chain.

math.ST