SearcharxivSearch

arXiv subjects

Benjamin Frot

Publications and source records attributed to Benjamin Frot.

4 recordsLinked to original sources

RSVP-graphs: Fast High-dimensional Covariance Matrix Estimation under Latent Confounding

In this work we consider the problem of estimating a high-dimensional $p \times p$ covariance matrix $\Sigma$, given $n$ observations of confounded data with covariance $\Sigma + \Gamma \Gamma^T$, where $\Gamma$ is an unknown $p \times q$ matrix of latent factor loadings. We propose a simple and scalable estimator based on the projection on to the right singular vectors of the observed data matrix, which we call RSVP. Our theoretical analysis of this method reveals that in contrast to PCA-based approaches, RSVP is able to cope well with settings where the smallest eigenvalue of $\Gamma^T \Gamma$ is close to the largest eigenvalue of $\Sigma$, as well as settings where the eigenvalues of $\Gamma^T \Gamma$ are diverging fast. It is also able to handle data that may have heavy tails and only requires that the data has an elliptical distribution. RSVP does not require knowledge or estimation of the number of latent factors $q$, but only recovers $\Sigma$ up to an unknown positive scale factor. We argue this suffices in many applications, for example if an estimate of the correlation matrix is desired. We also show that by using subsampling, we can further improve the performance of the method. We demonstrate the favourable performance of RSVP through simulation experiments and an analysis of gene expression datasets collated by the GTEX consortium.

stat.ME

Robust causal structure learning with some hidden variables

We introduce a new method to estimate the Markov equivalence class of a directed acyclic graph (DAG) in the presence of hidden variables, in settings where the underlying DAG among the observed variables is sparse, and there are a few hidden variables that have a direct effect on many of the observed ones. Building on the so-called low rank plus sparse framework, we suggest a two-stage approach which first removes the effect of the hidden variables, and then estimates the Markov equivalence class of the underlying DAG under the assumption that there are no remaining hidden variables. This approach is consistent in certain high-dimensional regimes and performs favourably when compared to the state of the art, both in terms of graphical structure recovery and total causal effect estimation.

stat.ME

Latent variable model selection for Gaussian conditional random fields

We consider the problem of learning a conditional Gaussian graphical model in the presence of latent variables. Building on recent advances in this field, we suggest a method that decomposes the parameters of a conditional Markov random field into the sum of a sparse and a low-rank matrix. We derive convergence bounds for this estimator and show that it is well-behaved in the high-dimensional regime as well as "sparsistent" (i.e. capable of recovering the graph structure). We then show how proximal gradient algorithms and semi-definite programming techniques can be employed to fit the model to thousands of variables. Through extensive simulations, we illustrate the conditions required for identifiability and show that there is a wide range of situations in which this model performs significantly better than its counterparts, for example, by accommodating more latent variables. Finally, the suggested method is applied to two datasets comprising individual level data on genetic variants and metabolites levels. We show our results replicate better than alternative approaches and show enriched biological signal.

stat.ME

Godel's Completeness Theorem and Deligne's Theorem

These notes were written for a presentation given at the university Paris VII in January 2012. The goal was to explain a proof of a famous theorem by P. Deligne about coherent topoi (coherent topoi have enough points) and to show how this theorem is equivalent to G\"odel's completeness theorem for first order logic. Because it was not possible to cover everything in only three hours, the focus was on Barr's and Deligne's theorems. This explains why the corresponding sections have been given more attention. Section 1 and the Appendix were added in an attempt to make this document self-contained and understandable by a reader with a good knowledge of topos theory. A coherent topos is a topos which is equivalent to a Grothendieck topos that admits a site $(C,J)$ such that $C$ has finite limits and there exists a base of $J$ with finite covering families. In order to prove that any coherent topos has enough points, it is enough to show that for any coherent topos $\mathcal{E}$ there is a surjective geometric morphism from a topos with enough points to $\mathcal{E}$. As sheaf topoi over topological spaces have enough points, any surjective geometric morphism from some $Sh(X)$, with $X$ a topological space, to $\mathcal{E}$ is sufficient. These are provided by Barr's theorem. This is the approach that will be taken here (following MacLane and Moerdijk, Sheaves in Geometry and Logic).

math.LO