SearcharxivSearch

arXiv subjects

Guilherme Ost

Publications and source records attributed to Guilherme Ost.

13 recordsLinked to original sources

Community detection for binary graphical models in high dimension

Let $N$ components be partitioned into two communities, denoted ${\cal P}_+$ and ${\cal P}_-$, possibly of different sizes. Assume that they are connected via a directed and weighted Erdös-Rényi (DWER) random graph with unknown parameter $ p \in (0, 1).$ The weights assigned to the existing connections are of mean-field-type, scaling as $N^{-1}$. At each time \modif{step}, we observe the state of each component: either it sends some signal to its successors (in the directed graph) or remains silent otherwise. In this paper, we show that it is possible to find the communities ${\cal P}_+$ and ${\cal P}_-$ based only on the activity of the $N$ components observed over $T$ time units. More specifically, we propose \modif{ two simple methods, an aggregated method and a spectral method, whose {\it misclassification rates} vanish as long as $T \gg N$ (up to log terms). This condition is proved to be near-optimal in the minimax sense. Moreover, under the stronger condition $T \gg N^2$ (up to log terms), the aggregated method is shown to achieve {\it exact recovery} with probability tending to $1$. } Interestingly, these simple \modif{methods} do not require any prior knowledge of the other model parameters (e.g. the edge probability $p$). The key step in our analysis is to derive an asymptotic approximation of the 1-lagged covariance matrix associated to the states of the $N$ components, as $N$ diverges. This asymptotic approximation relies on the study of the behavior of the solutions of a \modif{Stein-type} matrix equation satisfied by the simultaneous (0-lagged) covariance matrix associated to the states of the components. This study is challenging, especially because the simultaneous covariance matrix is random since it depends on the underlying DWER random graph.

math.ST

hdMTD: An R Package for High-Dimensional Mixture Transition Distribution Models

Several natural phenomena exhibit long-range conditional dependencies. High-order mixture transition distribution (MTD) are parsimonious non-parametric models to study these phenomena. An MTD is a Markov chain in which the transition probabilities are expressed as a convex combination of lower-order conditional distributions. Despite their generality, inference for MTD models has traditionally been limited by the need to estimate high-dimensional joint distributions. In particular, for a sample of size n, the feasible order d of the MTD is typically restricted to d approximately O(log n). To overcome this limitation, Ost and Takahashi (2023) recently introduced a computationally efficient non-parametric inference method that identifies the relevant lags in high-order MTD models, even when d is approximately O(n), provided that the set of relevant lags is sparse. In this article, we introduce hdMTD, an R package allowing us to estimate parameters of such high-dimensional Markovian models. Given a sample from an MTD chain, hdMTD can retrieve the relevant past set using the BIC algorithm or the forward stepwise and cut algorithm described in Ost and Takahashi (2023). The package also computes the maximum likelihood estimate for transition probabilities and estimates high-order MTD parameters through the expectation-maximization algorithm. Additionally, hdMTD also allows for simulating an MTD chain from its stationary invariant distribution using the perfect (exact) sampling algorithm, enabling Monte Carlo simulation of the model. We illustrate the package's capabilities through simulated data and a real-world application involving temperature records from Brazil.

stat.ME

Inferring the dependence graph density of binary graphical models in high dimension

We consider a system of binary interacting chains describing the dynamics of a group of $N$ components that, at each time unit, either send some signal to the others or remain silent otherwise. The interactions among the chains are encoded by a directed Erdös-Rényi random graph with unknown parameter $ p \in (0, 1) .$ Moreover, the system is structured within two populations (excitatory chains versus inhibitory ones) which are coupled via a mean field interaction on the underlying Erdös-Rényi graph. In this paper, we address the question of inferring the connectivity parameter $p$ based only on the observation of the interacting chains over $T$ time units. In our main result, we show that the connectivity parameter $p$ can be estimated with rate $N^{-1/2}+N^{1/2}/T+(\log(T)/T)^{1/2}$ through an easy-to-compute estimator. Our analysis relies on a precise study of the spatio-temporal decay of correlations of the interacting chains. This is done through the study of coalescing random walks defining a backward regeneration representation of the system. Interestingly, we also show that this backward regeneration representation allows us to perfectly sample the system of interacting chains (conditionally on each realization of the underlying Erdös-Rényi graph) from its stationary distribution. These probabilistic results have an interest in its own.

math.ST

Self-switching random walks on Erdös-Rényi random graphs feel the phase transition

We study random walks on Erdös-Rényi random graphs in which, every time the random walk returns to the starting point, first an edge probability is independently sampled according to a priori measure $μ$, and then an Erdös-Rényi random graph is sampled according to that edge probability. When the edge probability $p$ does not depend on the size of the graph $n$ (dense case), we show that the proportion of time the random walk spends on different values of $p$ -- {\it occupation measure} -- converges to the a priori measure $μ$ as $n$ goes to infinity. More interestingly, when $p=λ/n$ (sparse case), we show that the occupation measure converges to a limiting measure with a density that is a function of the survival probability of a Poisson branching process. This limiting measure is supported on the supercritial values for the Erdös-Rényi random graphs, showing that self-witching random walks can detect the phase transition.

math.PR

Neural Coding as a Statistical Testing Problem

We take the testing perspective to understand what the minimal discrimination time between two stimuli is for different types of rate coding neurons. Our main goal is to describe the testing abilities of two different encoding systems: place cells and grid cells. In particular, we show, through the notion of adaptation, that a fixed place cell system can have a minimum discrimination time that decreases when the stimuli are further away. This could be a considerable advantage for the place cell system that could complement the grid cell system, which is able to discriminate stimuli that are much closer than place cells.

math.ST

Sparse Markov Models for High-dimensional Inference

Finite order Markov models are theoretically well-studied models for dependent discrete data. Despite their generality, application in empirical work when the order is large is rare. Practitioners avoid using higher order Markov models because (1) the number of parameters grow exponentially with the order and (2) the interpretation is often difficult. Mixture of transition distribution models (MTD) were introduced to overcome both limitations. MTD represent higher order Markov models as a convex mixture of single step Markov chains, reducing the number of parameters and increasing the interpretability. Nevertheless, in practice, estimation of MTD models with large orders are still limited because of curse of dimensionality and high algorithm complexity. Here, we prove that if only few lags are relevant we can consistently and efficiently recover the lags and estimate the transition probabilities of high-dimensional MTD models. The key innovation is a recursive procedure for the selection of the relevant lags of the model. Our results are based on (1) a new structural result of the MTD and (2) an improved martingale concentration inequality. We illustrate our method using simulations and a weather data.

math.ST

Retrieving the structure of probabilistic sequences of auditory stimuli from EEG data

Using a new probabilistic approach we model the relationship between sequences of auditory stimuli generated by stochastic chains and the electroencephalographic (EEG) data acquired while 19 participants were exposed to those stimuli. The structure of the chains generating the stimuli are characterized by rooted and labeled trees whose leaves, henceforth called contexts, represent the sequences of past stimuli governing the choice of the next stimulus. A classical conjecture claims that the brain assigns probabilistic models to samples of stimuli. If this is true, then the context tree generating the sequence of stimuli should be encoded in the brain activity. Using an innovative statistical procedure we show that this context tree can effectively be extracted from the EEG data, thus giving support to the classical conjecture.

q-bio.NC

Fluctuations for Spatially Extended Hawkes Processes

In a previous paper, it has been shown that the mean-field limit of spatially extended Hawkes processes is characterized as the unique solution $u(t,x)$ of a neural field equation (NFE). The value $u(t,x)$ represents the membrane potential at time $t$ of a typical neuron located in position $x$, embedded in an infinite network of neurons. In the present paper, we complement this result by studying the fluctuations of such a stochastic system around its mean-field limit $u(t,x)$. Our first main result is a central limit theorem stating that the spatial distribution associated with these fluctuations converges to the unique solution of some stochastic differential equation driven by a Gaussian noise. In our second main result, we show that the solutions of this stochastic differential equation can be well approximated by a stochastic version of the neural field equation satisfied by $u(t,x)$. To the best of our knowledge, this result appears to be new in the literature.

math.PR

Sparse space-time models: Concentration Inequalities and Lasso

Inspired by Kalikow-type decompositions, we introduce a new stochastic model of infinite neuronal networks, for which we establish sharp oracle inequalities for Lasso methods and restricted eigenvalue properties for the associated Gram matrix with high probability. These results hold even if the network is only partially observed. The main argument rely on the fact that concentration inequalities can easily be derived whenever the transition probabilities of the underlying process admit a sparse space-time representation.

math.ST

Stability, convergence to equilibrium and simulation of non-linear Hawkes Processes with memory kernels given by the sum of Erlang kernels

Non-linear Hawkes processes with memory kernels given by the sum of Erlang kernels are considered. It is shown that their stability properties can be studied in terms of an associated class of piecewise deterministic Markov processes, called Markovian cascades of successive memory terms. Explicit conditions implying the positive Harris recurrence of these processes are presented. The proof is based on integration by parts with respect to the jump times. A crucial property is the non-degeneracy of the transition semigroup which is obtained thanks to the invertibility of an associated Vandermonde matrix. For Lipschitz continuous rate functions we also show that these Markovian cascades converge to equilibrium exponentially fast with respect to the Wasserstein distance. Finally, an extension of the classical thinning algorithm is proposed to simulate such Markovian cascades.

math.PR

Estimation of neuronal interaction graph from spike train data

One of the main current issues in Neurobiology concerns the understanding of interrelated spiking activity among multineuronal ensembles and differences between stimulus-driven and spontaneous activity in neurophysiological experiments. Multi electrode array recordings that are now commonly used monitor neuronal activity in the form of spike trains from many well identified neurons. A basic question when analyzing such data is the identification of the directed graph describing "synaptic coupling" between neurons. In this article we deal with this matter working with a high quality multielectrode array recording dataset (Pouzat et al., 2015) from the first olfactory relay of the locust, $Schistocerca$ $americana$. From a mathematical point of view this paper presents two novelties. First we propose a procedure allowing to deal with the small sample sizes met in actual datasets. Moreover we address the sensitive case of partially observed networks. Our starting point is the procedure introduced in Duarte et al. (2016). We evaluate the performance of both original and improved procedures through simulation studies, which are also used for parameter tuning and for exploring the effect of recording only a small subset of the neurons of a network.

q-bio.NC

A model for neural activity in the absence of external stimuli

We study a stochastic process describing the continuous time evolution of the membrane potentials of finite system of neurons in the absence of external stimuli. The values of the membrane potentials evolve under the effect of {\it chemical synapses}, {\it electrical synapses} and a \textit{leak current}. The evolution of the process can be informally described as follows. Each neuron spikes randomly following a point process with rate depending on its membrane potential. When a neuron spikes, its membrane potential is immediately reset to a resting value. Simultaneously, the membrane potential of the neurons which are influenced by it receive an additional positive value. Furthermore, between consecutive spikes, the system follows a deterministic motion due both to electrical synapses and the leak current. Electrical synapses push the system towards its average potential, while the leak current attracts the membrane potential of each neuron to the resting value. We show that in the absence leakage the process converges exponentially fast to an unique invariant measure, whenever the initial configuration is non null. More interesting, when leakage is present, we proved the system stops spiking after a finite amount of time almost surely. This implies that the unique invariant measure is supported only by the null configuration.

math.PR

Hydrodynamic Limits for Spatially Structured Interacting Neurons

In this paper we study the hydrodynamic limit for a stochastic process describing the time evolution of the membrane potentials of a system of neurons with spatial dependency. We do not impose on the neurons mean-field type interactions. The values of the membrane potentials evolve under the effect of chemical and electrical synapses and leak currents. The system consists of $ε^{-2}$ neurons embedded in $[0,1)^2$, each spiking randomly according to a point process with rate depending on both its membrane potential and position. When neuron $i$ spikes, its membrane potential is reset to a resting value while the membrane potential of $j$ is increased by a positive value $ε^2 a(i,j)$, if $i$ influences $j$. Furthermore, between consecutive spikes, the system follows a deterministic motion due both to electrical synapses and leak currents. The electrical synapses are involved in the synchronization of neurons. For each pair of neurons $(i,j)$, we modulate this synchronizing strength by $ε^2 b(i,j)$, where $b(i,j)$ is a nonnegative symmetric function. On the other hand, the leak currents inhibit the activity of all neurons, attracting simultaneously their membrane potentials to the resting value. In the main result of this paper is shown that the empirical distribution of the membrane potentials converges, as the parameter $ε$ goes to zero, to a probability density $ρ_t(u,r)$ which is proved to obey a non linear PDE of Hyperbolic type.

math.PR