Searcharxiv⌕ Search

arXiv subjects

Alejandro Lage-Castellanos

Publications and source records attributed to Alejandro Lage-Castellanos.

At least 19 recordsLinked to original sources

Epistatic strength, modularity, and locus heterogeneity shape the number of local optima in fitness landscapes

Fitness landscapes provide a quantitative framework for understanding how natural selection shapes evolutionary trajectories. A central feature of these landscapes is their number of local optima, which determines whether fitness-increasing evolution can proceed towards a global optimum or become trapped on suboptimal peaks. Although multiple peaks are known to require reciprocal sign epistasis, the quantitative relationship between epistasis and number of peaks remains incompletely understood. Here, we show that for a broad class of unstructured fitness landscapes, i.e. isotropic Gaussian random fields, the expected number of local optima is determined by a single local measure of epistasis: the correlation of fitness effects. This provides a baseline prediction for the number of peaks in typical unstructured landscapes and links peak density directly to the amount of reciprocal sign epistasis. This baseline changes when epistatic interactions are structured. We show that clustering interactions within blocks of loci slightly increases the number of local optima. In contrast, strong heterogeneity between loci, where only a small subset of loci participate in epistatic interactions, causes the number of peaks to collapse. These results show that the number of local optima is governed not only by the overall strength of epistasis, but also by how epistatic interactions are distributed across the genotype space. Our framework therefore reconciles the central role of reciprocal sign epistasis with the observation that landscapes with similar amounts of epistasis can differ substantially in ruggedness, and provides a guide to the range of peak numbers expected in typical landscapes.

q-bio.PE↗

Simple sign epistasis and evolutionary detours in fitness landscapes

In epistatic fitness landscapes, the fitness effect of a mutation depends on the genetic background and may even switch between deleterious and beneficial depending on the presence of another mutation. Epistatic interactions may cause both mutations to change the sign of each other's fitness effects (reciprocal sign epistasis) or only one mutation to do so (simple sign epistasis). Both these forms of epistasis influence evolutionary trajectories. While reciprocal sign epistasis has been associated with multi-peaked landscapes and their ruggedness, the role and relative frequency of simple sign epistasis in fitness landscapes have not been systematically investigated. Here, we prove that the presence of simple sign epistasis is associated with evolutionary detours, i.e., indirect, longer fitness-increasing paths to fitness peaks that include back-mutations. We also show that in experimentally resolved, weakly epistatic landscapes, simple sign epistasis occurs much more frequently than reciprocal sign epistasis. This result is consistent with the theoretical predictions we derive for most landscape models, with the exception of the block model and of landscapes dominated by pairwise allelic incompatibilities, such as RNA stability landscapes. Our results suggest that detours represent a general feature of evolutionary trajectories in weakly epistatic landscapes.

q-bio.PE↗

From informal markets to Limit Order Book dynamics: a mean field connection

We propose a unified mean-field framework that bridges the dynamics of informal financial markets and formal markets governed by Limit Order Books (LOBs). Both settings are modeled as interacting particle systems on a 1D price lattice, with temporal evolution described by master equations that account for new entries, cancellations, and executions. The key insight is the introduction of a preferential interaction parameter $Ψ$, which modulates the likelihood of transactions based on price compatibility: when $Ψ=0$, interactions are random and uncoordinated, reproducing the structure of informal markets; as $Ψ\to \infty$, only optimal (most mutually attractive) trades occur, recovering LOB-like dynamics. A grand-canonical interpretation is used to identify effective thermodynamic quantities -such as interaction energy and price-dependent chemical potentials -that underlie both systems. Most results are validated through numerical integration and simulations, although an analytical solution is shown to exist at least for the symmetric stationary case of the informal market.

cond-mat.stat-mech↗

Opinion formation by belief propagation: A heuristic to identify low-credible sources of information

With social media, the flow of uncertified information is constantly increasing, with the risk that more people will trust low-credible information sources. To design effective strategies against this phenomenon, it is of paramount importance to understand how people end up believing one source rather than another. To this end, we propose a realistic and cognitively affordable heuristic mechanism for opinion formation inspired by the well-known belief propagation algorithm. In our model, an individual observing a network of information sources must infer which of them are reliable and which are not. We study how the individual's ability to identify credible sources, and hence to form correct opinions, is affected by the noise in the system, intended as the amount of disorder in the relationships between the information sources in the network. We find numerically and analytically that there is a critical noise level above which it is impossible for the individual to detect the nature of the sources. Moreover, by comparing our opinion formation model with existing ones in the literature, we show under what conditions people's opinions can be reliable. Overall, our findings imply that the increasing complexity of the information environment is a catalyst for misinformation channels.

physics.soc-ph↗

City path tomography: reconstructing square road network from artificial users mobile phone data

Population mobility can be studied readily and cheaply using cellphone data, since people's mobility can be approximately mapped into tower-mobile registries. We model people moving in a grid-like city, where edges of the grid are weighted and paths are chosen according to overall weights between origin and destination. Cellphone users leave sparse signals in random nodes of the grid as they move by, mimicking the type of data collected from the tower-cellphone interactions. From this noisy data we seek to build a model of the city, {\it i.e.} to predict probabilities of paths from origin to destination. We focus on the simplest case where users move along shortest paths (no loops, no going backwards). In this simplified setting, we are able to infer the underlying weights of the edges (akin to road transitability) with an inverse statistical mechanic model.

physics.soc-ph↗

Ancestral Sequence Reconstruction for Co-evolutionary models

The ancestral sequence reconstruction problem is the inference, back in time, of the properties of common sequence ancestors from measured properties of contemporary populations. Standard algorithms for this problem assume independent (factorized) evolution of the characters of the sequences, which is generally wrong (e.g. proteins and genome sequences). In this work, we have studied this problem for sequences described by global co-evolutionary models, which reproduce the global pattern of cooperative interactions between the elements that compose it. For this, we first modeled the temporal evolution of correlated real valued characters by a multivariate Ornstein-Uhlenbeck process on a finite tree. This represents sequences as Gaussian vectors evolving in a quadratic potential, who describe selection forces acting on the evolving entities. Under a Bayesian framework, we developed a reconstruction algorithm for these sequences and obtained an analytical expression to quantify the quality of our estimation. We extend this formalism to discrete valued sequences by applying our method to a Potts model. We showed that for both continuous and discrete configurations, there is a wide range of parameters where, to properly reconstruct the ancestral sequences, intra-species correlations must be taken into account. We also demonstrated that, for sequences with discrete elements, our reconstruction algorithm outperforms traditional schemes based on independent site approximations.

cond-mat.dis-nn↗

Dynamics of epidemic models from cavity master equations

We apply the cavity master equation (CME) approach to epidemics models. We explore mostly the susceptible-infectious-susceptible (SIS) model, which can be readily treated with the CME as a two-state. We show that this approach is more accurate than individual based and pair based mean field methods, and a previously published dynamic message passing scheme. We explore average case predictions and extend the cavity master equation to SIR and SIRS models.

physics.soc-ph↗

Estimating undocumented Covid-19 infections in Cuba by means of a hybrid mechanistic-statistical approach

We adapt the hybrid mechanistic-statistical approach of Ref. [1] to estimate the total number of undocumented Covid-19 infections in Cuba. This scheme is based on the maximum likelihood estimation of a SIR-like model parameters for the infected population, assuming that the detection process matches a Bernoulli trial. Our estimations show that (a) 60% of the infections were undocumented, (b) the real epidemics behind the data peaked ten days before the reports suggested, and (c) the reproduction number swiftly vanishes after 80 epidemic days.

q-bio.PE↗

Contamination Source Detection in Water Distribution Networks using Belief Propagation

We present a Bayesian approach for the Contamination Source Detection problem in Water Distribution Networks. Given an observation of contaminants in one or more nodes in the network, we try to give probable explanation for it assuming that contamination is a rare event. We introduce extra variables to characterize the place and pattern of the first contamination event. Then we write down the posterior distribution for these extra variables given the observation obtained by the sensors. Our method relies on Belief Propagation for the evaluation of the marginals of this posterior distribution and the determination of the most likely origin. The method is implemented on a simplified binary forward-in-time dynamics. Simulations on data coming from the realistic simulation software EPANET on two networks show that the simplified model is nevertheless flexible enough to capture crucial information about contaminant sources.

physics.data-an↗

Contamination source inference in water distribution networks

We study the inference of the origin and the pattern of contamination in water distribution networks. We assume a simplified model for the dyanmics of the contamination spread inside a water distribution network, and assume that at some random location a sensor detects the presence of contaminants. We transform the source location problem into an optimization problem by considering discrete times and a binary contaminated/not contaminated state for the nodes of the network. The resulting problem is solved by Mixed Integer Linear Programming. We test our results on random networks as well as in the Modena city network.

physics.soc-ph↗

Low Auto-correlation Binary Sequences explored using Warning Propagation

The search of binary sequences with low auto-correlations (LABS) is a discrete combinatorial optimization problem contained in the NP-hard computational complexity class. We study this problem using Warning Propagation (WP) , a message passing algorithm, and compare the performance of the algorithm in the original problem and in two different disordered versions. We show that in all the cases Warning Propagation converges to low energy minima of the solution space. Our results highlight the importance of the local structure of the interaction graph of the variables for the convergence time of the algorithm and for the quality of the solutions obtained by WP. While in general the algorithm does not provide the optimal solutions in large systems it does provide, in polynomial time, solutions that are energetically similar to the optimal ones. Moreover, we designed hybrid models that interpolate between the standard LABS problem and the disordered versions of it, and exploit them to improved the convergence time of WP and the quality of the solutions.

cond-mat.dis-nn↗

Gauge-free cluster variational method by maximal messages and moment matching

We present a new implementation of the Cluster Variational Method (CVM) as a message passing algorithm. The kind of message passing algorithms used for CVM, usually named Generalized Belief Propagation, are a generalization of the Belief Propagation algorithm in the same way that CVM is a generalization of the Bethe approximation for estimating the partition function. However, the connection between fixed points of GBP and the extremal points of the CVM free-energy is usually not a one-to-one correspondence, because of the existence of a gauge transformation involving the GBP messages. Our contribution is twofold. Firstly we propose a new way of defining messages (fields) in a generic CVM approximation, such that messages arrive on a given region from all its ancestors, and not only from its direct parents, as in the standard Parent-to-Child GBP. We call this approach maximal messages. Secondly we focus on the case of binary variables, re-interpreting the messages as fields enforcing the consistency between the moments of the local (marginal) probability distributions. We provide a precise rule to enforce all consistencies, avoiding any redundancy, that would otherwise lead to a gauge transformation on the messages. This moment matching method is gauge free, i.e. it guarantees that the resulting GBP is not gauge invariant. We apply our maximal messages and moment matching GBP to obtain an analytical expression for the critical temperature of the Ising model in general dimensions at the level of plaquette-CVM. The values obtained outperform Bethe estimates, and are comparable with loop corrected Belief Propagation equations. The method allows for a straightforward generalization to disordered systems.

cond-mat.dis-nn↗

Random Field Ising Model in two dimensions: Bethe approximation, Cluster Variational Method and message passing algorithms

We study two free energy approximations (Bethe and plaquette-CVM) for the Random Field Ising Model in two dimensions. We compare results obtained by these two methods in single instances of the model on the square grid, showing the difficulties arising in defining a robust critical line. We also attempt average case calculations using a replica-symmetric ansatz, and compare the results with single instances. Both, Bethe and plaquette-CVM approximations present a similar panorama in the phase space, predicting long range order at low temperatures and fields. We show that plaquette-CVM is more precise, in the sense that predicts a lower critical line (the truth being no line at all). Furthermore, we give some insight on the non-trivial structure of the fixed points of different message passing algorithms.

cond-mat.dis-nn↗

A cavity approach to optimization and inverse dynamical problems

In these two lectures we shall discuss how the cavity approach can be used efficiently to study optimization problems with global (topological) constraints and how the same techniques can be generalized to study inverse problems in irreversible dynamical processes. These two classes of problems are formally very similar: they both require an efficient procedure to trace over all trajectories of either auxiliary variables which enforce global constraints, or directly dynamical variables defining the inverse dynamical problems. We will mention three basic examples, namely the Minimum Steiner Tree problem, the inverse threshold linear dynamical problem, and the patient-zero problem in epidemic cascades. All these examples are root problems in optimization and inference over networks. They appear in many modern applications and in a variety of different contexts. Credit for these results should be shared with A. Braunstein, A. Ramezanpour, F. Altarelli, L. Dall'Asta, I. Biazzo and A. Lage-Castellanos.

cond-mat.dis-nn↗

Bayesian inference of epidemics on networks via Belief Propagation

We study several bayesian inference problems for irreversible stochastic epidemic models on networks from a statistical physics viewpoint. We derive equations which allow to accurately compute the posterior distribution of the time evolution of the state of each node given some observations. At difference with most existing methods, we allow very general observation models, including unobserved nodes, state observations made at different or unknown times, and observations of infection times, possibly mixed together. Our method, which is based on the Belief Propagation algorithm, is efficient, naturally distributed, and exact on trees. As a particular case, we consider the problem of finding the "zero patient" of a SIR or SI epidemic given a snapshot of the state of the network at a later unknown time. Numerical simulations show that our method outperforms previous ones on both synthetic and real networks, often by a very large margin.

q-bio.QM↗

Stability of the replica symmetric solution in diluted perceptron learning

We study the role played by the dilution in the average behavior of a perceptron model with continuous coupling with the replica method. We analyze the stability of the replica symmetric solution as a function of the dilution field for the generalization and memorization problems. Thanks to a Gardner like stability analysis we show that at any fixed ratio $α$ between the number of patterns M and the dimension N of the perceptron ($α=M/N$), there exists a critical dilution field $h_c$ above which the replica symmetric ansatz becomes unstable.

cond-mat.dis-nn↗

Replica Cluster Variational Method: the Replica Symmetric solution for the 2D random bond Ising model

We present and solve the Replica Symmetric equations in the context of the Replica Cluster Variational Method for the 2D random bond Ising model (including the 2D Edwards-Anderson spin glass model). First we solve a linearized version of these equations to obtain the phase diagrams of the model on the square and triangular lattices. In both cases the spin-glass transition temperatures and the tricritical point estimations improve largely over the Bethe predictions. Moreover, we show that this phase diagram is consistent with the behavior of inference algorithms on single instances of the problem. Finally, we present a method to consistently find approximate solutions to the equations in the glassy phase. The method is applied to the triangular lattice down to T=0, also in the presence of an external field.

cond-mat.dis-nn↗

A very fast inference algorithm for finite-dimensional spin glasses: Belief Propagation on the dual lattice

Starting from a Cluster Variational Method, and inspired by the correctness of the paramagnetic Ansatz (at high temperatures in general, and at any temperature in the 2D Edwards-Anderson model) we propose a novel message passing algorithm --- the Dual algorithm --- to estimate the marginal probabilities of spin glasses on finite dimensional lattices. We show that in a wide range of temperatures our algorithm compares very well with Monte Carlo simulations, with the Double Loop algorithm and with exact calculation of the ground state of 2D systems with bimodal and Gaussian interactions. Moreover it is usually 100 times faster than other provably convergent methods, as the Double Loop algorithm.

cond-mat.dis-nn↗