SearcharxivSearch

arXiv subjects

M. Weigt

Publications and source records attributed to M. Weigt.

At least 19 recordsLinked to original sources

Aligning graphs and finding substructures by a cavity approach

We introduce a new distributed algorithm for aligning graphs or finding substructures within a given graph. It is based on the cavity method and is used to study the maximum-clique and the graph-alignment problems in random graphs. The algorithm allows to analyze large graphs and may find applications in fields such as computational biology. As a proof of concept we use our algorithm to align the similarity graphs of two interacting protein families involved in bacterial signal transduction, and to predict actually interacting protein partners between these families.

q-bio.QM

Identification of direct residue contacts in protein-protein interaction by message passing

Understanding the molecular determinants of specificity in protein-protein interaction is an outstanding challenge of postgenome biology. The availability of large protein databases generated from sequences of hundreds of bacterial genomes enables various statistical approaches to this problem. In this context covariance-based methods have been used to identify correlation between amino acid positions in interacting proteins. However, these methods have an important shortcoming, in that they cannot distinguish between directly and indirectly correlated residues. We developed a method that combines covariance analysis with global inference analysis, adopted from use in statistical physics. Applied to a set of >2,500 representatives of the bacterial two-component signal transduction system, the combination of covariance with global inference successfully and robustly identified residue pairs that are proximal in space without resorting to ad hoc tuning parameters, both for heterointeractions between sensor kinase (SK) and response regulator (RR) proteins and for homointeractions between RR proteins. The spectacular success of this approach illustrates the effectiveness of the global inference approach in identifying direct interaction based on sequence information alone. We expect this method to be applicable soon to interaction surfaces between proteins present in only 1 copy per genome as the number of sequenced genomes continues to expand. Use of this method could significantly increase the potential targets for therapeutic intervention, shed light on the mechanism of protein-protein interaction, and establish the foundation for the accurate prediction of interacting protein partners.

q-bio.BM

Inference algorithms for gene networks: a statistical mechanics analysis

The inference of gene regulatory networks from high throughput gene expression data is one of the major challenges in systems biology. This paper aims at analysing and comparing two different algorithmic approaches. The first approach uses pairwise correlations between regulated and regulating genes; the second one uses message-passing techniques for inferring activating and inhibiting regulatory interactions. The performance of these two algorithms can be analysed theoretically on well-defined test sets, using tools from the statistical physics of disordered systems like the replica method. We find that the second algorithm outperforms the first one since it takes into account collective effects of multiple regulators.

q-bio.QM

Gene-network inference by message passing

The inference of gene-regulatory processes from gene-expression data belongs to the major challenges of computational systems biology. Here we address the problem from a statistical-physics perspective and develop a message-passing algorithm which is able to infer sparse, directed and combinatorial regulatory mechanisms. Using the replica technique, the algorithmic performance can be characterized analytically for artificially generated data. The algorithm is applied to genome-wide expression data of baker's yeast under various environmental conditions. We find clear cases of combinatorial control, and enrichment in common functional annotations of regulated genes and their regulators.

q-bio.QM

Computational core and fixed-point organisation in Boolean networks

In this paper, we analyse large random Boolean networks in terms of a constraint satisfaction problem. We first develop an algorithmic scheme which allows to prune simple logical cascades and under-determined variables, returning thereby the computational core of the network. Second we apply the cavity method to analyse number and organisation of fixed points. We find in particular a phase transition between an easy and a complex regulatory phase, the latter one being characterised by the existence of an exponential number of macroscopically separated fixed-point clusters. The different techniques developed are reinterpreted as algorithms for the analysis of single Boolean networks, and they are applied to analysis and in silico experiments on the gene-regulatory networks of baker's yeast (saccaromices cerevisiae) and the segment-polarity genes of the fruit-fly drosophila melanogaster.

cond-mat.stat-mech

Core percolation and onset of complexity in Boolean networks

The determination and classification of fixed points of large Boolean networks is addressed in terms of constraint satisfaction problem. We develop a general simplification scheme that, removing all those variables and functions belonging to trivial logical cascades, returns the computational core of the network. The onset of an easy-to-complex regulatory phase is introduced as a function of the parameters of the model, identifying both theoretically and algorithmically the relevant regulatory variables.

cond-mat.dis-nn

Constraint Satisfaction by Survey Propagation

Survey Propagation is an algorithm designed for solving typical instances of random constraint satisfiability problems. It has been successfully tested on random 3-SAT and random $G(n,\frac{c}{n})$ graph 3-coloring, in the hard region of the parameter space. Here we provide a generic formalism which applies to a wide class of discrete Constraint Satisfaction Problems.

cond-mat.dis-nn

Glassy behavior induced by geometrical frustration in a hard-core lattice gas model

We introduce a hard-core lattice-gas model on generalized Bethe lattices and investigate analytically and numerically its compaction behavior. If compactified slowly, the system undergoes a first-order crystallization transition. If compactified much faster, the system stays in a meta-stable liquid state and undergoes a glass transition under further compaction. We show that this behavior is induced by geometrical frustration which appears due to the existence of short loops in the generalized Bethe lattices. We also compare our results to numerical simulations of a three-dimensional analog of the model.

cond-mat.stat-mech

Polynomial iterative algorithms for coloring and analyzing random graphs

We study the graph coloring problem over random graphs of finite average connectivity $c$. Given a number $q$ of available colors, we find that graphs with low connectivity admit almost always a proper coloring whereas graphs with high connectivity are uncolorable. Depending on $q$, we find the precise value of the critical average connectivity $c_q$. Moreover, we show that below $c_q$ there exist a clustering phase $c\in [c_d,c_q]$ in which ground states spontaneously divide into an exponential number of clusters. Furthermore, we extended our considerations to the case of single instances showing consistent results. This lead us to propose a new algorithm able to color in polynomial time random graphs in the hard but colorable region, i.e when $c\in [c_d,c_q]$.

cond-mat.dis-nn

Coloring random graphs

We study the graph coloring problem over random graphs of finite average connectivity $c$. Given a number $q$ of available colors, we find that graphs with low connectivity admit almost always a proper coloring whereas graphs with high connectivity are uncolorable. Depending on $q$, we find the precise value of the critical average connectivity $c_q$. Moreover, we show that below $c_q$ there exist a clustering phase $c\in [c_d,c_q]$ in which ground states spontaneously divide into an exponential number of clusters and where the proliferation of metastable states is responsible for the onset of complexity in local search algorithms.

cond-mat.stat-mech

Hiding solutions in random satisfiability problems: A statistical mechanics approach

A major problem in evaluating stochastic local search algorithms for NP-complete problems is the need for a systematic generation of hard test instances having previously known properties of the optimal solutions. On the basis of statistical mechanics results, we propose random generators of hard and satisfiable instances for the 3-satisfiability problem (3SAT). The design of the hardest problem instances is based on the existence of a first order ferromagnetic phase transition and the glassy nature of excited states. The analytical predictions are corroborated by numerical results obtained from complete as well as stochastic local algorithms.

cond-mat.dis-nn

A ferromagnet with a glass transition

We introduce a finite-connectivity ferromagnetic model with a three-spin interaction which has a crystalline (ferromagnetic) phase as well as a glass phase. The model is not frustrated, it has a ferromagnetic equilibrium phase at low temperature which is not reached dynamically in a quench from the high-temperature phase. Instead it shows a glass transition which can be studied in detail by a one step replica-symmetry broken calculation. This spin model exhibits the main properties of the structural glass transition at a solvable mean-field level.

cond-mat.dis-nn

Simplest random K-satisfiability problem

We study a simple and exactly solvable model for the generation of random satisfiability problems. These consist of $γN$ random boolean constraints which are to be satisfied simultaneously by $N$ logical variables. In statistical-mechanics language, the considered model can be seen as a diluted p-spin model at zero temperature. While such problems become extraordinarily hard to solve by local search methods in a large region of the parameter space, still at least one solution may be superimposed by construction. The statistical properties of the model can be studied exactly by the replica method and each single instance can be analyzed in polynomial time by a simple global solution method. The geometrical/topological structures responsible for dynamic and static phase transitions as well as for the onset of computational complexity in local search method are thoroughly analyzed. Numerical analysis on very large samples allows for a precise characterization of the critical scaling behaviour.

cond-mat.dis-nn

On the properties of small-world network models

We study the small-world networks recently introduced by Watts and Strogatz [Nature {\bf 393}, 440 (1998)], using analytical as well as numerical tools. We characterize the geometrical properties resulting from the coexistence of a local structure and random long-range connections, and we examine their evolution with size and disorder strength. We show that any finite value of the disorder is able to trigger a ``small-world'' behaviour as soon as the initial lattice is big enough, and study the crossover between a regular lattice and a ``small-world'' one. These results are corroborated by the investigation of an Ising model defined on the network, showing for every finite disorder fraction a crossover from a high-temperature region dominated by the underlying one-dimensional structure to a mean-field like low-temperature region. In particular there exists a finite-temperature ferromagnetic phase transition as soon as the disorder strength is finite.

cond-mat.dis-nn

Multifractal analysis of perceptron learning with errors

Random input patterns induce a partition of the coupling space of a perceptron into cells labeled by their output sequences. Learning some data with a maximal error rate leads to clusters of neighboring cells. By analyzing the internal structure of these clusters with the formalism of multifractals, we can handle different storage and generalization tasks for lazy students and absent-minded teachers within one unified approach. The results also allow some conclusions on the spatial distribution of cells.

cond-mat.dis-nn

A replica approach to products of random matrices

We analyse products of random $R\times R$ matrices by means of a variant of the replica trick which was recently introduced for one-dimensional disordered Ising models. The replicated transfer matrix can be block-diagonalized with help of irreducible representations of the permutation group. We show that the free energy (or the Lyapunov exponent) of the product corresponds to the replica symmetric representation, whereas non-trivial representations correspond to certain correlation functions.

cond-mat.dis-nn

Multifractality and percolation in the coupling space of perceptrons

The coupling space of perceptrons with continuous as well as with binary weights gets partitioned into a disordered multifractal by a set of $p=γN$ random input patterns. The multifractal spectrum $f(α)$ can be calculated analytically using the replica formalism. The storage capacity and the generalization behaviour of the perceptron are shown to be related to properties of $f(α)$ which are correctly described within the replica symmetric ansatz. Replica symmetry breaking is interpreted geometrically as a transition from percolating to non-percolating cells. The existence of empty cells gives rise to singularities in the multifractal spectrum. The analytical results for binary couplings are corroborated by numerical studies.

cond-mat.dis-nn

Replica structure of one--dimensional Ising models

We analyse the eigenvalue structure of the replicated transfer matrix of one-dimensional disordered Ising models. In the limit of $n \rightarrow 0$ replicas, an infinite sequence of transfer matrices is found, each corresponding to a different irreducible representation (labelled by a positive integer $ρ$) of the permutation group. We show that the free energy can be calculated from the replica symmetric subspace ($ρ=0$). The other ``replica symmetry broken'' representations ($ρ\ne 0$) are physically meaningful since their largest eigenvalues $λ^{(ρ)}$ control the disorder--averaged moments $\ll( \langle S_i S_j \rangle - \langle S_i \rangle \langle S_j \rangle )^ρ\gg \propto (λ^{(ρ)}) ^{|i-j|}$ of the connected two-points correlations.

cond-mat