SearcharxivSearch

arXiv subjects

Mikael Huss

Publications and source records attributed to Mikael Huss.

10 recordsLinked to original sources

Feature Space Saturation during Training

We propose layer saturation - a simple, online-computable method for analyzing the information processing in neural networks. First, we show that a layer's output can be restricted to the eigenspace of its variance matrix without performance loss. We propose a computationally lightweight method for approximating the variance matrix during training. From the dimension of its lossless eigenspace we derive layer saturation - the ratio between the eigenspace dimension and layer width. We show that saturation seems to indicate which layers contribute to network performance. We demonstrate how to alter layer saturation in a neural network by changing network depth, filter sizes and input resolution. Furthermore, we show that well-chosen input resolution increases network performance by distributing the inference process more evenly across the network.

cs.LG

Spectral Analysis of Latent Representations

We propose a metric, Layer Saturation, defined as the proportion of the number of eigenvalues needed to explain 99% of the variance of the latent representations, for analyzing the learned representations of neural network layers. Saturation is based on spectral analysis and can be computed efficiently, making live analysis of the representations practical during training. We provide an outlook for future applications of this metric by outlining the behaviour of layer saturation in different neural architectures and problems. We further show that saturation is related to the generalization and predictive performance of neural networks.

cs.LG

Atmospheric reaction systems as null-models to identify structural traces of evolution in metabolism

The metabolism is the motor behind the biological complexity of an organism. One problem of characterizing its large-scale structure is that it is hard to know what to compare it to. All chemical reaction systems are shaped by the same physics that gives molecules their stability and affinity to react. These fundamental factors cannot be captured by standard null-models based on randomization. The unique property of organismal metabolism is that it is controlled, to some extent, by an enzymatic machinery that is subject to evolution. In this paper, we explore the possibility that reaction systems of planetary atmospheres can serve as a null-model against which we can define metabolic structure and trace the influence of evolution. We find that the two types of data can be distinguished by their respective degree distributions. This is especially clear when looking at the degree distribution of the reaction network (of reaction connected to each other if they involve the same molecular species). For the Earth's atmospheric network and the human metabolic network, we look into more detail for an underlying explanation of this deviation. However, we cannot pinpoint a single cause of the difference, rather there are several concurrent factors. By examining quantities relating to the modular-functional organization of the metabolism, we confirm that metabolic networks have a more complex modular organization than the atmospheric networks, but not much more. We interpret the more variegated modular arrangement of metabolism as a trace of evolved functionality. On the other hand, it is quite remarkable how similar the structures of these two types of networks are, which emphasizes that the constraints from the chemical properties of the molecules has a larger influence in shaping the reaction system than does natural selection.

q-bio.MN

Substance graphs are optimal simple-graph representations of metabolism

One approach to studying the system-wide organization of biochemistry is to use statistical graph theory. Even in such a heavily simplified method, which disregards most of the dynamic aspects of biochemistry, one is faced with fundamental questions, such as how the chemical reaction systems should be reduced to a graph retaining as much functional information as possible from the original reaction system. In such graph representations, should the edges go between substrates and products, or substrates and substrates, or both? Should vertices represent substances or reactions? Different definitions encode different information about the reaction system. In this paper we evaluate four different graph representations of metabolism, applied to data from different organisms and databases. The graph representations are evaluated by comparing the overlap between clusters (network modules) and annotated functions, and also by comparing the set of identified currency metabolites with those that other authors have identified using qualitative biological arguments. We find that a "substance network," where all metabolites participating in a reaction are connected, is relatively better than others, evaluated both with respect to the functional overlap between modules and functions and to the number and identity of identified currency metabolites.

q-bio.MN

Comment on "Regularizing capacity of metabolic networks"

In a recent paper, Marr, Muller-Linow and Hutt [Phys. Rev. E 75, 041917 (2007)] investigate an artificial dynamic system on metabolic networks. They find a less complex time evolution of this dynamic system in real networks, compared to networks of reference models. The authors argue that this suggests that metabolic network structure is a major factor behind the stability of biochemical steady states. We reanalyze the same kind of data using a dynamic system modeling actual reaction kinetics. The conclusions about stability, from our analysis, are inconsistent with those of Marr et al. We argue that this issue calls for a more detailed type of modeling.

q-bio.MN

Currency and commodity metabolites: Their identification and relation to the modularity of metabolic networks

The large-scale shape and function of metabolic networks are intriguing topics of systems biology. Such networks are on one hand commonly regarded as modular (i.e. built by a number of relatively independent subsystems), but on the other hand they are robust in a way not expected of a purely modular system. To address this question we carefully discuss the partition of metabolic networks into subnetworks. The practice of preprocessing such networks by removing the most abundant substrates, "currency metabolites," is formalized into a network-based algorithm. We study partitions for metabolic networks of many organisms and find cores of currency metabolites and modular peripheries of what we call "commodity metabolites." The networks are found to be more modular than random networks but far from perfectly divisible into modules. We argue that cross-modular edges are the key for the robustness of metabolism.

q-bio.MN

Prediction of transcription factor binding to DNA using rule induction methods

The transcription of DNA into mRNA is initiated and aided by a number of transcription factors (TFs), proteins with DNA-binding regions that attach themselves to binding sites in the DNA (transcription factor binding sites, TFBSs). As it has become apparent that both TFs and TFBSs are highly variable, tools are needed to quantify the strength of the interaction resulting from a certain TF variant binding to a certain TFBS. We used a simple way to predict interactions between protein and DNA: given experimental cases from the literature where the interaction strength between two sequences has been quantified, we created training vectors for rule induction by regarding each amino acid and nucleotide position as a single feature in the example vector. The resulting interaction strength was used as the target class or value. These training vectors were then used to build a rule induction model. We applied the rule induction method to three protein families - transcription factors from the NF-kappaB, the early-growth-response (EGR), and the paired domain groups - and their corresponding DNA targets. These three prediction problems increase in complexity with regard to model building, and thus gave us a good range for validation. The main focus of the study was the most complex problem: paired domain-DNA binding. For this problem, we also found sequence/binding strength correlations using measures from information theory. Prediction results were uniformly good: the rule induction approach was able to correctly order all of nine unseen examples for one NF-kappaB protein, achieved a correlation coefficient of 0.52 on unseen (and noisy) examples in the EGR case, and reached a classification accuracy of 69.7% for the Paired domain as evaluated by cross-validation.

q-bio.BM

Role-similarity based functional prediction in networked systems: Application to the yeast proteome

We propose a general method to predict functions of vertices where: 1. The wiring of the network is somehow related to the vertex functionality. 2. A fraction of the vertices are functionally classified. The method is influenced by role-similarity measures of social network analysis. The two versions of our prediction scheme is tested on model networks were the functions of the vertices are designed to match their network surroundings. We also apply these methods to the proteome of the yeast Saccharomyces cerevisiae and find the results compatible with more specialized methods.

q-bio.MN

Discovery and analysis of biochemical subnetwork hierarchies

The representation of a biochemical network as a graph is the coarsest level of description in cellular biochemistry. By studying the network structure one can draw conclusions on the large scale organisation of the biochemical processes. We describe methods how one can extract hierarchies of subnetworks, how these can be interpreted and further deconstructed to find autonomous subnetworks. The large-scale organisation we find is characterised by a tightly connected core surrounded by increasingly loosely connected substrates.

q-bio.MN

Subnetwork hierarchies of biochemical pathways

We present a method to decompose biochemical networks into subnetworks based on the global geometry of the network. This method enables us to analyse the full hierarchical organisation of biochemical networks and is applied to 43 organisms from the WIT database. Two types of biochemical networks are considered: metabolic networks and whole-cellular networks (also including e.g. information processes). Conceptual and quantitative ways of describing the hierarchical ordering are discussed. The general picture of the metabolic networks arising from our study is that of a few core-clusters centred around the most highly connected substances enclosed by other substances in outer shells, and a few other well-defined subnetworks.

cond-mat.dis-nn