Searcharxiv⌕ Search

arXiv subjects

Ronan M. T. Fleming

Publications and source records attributed to Ronan M. T. Fleming.

18 recordsLinked to original sources

kgsteward: a tool for building, reproducing and maintaining distributed knowledge graphs

Collaborative research projects in life sciences increasingly need to integrate private, embargoed consortium data with public reference databases in order to reach statistically meaningful interpretations. The Resource Description Framework (RDF) is well suited to this task: it facilitates the integration of heterogeneous data sources, and allows researchers to keep data and their documentation as metadata in the same place, provided the knowledge graph itself remains private during the time course of the project. Nevertheless, the development and long-term maintenance of a scientific knowledge graph remains a challenging, labour-intensive endeavour owing to the state of constant flux of most public resources. To tackle this challenge, we present kgsteward, a Python command-line tool that builds and maintains knowledge graphs inside RDF stores from a single, version-controlled configuration file. kgsteward supports multiple triplestores, keeps the local graph up-to-date with its external sources possibly already in RDF, or transformed into it on the fly, and uses SPARQL 1.1 UPDATE commands to amend further imported RDF on the fly. It can also validate the resulting graph with SPARQL queries that double as usage examples for both human users and AI agents. kgsteward has already been used in several collaborative projects at the SIB Swiss Institute of Bioinformatics, and we demonstrate its applicability in two real-world international research projects: one that builds a library of plant extracts with chemical analyses and associated bio-activities, and a second that reconciles public reference resources for human metabolic-network reconstruction.

cs.DB↗

Variational kinetics: elementary reaction kinetics via conic optimisation

Genome-scale modelling methods primarily predict reaction fluxes, whereas established high throughput experimental technologies primarily measure molecular species concentrations. This apparently paradoxical situation has arisen because implementing the non-linear constraints that represent reaction kinetic rate equations is challenging without resorting to convenient yet inaccurate approximations or to expansions that are valid only near a reference state. We present a mathematically and computationally tractable solution to this problem. First, we introduce a mathematical reformulation of established knowledge of metabolic reactions and reaction kinetics in matrix-vector notation. We then present variational kinetics, a novel approach that satisfies steady state reaction kinetics at genome scale by exponential conic optimisation. The non-linear rate law constraints are relaxed to exponential cones, which renders the feasible set convex, and satisfaction of elementary kinetics is recovered by minimising a strictly concave merit function over that set, which attains zero if, and only if, every rate law holds. We establish that a particular sequence of conic optimisation problems converges to a stationary point of this merit function, and that every such stationary point is a steady state satisfying elementary kinetics. Moiety conservation, thermodynamic constraints on elementary kinetic parameters, regularised steady states and linear optimisation of external reaction rates are each accommodated within the same conic formulation. We demonstrate the approach computationally on a genome-scale metabolic model.

q-bio.MN↗

Characterisation of conserved and reacting moieties in chemical reaction networks

A detailed understanding of biochemical networks at the molecular level is essential for studying complex cellular processes. In this paper, we provide a comprehensive description of biochemical networks by considering individual atoms and chemical bonds. To address combinatorial complexity, we introduce a well-established approach to group similar types of information within biochemical networks. A conserved moiety is a set of atoms whose association is invariant across all reactions in a network. A reacting moiety is a set of bonds that are either broken, formed, or undergo a change in bond order in at least one reaction in the network. By mathematically identifying these moieties, we establish the biological significance of conserved and reacting moieties according to the mathematical properties of the stoichiometric matrix. We also present a novel decomposition of the stoichiometric matrix based on conserved moieties. This approach bridges the gap between graph theory, linear algebra, and biological interpretation, thus opening up new horizons in the study of chemical reaction networks.

q-bio.MN↗

DEMETER: Efficient simultaneous curation of genome-scale reconstructions guided by experimental data and refined gene annotations

Motivation: Manual curation of genome-scale reconstructions is laborious, yet existing automated curation tools typically do not take species-specific experimental data and manually refined genome annotations into account. Results: We developed DEMETER, a COBRA Toolbox extension that enables the efficient simultaneous refinement of thousands of draft genome-scale reconstructions while ensuring adherence to the quality standards in the field, agreement with available experimental data, and refinement of pathways based on manually refined genome annotations. Availability: DEMETER and tutorials are available at https://github.com/opencobra/cobratoolbox.

q-bio.GN↗

Local convergence of the Levenberg-Marquardt method under Hölder metric subregularity

We describe and analyse Levenberg-Marquardt methods for solving systems of nonlinear equations. More specifically, we propose an adaptive formula for the Levenberg-Marquardt parameter and analyse the local convergence of the method under Hölder metric subregularity of the function defining the equation and Hölder continuity of its gradient mapping. Further, we analyse the local convergence of the method under the additional assumption that the Łojasiewicz gradient inequality holds. We finally report encouraging numerical results confirming the theoretical findings for the problem of computing moiety conserved steady states in biochemical reaction networks. This problem can be cast as finding a solution of a system of nonlinear equations, where the associated mapping satisfies the Łojasiewicz gradient inequality assumption.

q-bio.MN↗

Finding Zeros of Hölder Metrically Subregular Mappings via Globally Convergent Levenberg-Marquardt Methods

We present two globally convergent Levenberg-Marquardt methods for finding zeros of Hölder metrically subregular mappings that may have non-isolated zeros. The first method unifies the Levenberg- Marquardt direction and an Armijo-type line search, while the second incorporates this direction with a nonmonotone trust-region technique. For both methods, we prove the global convergence to a first-order stationary point of the associated merit function. Furthermore, the worst-case global complexity of these methods are provided, indicating that an approximate stationary point can be computed in at most $\mathcal{O}(\varepsilon^{-2})$ function and gradient evaluations, for an accuracy parameter $\varepsilon>0$. We also study the conditions for the proposed methods to converge to a zero of the associated mappings. Computing a moiety conserved steady state for biochemical reaction networks can be cast as the problem of finding a zero of a Hölder metrically subregular mapping. We report encouraging numerical results for finding a zero of such mappings derived from real-world biological data, which supports our theoretical foundations.

math.OC↗

Creation and analysis of biochemical constraint-based models: the COBRA Toolbox v3.0

COnstraint-Based Reconstruction and Analysis (COBRA) provides a molecular mechanistic framework for integrative analysis of experimental data and quantitative prediction of physicochemically and biochemically feasible phenotypic states. The COBRA Toolbox is a comprehensive software suite of interoperable COBRA methods. It has found widespread applications in biology, biomedicine, and biotechnology because its functions can be flexibly combined to implement tailored COBRA protocols for any biochemical network. Version 3.0 includes new methods for quality controlled reconstruction, modelling, topological analysis, strain and experimental design, network visualisation as well as network integration of chemoinformatic, metabolomic, transcriptomic, proteomic, and thermochemical data. New multi-lingual code integration also enables an expansion in COBRA application scope via high-precision, high-performance, and nonlinear numerical optimisation solvers for multi-scale, multi-cellular and reaction kinetic modelling, respectively. This protocol can be adapted for the generation and analysis of a constraint-based model in a wide variety of molecular systems biology scenarios. This protocol is an update to the COBRA Toolbox 1.0 and 2.0. The COBRA Toolbox 3.0 provides an unparalleled depth of constraint-based reconstruction and analysis methods.

q-bio.QM↗

ARTENOLIS: Automated Reproducibility and Testing Environment for Licensed Software

Motivation: Automatically testing changes to code is an essential feature of continuous integration. For open-source code, without licensed dependencies, a variety of continuous integration services exist. The COnstraint-Based Reconstruction and Analysis (COBRA) Toolbox is a suite of open-source code for computational modelling with dependencies on licensed software. A novel automated framework of continuous integration in a semi-licensed environment is required for the development of the COBRA Toolbox and related tools of the COBRA community. Results: ARTENOLIS is a general-purpose infrastructure software application that implements continuous integration for open-source software with licensed dependencies. It uses a master-slave framework, tests code on multiple operating systems, and multiple versions of licensed software dependencies. ARTENOLIS ensures the stability, integrity, and cross-platform compatibility of code in the COBRA Toolbox and related tools. Availability and Implementation: The continuous integration server, core of the reproducibility and testing infrastructure, can be freely accessed under artenolis.lcsb.uni.lu. The continuous integration framework code is located in the /.ci directory and at the root of the repository freely available under github.com/opencobra/cobratoolbox.

cs.SE↗

DistributedFBA.jl: High-level, high-performance flux balance analysis in Julia

Motivation: Flux balance analysis, and its variants, are widely used methods for predicting steady-state reaction rates in biochemical reaction networks. The exploration of high dimensional networks with such methods is currently hampered by software performance limitations. Results: DistributedFBA.jl is a high-level, high-performance, open-source implementation of flux balance analysis in Julia. It is tailored to solve multiple flux balance analyses on a subset or all the reactions of large and huge-scale networks, on any number of threads or nodes. Availability: The code and benchmark data are freely available on http://github.com/opencobra/COBRA.jl. The documentation can be found at http://opencobra.github.io/COBRA.jl

q-bio.QM↗

Accelerating the DC algorithm for smooth functions

We introduce two new algorithms to minimise smooth difference of convex (DC) functions that accelerate the convergence of the classical DC algorithm (DCA). We prove that the point computed by DCA can be used to define a descent direction for the objective function evaluated at this point. Our algorithms are based on a combination of DCA together with a line search step that uses this descent direction. Convergence of the algorithms is proved and the rate of convergence is analysed under the Lojasiewicz property of the objective function. We apply our algorithms to a class of smooth DC programs arising in the study of biochemical reaction networks, where the objective function is real analytic and thus satisfies the Lojasiewicz property. Numerical tests on various biochemical models clearly show that our algorithms outperforms DCA, being on average more than four times faster in both computational time and the number of iterations. Numerical experiments show that the algorithms are globally convergent to a non-equilibrium steady state of various biochemical networks, with only chemically consistent restrictions on the network topology.

math.OC↗

Reliable and efficient solution of genome-scale models of Metabolism and macromolecular Expression

Constraint-Based Reconstruction and Analysis (COBRA) is currently the only methodology that permits integrated modeling of Metabolism and macromolecular Expression (ME) at genome-scale. Linear optimization computes steady-state flux solutions to ME models, but flux values are spread over many orders of magnitude. Standard double-precision solvers may return inaccurate solutions or report that no solution exists. Exact simplex solvers are extremely slow and hence not practical for ME models that currently have 70,000 constraints and variables and will grow larger. We have developed a quadruple-precision version of our linear and nonlinear optimizer MINOS, and a solution procedure (DQQ) involving Double and Quad MINOS that achieves efficiency and reliability for ME models. DQQ enables extensive use of large, multiscale, linear and nonlinear models in systems biology and many other applications.

q-bio.MN↗

MetaboTools: A comprehensive toolbox for analysis of genome-scale metabolic models

Metabolomic data sets provide a direct read-out of cellular phenotypes and are increasingly generated to study biological questions. Our previous work revealed the potential of analyzing extracellular metabolomic data in the context of the metabolic model using constraint-based modeling. Through this work, which consists of a protocol, a toolbox, and tutorials of two use cases, we make our methods available to the broader scientific community. The protocol describes, in a step-wise manner, the workflow of data integration and computational analysis. The MetaboTools comprise the Matlab code required to complete the workflow described in the protocol. Tutorials explain the computational steps for integration of two different data sets and demonstrate a comprehensive set of methods for the computational analysis of metabolic models and stratification thereof into different phenotypes. The presented workflow supports integrative analysis of multiple omics data sets. Importantly, all analysis tools can be applied to metabolic models without performing the entire workflow. Taken together, this protocol constitutes a comprehensive guide to the intra-model analysis of extracellular metabolomic data and a resource offering a broad set of computational analysis tools for a wide biomedical and non-biomedical research community.

q-bio.MN↗

ReconMap: An interactive visualisation of human metabolism

A genome-scale reconstruction of human metabolism, Recon 2, is available but no interface exists to interactively visualise its content integrated with omics data and simulation results. We manually drew a comprehensive map, ReconMap 2.0, that is consistent with the content of Recon 2. We present it within a web interface that allows content query, visualization of custom datasets and submission of feedback to manual curators. ReconMap can be accessed via http://vmh.uni.lu, with network export in a Systems Biology Graphical Notation compliant format. A Constraint-Based Reconstruction and Analysis (COBRA) Toolbox extension to interact with ReconMap is available via https://github.com/opencobra/cobratoolbox.

q-bio.MN↗

Identification of conserved moieties in metabolic networks by graph theoretical analysis of atom transition networks

Conserved moieties are groups of atoms that remain intact in all reactions of a metabolic network. Identification of conserved moieties gives insight into the structure and function of metabolic networks and facilitates metabolic modelling. All moiety conservation relations can be represented as nonnegative integer vectors in the left null space of the stoichiometric matrix corresponding to a biochemical network. Algorithms exist to compute such vectors based only on reaction stoichiometry but their computational complexity has limited their application to relatively small metabolic networks. Moreover, the vectors returned by existing algorithms do not, in general, represent conservation of a specific moiety with a defined atomic structure. Here, we show that identification of conserved moieties requires data on reaction atom mappings in addition to stoichiometry. We present a novel method to identify conserved moieties in metabolic networks by graph theoretical analysis of their underlying atom transition networks. Our method returns the exact group of atoms belonging to each conserved moiety as well as the corresponding vector in the left null space of the stoichiometric matrix. It can be implemented as a pipeline of polynomial time algorithms. Our implementation completes in under five minutes on a metabolic network with more than 4,000 mass balanced reactions. The scalability of the method enables extension of existing applications for moiety conservation relations to genome-scale metabolic networks. We also give examples of new applications made possible by elucidating the atomic structure of conserved moieties.

q-bio.MN↗

Conditions for duality between fluxes and concentrations in biochemical networks

Mathematical and computational modelling of biochemical networks is often done in terms of either the concentrations of molecular species or the fluxes of biochemical reactions. When is mathematical modelling from either perspective equivalent to the other? Mathematical duality translates concepts, theorems or mathematical structures into other concepts, theorems or structures, in a one-to-one manner. We present a novel stoichiometric condition that is necessary and sufficient for duality between unidirectional fluxes and concentrations. Our numerical experiments, with computational models derived from a range of genome-scale biochemical networks, suggest that this flux-concentration duality is a pervasive property of biochemical networks. We also provide a combinatorial characterisation that is sufficient to ensure flux-concentration duality. That is, for every two disjoint sets of molecular species, there is at least one reaction complex that involves species from only one of the two sets. When unidirectional fluxes and molecular species concentrations are dual vectors, this implies that the behaviour of the corresponding biochemical network can be described entirely in terms of either concentrations or unidirectional fluxes.

q-bio.MN↗

Mass conserved elementary kinetics is sufficient for the existence of a non-equilibrium steady state concentration

Living systems are forced away from thermodynamic equilibrium by exchange of mass and energy with their environment. In order to model a biochemical reaction network in a non-equilibrium state one requires a mathematical formulation to mimic this forcing. We provide a general formulation to force an arbitrary large kinetic model in a manner that is still consistent with the existence of a non-equilibrium steady state. We can guarantee the existence of a non-equilibrium steady state assuming only two conditions; that every reaction is mass balanced and that continuous kinetic reaction rate laws never lead to a negative molecule concentration. These conditions can be verified in polynomial time and are flexible enough to permit one to force a system away from equilibrium. In an expository biochemical example we show how a reversible, mass balanced perpetual reaction, with thermodynamically infeasible kinetic parameters, can be used to perpetually force a kinetic model of anaerobic glycolysis in a manner consistent with the existence of a steady state. Easily testable existence conditions are foundational for efforts to reliably compute non-equilibrium steady states in genome-scale biochemical kinetic models.

q-bio.MN↗

Existence of Positive Steady States for Mass Conserving and Mass-Action Chemical Reaction Networks with a Single Terminal-Linkage Class

We establish that mass conserving single terminal-linkage networks of chemical reactions admit positive steady states regardless of network deficiency and the choice of reaction rate constants. This result holds for closed systems without material exchange across the boundary, as well as for open systems with material exchange at rates that satisfy a simple sufficient and necessary condition. Our proof uses a fixed point of a novel convex optimization formulation to find the steady state behavior of chemical reaction networks that satisfy the law of mass-action kinetics. A fixed point iteration can be used to compute these steady states, and we show that it converges for weakly reversible homogeneous systems. We report the results of our algorithm on numerical experiments.

q-bio.MN↗

A variational principle for computing nonequilibrium fluxes and potentials in genome-scale biochemical networks

We derive a convex optimization problem on a steady-state nonequilibrium network of biochemical reactions, with the property that energy conservation and the second law of thermodynamics both hold at the problem solution. This suggests a new variational principle for biochemical networks that can be implemented in a computationally tractable manner. We derive the Lagrange dual of the optimization problem and use strong duality to demonstrate that a biochemical analogue of Tellegen's theorem holds at optimality. Each optimal flux is dependent on a free parameter that we relate to an elementary kinetic parameter when mass action kinetics is assumed.

q-bio.MN↗