Searcharxiv⌕ Search

arXiv subjects

Niall M Mangan

Publications and source records attributed to Niall M Mangan.

3 recordsLinked to original sources

Ill-Conditioning in Dictionary-Based Dynamic-Equation Learning: A Systems Biology Case Study

Data-driven discovery of governing equations from time-series data provides a powerful framework for understanding complex biological systems. Library-based approaches that use sparse regression over candidate functions have shown considerable promise, but they face a critical challenge when candidate functions become strongly correlated: numerical ill-conditioning. Poor or restricted sampling, together with particular choices of candidate libraries, can produce strong multicollinearity and numerical instability. In such cases, measurement noise may lead to widely different recovered models, obscuring the true underlying dynamics and hindering accurate system identification. Although sparse regularization promotes parsimonious solutions and can partially mitigate conditioning issues, strong correlations may persist, regularization may bias the recovered models, and the regression problem may remain highly sensitive to small perturbations in the data. We present a systematic analysis of how ill-conditioning affects sparse identification of biological dynamics using benchmark models from systems biology. We show that combinations involving as few as two or three terms can already exhibit strong multicollinearity and extremely large condition numbers. We further show that orthogonal polynomial bases do not consistently resolve ill-conditioning and can perform worse than monomial libraries when the data distribution deviates from the weight function associated with the orthogonal basis. Finally, we demonstrate that when data are sampled from distributions aligned with the appropriate weight functions corresponding to the orthogonal basis, numerical conditioning improves, and orthogonal polynomial bases can yield improved model recovery accuracy across two baseline models.

q-bio.QM↗

Mathematical modeling of 1,2-propanediol utilization bacterial microcompartments in vivo activity

On exposure to 1,2-propanediol (1,2-PD), Salmonella enterica serovar Typhimurium LT2 produces 1,2-PD utilization (Pdu) microcompartments (MCPs), nanoscale protein-bound shells that encapsulate metabolic enzymes. MCPs serve as a bioengineering platform to study reaction organization and enhance flux through specific pathways. However, a recently published assay of purified wild-type (WT) MCPs reported metabolic activity that differed markedly from that observed in vivo. Using kinetic modeling, we attribute these discrepancies to in vivo cell growth and to the cytosolic presence of MCP-associated enzymes and promiscuous alcohol dehydrogenases, which are not present in the purified MCPs. Assays of purified MCPs in E. coli lysate, together with an LT2 growth assay in which the native Pdu MCP-associated alcohol dehydrogenase, PduQ, was knocked out, support the conclusion that exogenous Pdu cytosolic enzyme activity can narrow the gap between in vitro and in vivo experiments. Our modeling further suggests that MCP-localized enzymes contribute little to in vivo metabolic flux downstream of PduCDE. We therefore propose a revised in vivo model of WT growth on 1,2-PD in which PduCDE is fully encapsulated, while much of the downstream Pdu activity occurs in the cytosol.

q-bio.MN↗

Model selection for hybrid dynamical systems via sparse regression

Hybrid systems are traditionally difficult to identify and analyze using classical dynamical systems theory. Moreover, recently developed model identification methodologies largely focus on identifying a single set of governing equations solely from measurement data. In this article, we develop a new methodology, Hybrid-Sparse Identification of Nonlinear Dynamics (Hybrid-SINDy), which identifies separate nonlinear dynamical regimes, employs information theory to manage uncertainty, and characterizes switching behavior. Specifically, we utilize the nonlinear geometry of data collected from a complex system to construct a set of coordinates based on measurement data and augmented variables. Clustering the data in these measurement-based coordinates enables the identification of nonlinear hybrid systems. This methodology broadly empowers nonlinear system identification without constraining the data locally in time and has direct connections to hybrid systems theory. We demonstrate the success of this method on numerical examples including a mass-spring hopping model and an infectious disease model. Characterizing complex systems that switch between dynamic behaviors is integral to overcoming modern challenges such as eradication of infectious diseases, the design of efficient legged robots, and the protection of cyber infrastructures.

math.DS↗