SearcharxivSearch

arXiv subjects

Miao Guo

Publications and source records attributed to Miao Guo.

16 recordsLinked to original sources

PyOMES: an open-source framework for biochemical process modelling

PyOMES is a Python-based, Open-source Modelling Environment for (bio)chemical process Simulation that aims to simplify the modelling of dynamic (including steady state) processes. This is done in a generalied, modular way to facilitate modelling a broad range of biological, chemical, and biochemical systems under a single modelling framework. PyOMES has been built to be accessible to a broad range of users - ranging from those with little modelling experience, such as experimentalists and students, through to more experienced power-users. Here, an introduction is provided to the PyOMES software including a summary of the design, architecture and vision. Use cases are then provided to demonstrate applicability of PyOMES to a number of (bio)chemical process modelling scenarios. Comparisons of predictions against existing benchmark software (i.e. PHREEQC) demonstrate the robustness of the package. Finally, a summary is provided of future directions for the PyOMES package - highlighting its establishment as a unified modelling framework and its potential for community-driven improvement as future developments.

q-bio.QM

Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization

Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choice, but they scale poorly as data accumulate and assume a smooth landscape that molecular and materials search spaces routinely violate. Transfer learning offers an alternative suited to this regime: it learns a representation from abundant cheap data and adapts it to sparse expensive data. Despite its use in property prediction, transfer learning has not been tested as the engine of a closed-loop optimization. Here we benchmark eleven transfer-learning surrogates against four GP methods under an identical selection rule, fidelity budget, and model size, across nine tasks spanning synthetic functions to real chemistry and materials problems. GPs win on smooth, low-dimensional functions but perform worst on molecular and materials problems, where transfer-learning surrogates reach substantially better solutions using far less computation. Because acquisition policy is held fixed across surrogates, this advantage is attributable to the surrogate itself. Uncertainty-driven exploration is not reliably beneficial, and calibration does not predict optimization performance, so greedy exploitation of the transfer-learned mean is the more robust default. Transfer learning is therefore the surrogate of choice for molecular and materials MFBO.

cs.LG

Valorisation of Fermentation Side-Stream for Waste-to-Mycoprotein: Nutrient Composition, Metabolic Insights and Process Optimisation

Fermentation-derived side streams represent an underutilised resource for sustainable protein production. This study investigates the potential of centrate from industrial Fusarium venenatum fermentation as a nutrient source for fungal biomass generation. Following compositional characterisation, a synthetic centrate medium was formulated and evaluated using a Box-Behnken design combined with response surface methodology. Across 46 experimental runs, cell dry weight (CDW) ranged from 0.22 to 3.87 g per liter, demonstrating a strong dependence on nutrient composition. Ammonia and glucose were identified as the dominant factors influencing biomass production, with significant nonlinear effects. The model predicted a maximum CDW of 4.17 g per liter under optimised conditions, which was experimentally validated at 3.99 g per liter. Carbon conversion efficiency reached up to 29.02%, indicating effective substrate utilisation. These findings demonstrate that fermentation-derived centrate can support substantial fungal growth, while highlighting its potential to enhance nutrient recovery and influence the biochemical composition of sustainable mycoprotein.

eess.SY

Kinetics of Mycoprotein Production from Alternative Carbon Substrates

High throughput screening was used to study of the biokinetics of F. venenatum A3/5 cultivation on alternative carbon substrates, including monosaccharides, disaccharides and mixtures relevant to food & beverage, dairy and agricultural waste streams. Expired functional drink from the beverage sector was also assessed as the primary carbon source for mycoprotein production. Growth data was analysed using modified single and multiphase Gompertz models for comparison of maximum specific growth rate and progression milestones across diverse growth regimes. Time-series substrate and byproduct data was analysed using comparative metrics, providing an explanatory basis for the different growth phenotypes observed. Substrate type strongly influenced the apparent carbon allocation strategies, with rapidly consumed sugars such as glucose and sucrose supporting high growth rates, low biomass yield and a high degree of fermentative byproduct formation. Fructose and xylose cultivations led to slower overall growth but higher biomass yield and lower byproduct formation. Galactose and lactose showed distinct dynamics that suggested co-existence of transport and metabolic induction limitations. In all dual-substrate systems, sequential utilisation was observed. However, metabolic inheritance and environmental shift effects were highlighted as potential kinetic limitations. These conditions exhibited stunted diauxic growth and low yield from secondary sugars, with glucose-dominated primary growth significantly reshaping secondary substrate efficiencies relative to their study in silo. The expired functional drink supported highly rapid growth and achieved the highest maximum specific growth rate and biomass titre of all conditions examined, alongside reduced fermentative overflow and enhanced ethanol reassimilation relative to a compositionally matched synthetic control.

physics.bio-ph

Literature Mining System for Nutraceutical Biosynthesis: From AI Framework to Biological Insight

The extraction of structured knowledge from scientific literature remains a major bottleneck in nutraceutical research, particularly when identifying microbial strains involved in compound biosynthesis. This study presents a domain-adapted system powered by large language models (LLMs) and guided by advanced prompt engineering techniques to automate the identification of nutraceutical-producing microbes from unstructured scientific text. By leveraging few-shot prompting and tailored query designs, the system demonstrates robust performance across multiple configurations, with DeepSeekV3 outperforming LLaMA2 in accuracy, especially when domain-specific strain information is included. A structured and validated dataset comprising 35 nutraceutical-strain associations was generated, spanning amino acids, fibers, phytochemicals, and vitamins. The results reveal significant microbial diversity across monoculture and co-culture systems, with dominant contributions from Corynebacterium glutamicum, Escherichia coli, and Bacillus subtilis, alongside emerging synthetic consortia. This AI-driven framework not only enhances the scalability and interpretability of literature mining but also provides actionable insights for microbial strain selection, synthetic biology design, and precision fermentation strategies in the production of high-value nutraceuticals.

q-bio.QM

Beyond the Expiry Date: Uncovering Hidden Value in Functional Drink Waste for a Circular Future

Expired functional drinks have great valorisation potential due to the high concentration of organic molecules present. However, detailed information of the resources in these expired functional drinks is limited, hindering the rational design of a recovery system. To address this gap, we present here a study that comprehensively characterises the chemical composition of functional drinks and discus their potential use as feedstocks for biomethane production. The example functional drinks were abundant in sugars, organic acids, and amino acids, and were especially rich in glucose, fructose, and alanine. Our studies revealed that functional drinks with high COD values that corresponded to high proportions of sugar and organic acid and low proportions of sorbitol and amino acids could realise profitable recovery through anaerobic digestion, with a minimum biomethane yield of 11.72 mL CH4 / mL drink. To assess utility further we also examined the dynamic composition of functional drinks up to 16 weeks (at 4 {\deg}C) after expiration to capture the shift in resources during deterioration. In doing so, we identified 4 distinct periods of carbon resource variation: 1) chemically stable period, 2) sorbitol degradation period, 3) sugar degradation period, and 4) acidification period. Based on the time-course biomethane production experiments for expired functional drinks, the optimal operating time window for biomethane production from drinks without ascorbic acid would be after sorbitol degradation period in terms of its economic performance through convenient natural deterioration. Therefore, this comprehensive study on dynamic chemical composition in expired functional drinks and their biomethane production potential could facilitate a rational design of resource recovery system for soft drink field.

eess.SY

Prediction, Generation of WWTPs microbiome community structures and Clustering of WWTPs various feature attributes using DE-BP model, SiTime-GAN model and DPNG-EPMC ensemble clustering algorithm with modulation of microbial ecosystem health

Microbiomes not only underpin Earth's biogeochemical cycles but also play crucial roles in both engineered and natural ecosystems, such as the soil, wastewater treatment, and the human gut. However, microbiome engineering faces significant obstacles to surmount to deliver the desired improvements in microbiome control. Here, we use the backpropagation neural network (BPNN), optimized through differential evolution (DE-BP), to predict the microbial composition of activated sludge (AS) systems collected from wastewater treatment plants (WWTPs) located worldwide. Furthermore, we introduce a novel clustering algorithm termed Directional Position Nonlinear Emotional Preference Migration Behavior Clustering (DPNG-EPMC). This method is applied to conduct a clustering analysis of WWTPs across various feature attributes. Finally, we employ the Similar Time Generative Adversarial Networks (SiTime-GAN), to synthesize novel microbial compositions and feature attributes data. As a result, we demonstrate that the DE-BP model can provide superior predictions of the microbial composition. Additionally, we show that the DPNG-EPMC can be applied to the analysis of WWTPs under various feature attributes. Finally, we demonstrate that the SiTime-GAN model can generate valuable incremental synthetic data. Our results, obtained through predicting the microbial community and conducting analysis of WWTPs under various feature attributes, develop an understanding of the factors influencing AS communities.

cs.LG

Comparison of Optimised Geometric Deep Learning Architectures, over Varying Toxicological Assay Data Environments

Geometric deep learning is an emerging technique in Artificial Intelligence (AI) driven cheminformatics, however the unique implications of different Graph Neural Network (GNN) architectures are poorly explored, for this space. This study compared performances of Graph Convolutional Networks (GCNs), Graph Attention Networks (GATs) and Graph Isomorphism Networks (GINs), applied to 7 different toxicological assay datasets of varying data abundance and endpoint, to perform binary classification of assay activation. Following pre-processing of molecular graphs, enforcement of class-balance and stratification of all datasets across 5 folds, Bayesian optimisations were carried out, for each GNN applied to each assay dataset (resulting in 21 unique Bayesian optimisations). Optimised GNNs performed at Area Under the Curve (AUC) scores ranging from 0.728-0.849 (averaged across all folds), naturally varying between specific assays and GNNs. GINs were found to consistently outperform GCNs and GATs, for the top 5 of 7 most data-abundant toxicological assays. GATs however significantly outperformed over the remaining 2 most data-scarce assays. This indicates that GINs are a more optimal architecture for data-abundant environments, whereas GATs are a more optimal architecture for data-scarce environments. Subsequent analysis of the explored higher-dimensional hyperparameter spaces, as well as optimised hyperparameter states, found that GCNs and GATs reached measurably closer optimised states with each other, compared to GINs, further indicating the unique nature of GINs as a GNN algorithm.

q-bio.QM

Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges

The global demand for sustainable protein sources has accelerated the need for intelligent tools that can rapidly process and synthesise domain-specific scientific knowledge. In this study, we present a proof-of-concept multi-agent Artificial Intelligence (AI) framework designed to support sustainable protein production research, with an initial focus on microbial protein sources. Our Retrieval-Augmented Generation (RAG)-oriented system consists of two GPT-based LLM agents: (1) a literature search agent that retrieves relevant scientific literature on microbial protein production for a specified microbial strain, and (2) an information extraction agent that processes the retrieved content to extract relevant biological and chemical information. Two parallel methodologies, fine-tuning and prompt engineering, were explored for agent optimisation. Both methods demonstrated effectiveness at improving the performance of the information extraction agent in terms of transformer-based cosine similarity scores between obtained and ideal outputs. Mean cosine similarity scores were increased by up to 25%, while universally reaching mean scores of $\geq 0.89$ against ideal output text. Fine-tuning overall improved the mean scores to a greater extent (consistently of $\geq 0.94$) compared to prompt engineering, although lower statistical uncertainties were observed with the latter approach. A user interface was developed and published for enabling the use of the multi-agent AI system, alongside preliminary exploration of additional chemical safety-based search capabilities

cs.AI

Temporal Dynamics of Microbial Communities in Anaerobic Digestion: Influence of Temperature and Feedstock Composition on Reactor Performance and Stability

Anaerobic digestion (AD) offers a sustainable biotechnology to recover resources from carbon-rich wastewater, such as food-processing wastewater. Despite crude wastewater characterisation, the impact of detailed chemical fingerprinting on AD remains underexplored. This study investigated the influence of fermentation-wastewater composition and operational parameters on AD over time to identify critical factors influencing reactor biodiversity and performance. Eighteen reactors were operated under various operational conditions using mycoprotein fermentation wastewater. Detailed chemical analysis fingerprinted the molecules in the fermentation wastewater throughout AD including sugars, sugar alcohols and volatile fatty acids (VFAs). Sequencing revealed distinct microbiome profiles linked to temperature and reactor configuration, with mesophilic conditions supporting a more diverse and densely connected microbiome. Significant elevations in Methanomassiliicoccus were correlated to high butyric acid concentrations and decreased biogas production, further elucidating the role of this newly discovered methanogen. Dissimilarity analysis demonstrated the importance of individual molecules on microbiome diversity, highlighting the need for detailed chemical fingerprinting in AD studies of microbial trends. Machine learning (ML) models predicting reactor performance achieved high accuracy based on operational parameters and microbial taxonomy. Operational parameters had the most substantial influence on chemical oxygen demand removal, whilst Oscillibacter and two Clostridium sp. were highlighted as key factors in biogas production. By integrating detailed chemical and biological fingerprinting with ML models this research presents a novel approach to advance our understanding of AD microbial ecology, offering insights for industrial applications of sustainable waste-to-energy systems.

q-bio.QM

A KAN-based Interpretable Framework for Process-Informed Prediction of Global Warming Potential

Accurate prediction of Global Warming Potential (GWP) is essential for assessing the environmental impact of chemical processes and materials. Traditional GWP prediction models rely predominantly on molecular structure, overlooking critical process-related information. In this study, we present an integrative GWP prediction model that combines molecular descriptors (MACCS keys and Mordred descriptors) with process information (process title, description, and location) to improve predictive accuracy and interpretability. Using a deep neural network (DNN) model, we achieved an R-squared of 86% on test data with Mordred descriptors, process location, and description information, representing a 25% improvement over the previous benchmark of 61%; XAI analysis further highlighted the significant role of process title embeddings in enhancing model predictions. To enhance interpretability, we employed a Kolmogorov-Arnold Network (KAN) to derive a symbolic formula for GWP prediction, capturing key molecular and process features and providing a transparent, interpretable alternative to black-box models, enabling users to gain insights into the molecular and process factors influencing GWP. Error analysis showed that the model performs reliably in densely populated data ranges, with increased uncertainty for higher GWP values. This analysis allows users to manage prediction uncertainty effectively, supporting data-driven decision-making in chemical and process design. Our results suggest that integrating both molecular and process-level information in GWP prediction models yields substantial gains in accuracy and interpretability, offering a valuable tool for sustainability assessments. Future work may extend this approach to additional environmental impact categories and refine the model to further enhance its predictive reliability.

cs.LG

Exploiting Dependency-Aware Priority Adjustment for Mixed-Criticality TSN Flow Scheduling

Time-Sensitive Networking (TSN) serves as a one-size-fits-all solution for mixed-criticality communication, in which flow scheduling is vital to guarantee real-time transmissions. Traditional approaches statically assign priorities to flows based on their associated applications, resulting in significant queuing delays. In this paper, we observe that assigning different priorities to a flow leads to varying delays due to different shaping mechanisms applied to different flow types. Leveraging this insight, we introduce a new scheduling method in mixed-criticality TSN that incorporates a priority adjustment scheme among diverse flow types to mitigate queuing delays and enhance schedulability. Specifically, we propose dependency-aware priority adjustment algorithms tailored to different link-overlapping conditions. Experiments in various settings validate the effectiveness of the proposed method, which enhances the schedulability by 20.57% compared with the SOTA method.

cs.NI

High Throughput Parameter Estimation and Uncertainty Analysis Applied to the Production of Mycoprotein from Synthetic Lignocellulosic Hydrolysates

The current global food system produces substantial waste and carbon emissions while exacerbating the effects of global hunger and protein deficiency. This study aims to address these challenges by exploring the use of lignocellulosic agricultural residues as feedstocks for microbial protein fermentation, focusing on Fusarium venenatum A3/5, a mycelial strain known for its high protein yield and quality. We propose a high throughput microlitre batch fermentation system paired with analytical chemistry to generate time-series data of microbial growth and substrate utilisation. An unstructured biokinetic model was developed using a bootstrap sampling approach to quantify uncertainty in the parameter estimates. The model was validated against an independent dataset of a different glucose-xylose composition to assess the predictive performance. Our results indicate a robust model fit with high coefficients of determination and low root mean squared errors for biomass, glucose, and xylose concentrations. Estimated parameter values provided insights into the resource utilisation strategies of Fusarium venenatum A3/5 in mixed substrate cultures, aligning well with previous research findings. Significant correlations between estimated parameters were observed, highlighting challenges in parameter identifiability. This work provides a foundational model for optimising the production of microbial protein from lignocellulosic waste, contributing to a more sustainable global food system.

physics.bio-ph

Surrogate-based optimisation of process systems to recover resources from wastewater

Wastewater systems are transitioning towards integrative process systems to recover multiple resources whilst simultaneously satisfying regulations on final effluent quality. This work contributes to the literature by bringing a systems-thinking approach to resource recovery from wastewater, harnessing surrogate modelling and mathematical optimisation techniques to highlight holistic process systems. A surrogate-based process synthesis methodology was presented to harness high-fidelity data from black box process simulations, embedding first principles models, within a superstructure optimisation framework. Modelling tools were developed to facilitate tailored derivative-free optimisation solutions widely applicable to black box optimisation problems. The optimisation of a process system to recover energy and nutrients from a brewery wastewater reveals significant scope to reduce the environmental impacts of food and beverage production systems. Additionally, the application demonstrates the capabilities of the modelling methodology to highlight optimal processes to recover carbon, nitrogen, and phosphorous resources whilst also accounting for uncertainties inherent to wastewater systems.

eess.SY

A sustainable waste-to-protein system to maximise waste resource utilisation for developing food- and feed-grade protein solutions

A waste-to-protein system that integrates a range of waste-to-protein upgrading technologies has the potential to converge innovations on zero-waste and protein security to ensure a sustainable protein future. We present a global overview of food-safe and feed-safe waste resource potential and technologies to sort and transform such waste streams with compositional quality characteristics into food-grade or feed-grade protein. The identified streams are rich in carbon and nutrients and absent of pathogens and hazardous contaminants, including food waste streams, lignocellulosic waste from agricultural residues and forestry, and contaminant-free waste from the food and drink industry. A wide range of chemical, physical, and biological treatments can be applied to extract nutrients and convert waste-carbon to fermentable sugars or other platform chemicals for subsequent conversion to protein. Our quantitative analyses suggest that the waste-to-protein system has the potential to maximise recovery of various low-value resources and catalyse the transformative solutions toward a sustainable protein future. However, novel protein regulation processes remain expensive and resource intensive in many countries, with protracted timelines for approval. This poses a significant barrier to market expansion, despite accelerated research and development in waste-to-protein technologies and novel protein sources. Thus, the waste-to-protein system is an important initiative to promote metabolic health across the lifespan and tackle the global hunger crisis.

physics.soc-ph

Optimisation of Wastewater Treatment Strategies in Eco-Industrial Parks: Technology, Location and Transport

The expanding population and rapid urbanisation, in particular in the Global South, are leading to global challenges on resource supply stress and rising waste generation. A transformation to resource-circular systems and sustainable recovery of carbon-containing and nutrient-rich waste offers a way to tackle such challenges. Eco-industrial parks have the potential to capture symbioses across individual waste producers, leading to more effective waste-recovery schemes. With whole-system design, economically attractive approaches can be achieved, reducing the nvironmental impacts while increasing the recovery of high-value resources. In this paper, an optimisation framework is developed to enable such design, allowing for wide ranging treatment options to be considered capturing both technological and fnancial detail. As well as technology selection, the framework also accounts for spatial aspects, with the design of suitable transport networks playing a key role. A range of scenarios are investigated using the network, highlighting the multi-faceted nature of the problem. The need to incorporate the impact of resource recovery at the design stage is shown to be of particular importance.

math.OC