SearcharxivSearch

arXiv subjects

Marianne Rooman

Publications and source records attributed to Marianne Rooman.

13 recordsLinked to original sources

Engineering T7 RNA Polymerase for High-Purity In Vitro Transcription

In vitro transcription using bacteriophage T7 RNA polymerase (T7 RNAP) is the gold-standard platform for RNA production in both research and therapeutic applications. Despite its high processivity and promoter specificity, T7 RNAP generates multiple RNA by-products, including double-stranded RNA, 3'-extended transcripts, abortive RNAs, and prematurely terminated products. These impurities reduce RNA yield, complicate downstream purification, and raise safety concerns for RNA-based therapeutics by activating adverse innate immune pathways. Although reaction optimization and downstream purification strategies can mitigate these issues, they typically involve trade-offs between RNA purity and yield. Enzyme engineering has therefore emerged as a powerful upstream strategy to suppress by-product formation at its molecular origin. Here, we synthesize current knowledge on the structural and mechanistic basis of T7 RNAP by-product formation and systematically review engineering strategies to improve RNA purity. T7 RNAP variants are classified according to their underlying mechanisms of action, including enhanced thermostability, reduced non-specific template binding, smoother initiation-to-elongation transition, reduced premature termination, and template-biased polymerase designs. This analysis identifies general principles governing the trade-off between specificity and processivity and highlights synergistic combinations of mutations that improve RNA purity without compromising transcriptional efficiency. We conclude by discussing the remaining challenges for engineering T7 RNAP to meet the stringent purity requirements of next-generation RNA therapeutics.

q-bio.BM

Critical review of conformational B-cell epitope prediction methods

Accurate in-silico prediction of conformational B-cell epitopes would lead to major improvements in disease diagnostics, drug design and vaccine development. A variety of computational methods, mainly based on machine learning approaches, have been developed in the last decades to tackle this challenging problem. Here, we rigorously benchmarked nine state-of-the-art conformational B-cell epitope prediction webservers, including generic and antibody-specific methods, on a dataset of over 250 antibody-antigen structures. The results of our assessment and statistical analyses show that all the methods achieve very low performances, and some do not perform better than randomly generated patches of surface residues. In addition, we also found that commonly used consensus strategies that combine the results from multiple webservers are at best only marginally better than random. Finally, we applied all the predictors to the SARS-CoV-2 spike protein as an independent case study, and showed that they perform poorly in general, which largely recapitulates our benchmarking conclusions. We hope that these results will lead to greater caution when using these tools until the biases and issues that limit current methods have been addressed, promote the use of state-of-the-art evaluation methodologies in future publications, and suggest new strategies to improve the performance of conformational B-cell epitope prediction methods.

q-bio.BM

AI challenges for predicting the impact of mutations on protein stability

Stability is a key ingredient of protein fitness and its modification through targeted mutations has applications in various fields such as protein engineering, drug design and deleterious variant interpretation. Many studies have been devoted over the past decades to building new, more effective methods for predicting the impact of mutations on protein stability, based on the latest developments in artificial intelligence (AI). We discuss their features, algorithms, computational efficiency, and accuracy estimated on an independent test set. We focus on a critical analysis of their limitations, the recurrent biases towards the training set, their generalizability and interpretability. We found that the accuracy of the predictors has stagnated at around 1 kcal/mol for over 15 years. We conclude by discussing the challenges that need to be addressed to reach improved performance.

q-bio.MN

Modeling the molecular impact of SARS-CoV-2 infection on the renin-angiotensin system

SARS-CoV-2 coronavirus infection is mediated by the binding of its spike protein to the angiotensin-converting enzyme 2 (ACE2), which plays a pivotal role in the renin-angiotensin system (RAS). The study of RAS dysregulation due to SARS-CoV-2 infection is fundamentally important for a better understanding of the pathogenic mechanisms and risk factors associated with COVID-19 coronavirus disease, and to design effective therapeutic strategies. In this context, we developed a mathematical model of RAS based on data regarding protein and peptide concentrations; the model was tested on clinical data from healthy normotensive and hypertensive individuals. We then used our model to analyze the impact of SARS-CoV-2 infection on RAS, which we modeled through a down-regulation of ACE2 as a function of viral load. We also used it to predict the effect of RAS-targeting drugs, such as RAS-blockers, human recombinant ACE2, and angiotensin 1-7 peptide, on COVID-19 patients; the model predicted an improvement of the clinical outcome for some drugs and a worsening for others.

q-bio.MN

Deciphering noise amplification and reduction in open chemical reaction networks

The impact of random fluctuations on the dynamical behavior a complex biological systems is a longstanding issue, whose understanding would shed light on the evolutionary pressure that nature imposes on the intrinsic noise levels and would allow rationally designing synthetic networks with controlled noise. Using the Itō stochastic differential equation formalism, we performed both analytic and numerical analyses of several model systems containing different molecular species in contact with the environment and interacting with each other through mass-action kinetics. These systems represent for example biomolecular oligomerization processes, complex-breakage reactions, signaling cascades or metabolic networks. For chemical reaction networks with zero deficiency values, which admit a detailed- or complex-balanced steady state, all molecular species are uncorrelated. The number of molecules of each species follow a Poisson distribution and their Fano factors, which measure the intrinsic noise, are equal to one. Systems with deficiency one have an unbalanced non-equilibrium steady state and a non-zero S-flux, defined as the flux flowing between the complexes multiplied by an adequate stoichiometric coefficient. In this case, the noise on each species is reduced if the flux flows from the species of lowest to highest complexity, and is amplified is the flux goes in the opposite direction. These results are generalized to systems of deficiency two, which possess two independent non-vanishing S-fluxes, and we conjecture that a similar relation holds for higher deficiency systems.

q-bio.MN

Insights into the relation between noise and biological complexity

Understanding under which conditions the increase of systems complexity is evolutionary advantageous, and how this trend is related to the modulation of the intrinsic noise, are fascinating issues of utmost importance for synthetic and systems biology. To get insights into these matters, we analyzed chemical reaction networks with different topologies and degrees of complexity, interacting or not with the environment. We showed that the global level of fluctuations at the steady state, as measured by the sum of the Fano factors of the number of molecules of all species, is directly related to the topology of the network. For systems with zero deficiency, this sum is constant and equal to the rank of the network. For higher deficiencies, we observed an increase or decrease of the fluctuation levels according to the values of the reaction fluxes that link internal species, multiplied by the associated stoichiometry. We showed that the noise is reduced when the fluxes all flow towards the species of higher complexity, whereas it is amplified when the fluxes are directed towards lower complexity species.

q-bio.MN

A little walk from physical to biological complexity: protein folding and stability

As an example of topic where biology and physics meet, we present the issue of protein folding and stability, and the development of thermodynamics-based bioinformatics tools that predict the stability and thermal resistance of proteins and the change of these quantities upon amino acid substitutions. These methods are based on knowledge-driven statistical potentials, derived from experimental protein structures using the inverse Boltzmann law. We also describe an application of these predictors, which contributed to the understanding of the mechanisms of aggregation of a particular protein known to cause a neuronal disease.

q-bio.BM

Stochastic noise reduction upon complexification: positively correlated birth-death type systems

Cell systems consist of a huge number of various molecules that display specific patterns of interactions, which have a determining influence on the cell's functioning. In general, such complexity is seen to increase with the complexity of the organism, with a concomitant increase of the accuracy and specificity of the cellular processes. The question thus arises how the complexification of systems - modeled here by simple interacting birth-death type processes - can lead to a reduction of the noise - described by the variance of the number of molecules. To gain understanding of this issue, we investigated the difference between a single system containing molecules that are produced and degraded, and the same system - with the same average number of molecules - connected to a buffer. We modeled these systems using Ito stochastic differential equations in discrete time, as they allow straightforward analytical developments. In general, when the molecules in the system and the buffer are positively correlated, the variance on the number of molecules in the system is found to decrease compared to the equivalent system without a buffer. Only buffers that are too noisy by themselves tend to increase the noise in the main system. We tested this result on two model cases, in which the system and the buffer contain proteins in their active and inactive state, or protein monomers and homodimers. We found that in the second test case, where the interconversion terms are non-linear in the number of molecules, the noise reduction is much more pronounced; it reaches up to 20% reduction of the Fano factor with the parameter values tested in numerical simulations on an unperturbed birth-death model. We extended our analysis to two arbitrary interconnected systems.

q-bio.BM

Dynamic modeling of gene expression in prokaryotes: application to glucose-lactose diauxie in Escherichia coli

Coexpression of genes or, more generally, similarity in the expression profiles poses an unsurmountable obstacle to inferring the gene regulatory network (GRN) based solely on data from DNA microarray time series. Clustering of genes with similar expression profiles allows for a course-grained view of the GRN and a probabilistic determination of the connectivity among the clusters. We present a model for the temporal evolution of a gene cluster network which takes into account interactions of gene products with genes and, through a non-constant degradation rate, with other gene products. The number of model parameters is reduced by using polynomial functions to interpolate temporal data points. In this manner, the task of parameter estimation is reduced to a system of linear algebraic equations, thus making the computation time shorter by orders of magnitude. To eliminate irrelevant networks, we test each GRN for stability with respect to parameter variations, and impose restrictions on its behavior near the steady state. We apply our model and methods to DNA microarray time series' data collected on Escherichia coli during glucose-lactose diauxie and infer the most probable cluster network for different phases of the experiment.

q-bio.MN

The first peptides: the evolutionary transition between prebiotic amino acids and early proteins

The issues we attempt to tackle here are what the first peptides did look like when they emerged on the primitive earth, and what simple catalytic activities they fulfilled. We conjecture that the early functional peptides were short (3 to 8 amino acids long), were made of those amino acids, Gly, Ala, Val and Asp, that are abundantly produced in many prebiotic synthesis experiments and observed in meteorites, and that the neutralization of Asp's negative charge is achieved by metal ions. We further assume that some traces of these prebiotic peptides still exist, in the form of active sites in present-day proteins. Searching these proteins for prebiotic peptide candidates led us to identify three main classes of motifs, bound mainly to Mg^{2+} ions: D(F/Y)DGD corresponding to the active site in RNA polymerases, DGD(G/A)D present in some kinds of mutases, and DAKVGDGD in dihydroxyacetone kinase. All three motifs contain a DGD submotif, which is suggested to be the common ancestor of all active peptides. Moreover, all three manipulate phosphate groups, which was probably a very important biological function in the very first stages of life. The statistical significance of our results is supported by the frequency of these motifs in today's proteins, which is three times higher than expected by chance, with a P-value of 3 10^{-2}. The implications of our findings in the context of the appearance of life and the possibility of an experimental validation are discussed.

q-bio.BM

Noncommutative Locally Anti-de Sitter Black Holes

We give a review of our joint work on strict deformation of BHTZ 2+1 black holes \cite{BRS02,BDHRS03}. However some results presented here are not published elsewhere, and an effort is made for enlightening the instrinsical aspect of the constructions. This shows in particular that the three dimensional case treated here could be generalized to an anti-de Sitter space of arbitrary dimension provided one disposes of a universal deformation formula for the actions of a parabolic subgroup of its isometry group.

math.QA

Star products on extended massive non-rotating BTZ black holes

$AdS_3$ space-time admits a foliation by two-dimensional twisted conjugacy classes, stable under the identification subgroup yielding the non-rotating massive BTZ black hole. Each leaf constitutes a classical solution of the space-time Dirac-Born-Infeld action, describing an open D-string in $AdS_3$ or a D-string winding around the black hole. We first describe two nonequivalent maximal extensions of the non-rotating massive BTZ space-time and observe that in one of them, each D-string worldsheet admits an action of a two-parameter subgroup ($\ca \cn$) of $\SL$. We then construct non-formal, $\ca \cn$-invariant, star products that deform the classical algebra of functions on the D-string worldsheets and on their embedding space-times. We end by giving the first elements towards the definition of a Connes spectral triple on non-commutative $AdS$ space-times.

hep-th

Optimality of the genetic code with respect to protein stability and amino acid frequencies

How robust is the natural genetic code with respect to mistranslation errors? It has long been known that the genetic code is very efficient in limiting the effect of point mutation. A misread codon will commonly code either for the same amino acid or for a similar one in terms of its biochemical properties, so the structure and function of the coded protein remain relatively unaltered. Previous studies have attempted to address this question more quantitatively, namely by statistically estimating the fraction of randomly generated codes that do better than the genetic code regarding its overall robustness. In this paper, we extend these results by investigating the role of amino acid frequencies in the optimality of the genetic code. When measuring the relative fitness of the natural code with respect to a random code, it is indeed natural to assume that a translation error affecting a frequent amino acid is less favorable than that of a rare one, at equal mutation cost. We find that taking the amino acid frequency into account accordingly decreases the fraction of random codes that beat the natural code, making the latter comparatively even more robust. This effect is particularly pronounced when more refined measures of the amino acid substitution cost are used than hydrophobicity. To show this, we devise a new cost function by evaluating with computer experiments the change in folding free energy caused by all possible single-site mutations in a set of known protein structures. With this cost function, we estimate that of the order of one random code out of 100 millions is more fit than the natural code when taking amino acid frequencies into account. The genetic code seems therefore structured so as to minimize the consequences of translation errors on the 3D structure and stability of proteins.

physics.bio-ph