SearcharxivSearch

arXiv subjects

Anna Gambin

Publications and source records attributed to Anna Gambin.

8 recordsLinked to original sources

Jaccard/Tanimoto similarity test and estimation methods

Binary data are used in a broad area of biological sciences. Using binary presence-absence data, we can evaluate species co-occurrences that help elucidate relationships among organisms and environments. To summarize similarity between occurrences of species, we routinely use the Jaccard/Tanimoto coefficient, which is the ratio of their intersection to their union. It is natural, then, to identify statistically significant Jaccard/Tanimoto coefficients, which suggest non-random co-occurrences of species. However, statistical hypothesis testing using this similarity coefficient has been seldom used or studied. We introduce a hypothesis test for similarity for biological presence-absence data, using the Jaccard/Tanimoto coefficient. Several key improvements are presented including unbiased estimation of expectation and centered Jaccard/Tanimoto coefficients, that account for occurrence probabilities. We derived the exact and asymptotic solutions and developed the bootstrap and measurement concentration algorithms to compute statistical significance of binary similarity. Comprehensive simulation studies demonstrate that our proposed methods produce accurate p-values and false discovery rates. The proposed estimation methods are orders of magnitude faster than the exact solution. The proposed methods are implemented in an open source R package called jaccard (https://cran.r-project.org/package=jaccard). We introduce a suite of statistical methods for the Jaccard/Tanimoto similarity coefficient, that enable straightforward incorporation of probabilistic measures in analysis for species co-occurrences. Due to their generality, the proposed methods and implementations are applicable to a wide range of binary data arising from genomics, biochemistry, and other areas of science.

stat.ME

Assigning peaks and modeling ETD in top-down mass spectrometry

Among many techniques of modern mass spectrometry, the top down methods are becoming continuously more popular in the overall strive to describe the proteome. These techniques are based on fragmentation of ions inside mass spectrometers instead of being proteolytically digested. In some of these techniques, the fragmentation is induced by electron transfer. It can trigger several concurring reactions: electron transfer dissociation, electron transfer without dissociation, and proton transfer reaction. The evaluation of the extent of these reactions is important for the proper understanding of the functioning of the instrument and, what is even more important, to know if it can be used to reveal important structural information. We present a workflow for assigning peaks and interpreting the results of electron transfer driven reactions. We also present software written in Python and available under GNU v3 license.

stat.AP

Computational model of sphingolipids metabolism: a case study of Alzheimer's disease

Background: Sphingolipids - as suggested by the prefix in their name - are mysterious molecules, which play surprisingly various roles in opposable cellular processes, like autophagy, apoptosis, proliferation and differentiation. Recently they have been also recognized as important messengers in cellular signalling pathways. More importantly, sphingolipid metabolism disorders were observed in various pathological conditions such as cancer and neurodegeneration. Results: Existing formal models of sphingolipids metabolism concentrates mostly on de novo ceramide synthesis or restrict their focus to biochemical transformations of a particular subspecies. We propose first comprehensive computational model of sphingolipid metabolism in human tissue. In contrast to previous approaches we explicitly model compartmentalization what allows emphasizing the differences among individual organelles. Conclusions: Presented here model was validated by means of recently proposed model analysis technics allowing for detection of most sensitive and experimentally non-identifiable parameters and determination of main sources of model variance. Moreover, we demonstrate the utility of the model for the study of molecular processes underlying Alzheimer's disease.

q-bio.MN

Law of Localised Fine Structure with application in mass spectrometry

This paper presents a brand new methodology to deal with isotopic fine structure calculations. By using the Poisson approximation in an entirely novel way, we introduce mathematical elegance into the discussion on the trade-off between resolution and tractability. Our considerations unify the concepts of fine-structure, equatransneutronic configurations, and aggregate isotopic structure in a natural and simple way. We show how to boost the theoretical resolution in a seemingly costless way by several orders of magnitude with respect to the already very efficient algorithms operating on isotopic aggregates. We also develop an effective new way to obtain the important peaks in the most disaggregated isotopic structure localised in a precise region in the mass domain.

physics.chem-ph

StochDecomp - Matlab package for noise decomposition in stochastic biochemical systems

Stochasticity is an indispensable aspect of biochemical processes at the cellular level. Studies on how the noise enters and propagates in biochemical systems provided us with nontrivial insights into the origins of stochasticity, in total however they constitute a patchwork of different theoretical analyses. Here we present a flexible and generally applicable noise decomposition tool, that allows us to calculate contributions of individual reactions to the total variability of a system's output. With the package it is therefore possible to quantify how the noise enters and propagates in biochemical systems. We also demonstrate and exemplify using the JAK-STAT signalling pathway that it is possible to infer noise contributions resulting from individual reactions directly from experimental data. This is the first computational tool that allows to decompose noise into contributions resulting from individual reactions.

q-bio.QM

Modelling the efficacy of hyperthermia treatment

Multimodal oncological strategies which combine chemotherapy or radiotherapy with hyperthermia have a potential of improving the efficacy of the non-surgical methods of cancer treatment. Hyperthermia engages the heat-shock response mechanism (HSR), main component of which are heat-shock proteins (HSP). Cancer cells have already partially activated HSR, thereby, hyperthermia may be more toxic to them relative to normal cells. On the other hand, HSR triggers thermotolerance, i.e. hyperthermia treated cells show an impairment in their susceptibility to a subsequent heat-induced stress. This poses questions about efficacy and optimal strategy of the anti-cancer therapy combined with hyperthermia treatment. To address these questions, we adapt our previous HSR model and propose its stochastic extension. We formalise the notion of a HSP-induced thermotolerance. Next, we estimate the intensity and the duration of the thermotolerance. Finally, we quantify the effect of a multimodal therapy based on hyperthermia and a cytotoxic effect of bortezomib, a clinically approved proteasome inhibitor. Consequently, we propose an optimal strategy for combining hyperthermia and proteasome inhibition modalities. In summary, by a proof of concept mathematical analysis of HSR we are able to support the common belief that the combination of cancer treatment strategies increases therapy efficacy. thermotolerance.

q-bio.SC

On subset seeds for protein alignment

We apply the concept of subset seeds proposed in [1] to similarity search in protein sequences. The main question studied is the design of efficient seed alphabets to construct seeds with optimal sensitivity/selectivity trade-offs. We propose several different design methods and use them to construct several alphabets. We then perform a comparative analysis of seeds built over those alphabets and compare them with the standard BLASTP seeding method [2], [3], as well as with the family of vector seeds proposed in [4]. While the formalism of subset seeds is less expressive (but less costly to implement) than the cumulative principle used in BLASTP and vector seeds, our seeds show a similar or even better performance than BLASTP on Bernoulli models of proteins compatible with the common BLOSUM62 matrix. Finally, we perform a large-scale benchmarking of our seeds against several main databases of protein alignments. Here again, the results show a comparable or better performance of our seeds vs. BLASTP.

q-bio.QM

Efficient seeding techniques for protein similarity search

We apply the concept of subset seeds proposed in [1] to similarity search in protein sequences. The main question studied is the design of efficient seed alphabets to construct seeds with optimal sensitivity/selectivity trade-offs. We propose several different design methods and use them to construct several alphabets.We then perform an analysis of seeds built over those alphabet and compare them with the standard Blastp seeding method [2,3], as well as with the family of vector seeds proposed in [4]. While the formalism of subset seed is less expressive (but less costly to implement) than the accumulative principle used in Blastp and vector seeds, our seeds show a similar or even better performance than Blastp on Bernoulli models of proteins compatible with the common BLOSUM62 matrix.

q-bio.QM