SearcharxivSearch

arXiv subjects

Rosemary Braun

Publications and source records attributed to Rosemary Braun.

12 recordsLinked to original sources

Modeling Transient Changes in Circadian Rhythms

The circadian clock can adapt itself to external cues, but the molecular mechanisms and regulatory networks governing circadian oscillations' transient adjustments are still largely unknown. Here we consider the specific case of circadian oscillations transiently responding to a temperature change. Using a framework motivated by Floquet theory, we model the mRNA expression level of the fat body from Drosophila melanogaster following a step change from 25C to 18C. Using the method we infer the adaptation rates of individual genes as they adapt to the new temperature. To deal with heteroskedastic noise and outliers present in the expression data we employ quantile regression and wild bootstrap for significance testing. Model selection with finite-sample corrected Akaike Information Criterion (AICc) is performed additionally for robust inference. We identify several genes with fast transition rates as potential sources of temperature-mediated responses in the circadian system of fruit flies, and the constructed network suggests that the proteasome may play important roles in governing these responses.

q-bio.MN

A minimal model of peripheral clocks reveals differential circadian re-entrainment in aging

The mammalian circadian system comprises a network of cell-autonomous oscillators, spanning from the central clock in the brain to peripheral clocks in other organs. These clocks are tightly coordinated to orchestrate rhythmic physiological and behavioral functions. Dysregulation of these rhythms is a hallmark of aging, yet it remains unclear how age-related changes lead to more easily disrupted circadian rhythms. Using a two-population model of coupled oscillators that integrates the central clock and the peripheral clocks, we derive simple mean-field equations that can capture many aspects of the rich behavior found in the mammalian circadian system. We focus on three age-associated effects which have been posited to contribute to circadian misalignment: attenuated input from the sympathetic pathway, reduced responsiveness to light, and a decline in the expression of neurotransmitters. We find that the first two factors can significantly impede re-entrainment of the clocks following a perturbation, while a weaker coupling within the central clock does not affect the recovery rate. Moreover, using our minimal model, we demonstrate the potential of using the feed-fast cycle as an effective intervention to accelerate circadian re-entrainment. These results highlight the importance of peripheral clocks in regulating the circadian rhythm and provide fresh insights into the complex interplay between aging and the resilience of the circadian system.

nlin.AO

fasano.franceschini.test: An Implementation of a Multidimensional KS Test in R

The Kolmogorov-Smirnov (KS) test is a nonparametric statistical test used to test for differences between univariate probability distributions. The versatility of the KS test has made it a cornerstone of statistical analysis across many scientific disciplines. However, the test proposed by Kolmogorov and Smirnov does not easily extend to multidimensional distributions. Here we present the fasano.franceschini.test package, an R implementation of a multidimensional two-sample KS test described by Fasano and Franceschini (1987). The fasano.franceschini.test package provides a test that is computationally efficient, applicable to data of any dimension and type (continuous, discrete, or mixed), and that performs competitively with similar R packages.

stat.ME

Network-based identification of disease genes in expression data: the GeneSurrounder method

The advent of high--throughput transcription profiling technologies has enabled identification of genes and pathways associated with disease, providing new avenues for precision medicine. A key challenge is to analyze this data in the context of the regulatory networks and pathways that control cellular processes, while still obtaining insights that can be used to design new diagnostic and therapeutic interventions. While classical differential expression analysis provides specific and hence targetable gene-level insights, it does not include any systems-level information. On the other hand, pathway analyses integrate systems-level information with expression data, but are often limited in their ability to indicate specific molecular targets. We introduce GeneSurrounder, an analysis method that takes into account the complex structure of interaction networks to identify specific genes that disrupt pathway activity in a disease-specific manner. GeneSurrounder integrates transcriptomic data and pathway network information in a novel two-step procedure to detect genes that (i) appear to influence the expression of other genes local to it in the network and (ii) are part of a subnetwork of differentially expressed genes. Combined, this evidence can be used to pinpoint specific genes that have a mechanistic role in the phenotype of interest. Applying GeneSurrounder to three distinct ovarian cancer studies using a global KEGG network, we show that our method is able to identify biologically relevant genes and genes missed by single-gene association tests, integrate pathway and expression data, and yield more consistent results across multiple studies of the same phenotype than competing methods.

q-bio.QM

Time-lagged Ordered Lasso for network inference

Accurate gene regulatory networks can be used to explain the emergence of different phenotypes, disease mechanisms, and other biological functions. Many methods have been proposed to infer networks from gene expression data but have been hampered by problems such as low sample size, inaccurate constraints, and incomplete characterizations of regulatory dynamics. Since expression regulation is dynamic, time-course data can be used to infer causality, but these datasets tend to be short or sparsely sampled. In addition, temporal methods typically assume that the expression of a gene at a time point depends on the expression of other genes at only the immediately preceding time point, while other methods include additional time points without any constraints to account for their temporal distance. These limitations can contribute to inaccurate networks with many missing and anomalous links. We adapted the time-lagged Ordered Lasso, a regularized regression method with temporal monotonicity constraints, for \textit{de novo} reconstruction. We also developed a semi-supervised method that embeds prior network information into the Ordered Lasso to discover novel regulatory dependencies in existing pathways. We evaluated these approaches on simulated data for a repressilator, time-course data from past DREAM challenges, and a HeLa cell cycle dataset to show that they can produce accurate networks subject to the dynamics and assumptions of the time-lagged Ordered Lasso regression.

q-bio.QM

Single nucleotide polymorphisms that modulate microRNA regulation of gene expression in tumors

Genome-wide association studies (GWAS) have identified single nucleotide polymorphisms (SNPs) associated with trait diversity and disease susceptibility, yet the functional properties of many genetic variants and their molecular interactions remains unclear. It has been hypothesized that SNPs in microRNA binding sites may disrupt gene regulation by microRNAs (miRNAs), short non-coding RNAs that bind to mRNA and downregulate the target gene. While a number of studies have been conducted to predict the location of SNPs in miRNA binding sites, to date there has been no comprehensive analysis of how SNP variants may impact miRNA regulation of genes. Here we investigate the functional properties of genetic variants and their effects on miRNA regulation of gene expression in cancer. Our analysis is motivated by the hypothesis that distinct alleles may cause differential binding (from miRNAs to mRNAs or from transcription factors to DNA) and change the expression of genes. We previously identified pathways--systems of genes conferring specific cell functions--that are dysregulated by miRNAs in cancer, by comparing miRNA-pathway associations between healthy and tumor tissue. We draw on these results as a starting point to assess whether SNPs in genes on dysregulated pathways are responsible for miRNA dysregulation of individual genes in tumors. Using an integrative analysis that incorporates miRNA expression, mRNA expression, and SNP genotype data, we identify SNPs that appear to influence the association between miRNAs and genes, which we term "regulatory QTLs (regQTLs)": loci whose alleles impact the regulation of genes by miRNAs. We describe the method, apply it to analyze four cancer types (breast, liver, lung, prostate) using data from The Cancer Genome Atlas (TCGA), and provide a tool to explore the findings.

q-bio.GN

Semi-supervised network inference using simulated gene expression dynamics

Motivation: Inferring the structure of gene regulatory networks from high--throughput datasets remains an important and unsolved problem. Current methods are hampered by problems such as noise, low sample size, and incomplete characterizations of regulatory dynamics, leading to networks with missing and anomalous links. Integration of prior network information (e.g., from pathway databases) has the potential to improve reconstructions. Results: We developed a semi--supervised network reconstruction algorithm that enables the synthesis of information from partially known networks with time course gene expression data. We adapted PLS-VIP for time course data and used reference networks to simulate expression data from which null distributions of VIP scores are generated and used to estimate edge probabilities for input expression data. By using simulated dynamics to generate reference distributions, this approach incorporates previously known regulatory relationships and links the network to the dynamics to form a semi-supervised approach that discovers novel and anomalous connections. We applied this approach to data from a sleep deprivation study with KEGG pathways treated as prior networks, as well as to synthetic data from several DREAM challenges, and find that it is able to recover many of the true edges and identify errors in these networks, suggesting its ability to derive posterior networks that accurately reflect gene expression dynamics.

q-bio.QM

Integrative analysis reveals disrupted pathways regulated by microRNAs in cancer

MicroRNAs (miRNAs) are small endogenous regulatory molecules that modulate gene expression post-transcriptionally. Although differential expression of miRNAs have been implicated in many diseases (including cancers), the underlying mechanisms of action remain unclear. Because each miRNA can target multiple genes, miRNAs may potentially have functional implications for the overall behavior of entire pathways. Here we investigate the functional consequences of miRNA dysregulation through an integrative analysis of miRNA and mRNA expression data using a novel approach that incorporates pathway information a priori. By searching for miRNA-pathway associations that differ between healthy and tumor tissue, we identify specific relationships at the systems-level which are disrupted in cancer. Our approach is motivated by the hypothesis that if a miRNA and pathway are associated, then the expression of the miRNA and the collective behavior of the genes in a pathway will be correlated. As such, we first obtain an expression-based summary of pathway activity using Isomap, a dimension reduction method which can articulate nonlinear structure in high-dimensional data. We then search for miRNAs that exhibit differential correlations with the pathway summary between phenotypes as a means of finding aberrant miRNA-pathway coregulation in tumors. We apply our method to cancer data using gene and miRNA expression datasets from The Cancer Genome Atlas (TCGA) and compare ${\sim}10^5$ miRNA-pathway relationships between healthy and tumor samples from four tissues (breast, prostate, lung, and liver). Many of the flagged pairs we identify have a biological basis for disruption in cancer.

q-bio.GN

Network Methods for Pathway Analysis of Genomic Data

Rapid advances in high-throughput technologies have led to considerable interest in analyzing genome-scale data in the context of biological pathways, with the goal of identifying functional systems that are involved in a given phenotype. In the most common approaches, biological pathways are modeled as simple sets of genes, neglecting the network of interactions comprising the pathway and treating all genes as equally important to the pathway's function. Recently, a number of new methods have been proposed to integrate pathway topology in the analyses, harnessing existing knowledge and enabling more nuanced models of complex biological systems. However, there is little guidance available to researches choosing between these methods. In this review, we discuss eight topology-based methods, comparing their methodological approaches and appropriate use cases. In addition, we present the results of the application of these methods to a curated set of ten gene expression profiling studies using a common set of pathway annotations. We report the computational efficiency of the methods and the consistency of the results across methods and studies to help guide users in choosing a method. We also discuss the challenges and future outlook for improved network analysis methodologies.

q-bio.QM

Partition Decoupling for Multi-gene Analysis of Gene Expression Profiling Data

We present the extention and application of a new unsupervised statistical learning technique--the Partition Decoupling Method--to gene expression data. Because it has the ability to reveal non-linear and non-convex geometries present in the data, the PDM is an improvement over typical gene expression analysis algorithms, permitting a multi-gene analysis that can reveal phenotypic differences even when the individual genes do not exhibit differential expression. Here, we apply the PDM to publicly-available gene expression data sets, and demonstrate that we are able to identify cell types and treatments with higher accuracy than is obtained through other approaches. By applying it in a pathway-by-pathway fashion, we demonstrate how the PDM may be used to find sets of mechanistically-related genes that discriminate phenotypes.

q-bio.QM

Pathways of Distinction Analysis: a new technique for multi-SNP analysis of GWAS data

Genome-wide association studies have become increasingly common due to advances in technology and have permitted the identification of differences in single nucleotide polymorphism (SNP) alleles that are associated with diseases. However, while typical GWAS analysis techniques treat markers individually, complex diseases are unlikely to have a single causative gene. There is thus a pressing need for multi-SNP analysis methods that can reveal system-level differences in cases and controls. Here, we present a novel multi-SNP GWAS analysis method called Pathways of Distinction Analysis (PoDA). The method uses GWAS data and known pathway-gene and gene-SNP associations to identify pathways that permit, ideally, the distinction of cases from controls. The technique is based upon the hypothesis that if a pathway is related to disease risk, cases will appear more similar to other cases than to controls for the SNPs associated with that pathway. By systematically applying the method to all pathways of potential interest, we can identify those for which the hypothesis holds true, i.e., pathways containing SNPs for which the samples exhibit greater within-class similarity than across classes. Importantly, PoDA improves on existing single-SNP and SNP-set enrichment analyses in that it does not require the SNPs in a pathway to exhibit independent main effects. This permits PoDA to reveal pathways in which epistatic interactions drives risk. In this paper, we detail the PoDA method and apply it to two GWA studies: one of breast cancer, and the other of liver cancer. The results obtained strongly suggest that there exist pathway-wide genomic differences that contribute to disease susceptibility. PoDA thus provides an analytical tool that is complementary to existing techniques and has the power to enrich our understanding of disease genomics at the systems-level.

q-bio.QM

Needles in the Haystack: Identifying Individuals Present in Pooled Genomic Data

Recent publications have described and applied a novel metric that quantifies the genetic distance of an individual with respect to two population samples, and have suggested that the metric makes it possible to infer the presence of an individual of known genotype in a sample for which only the marginal allele frequencies are known. However, the assumptions, limitations, and utility of this metric remained incompletely characterized. Here we present an exploration of the strengths and limitations of that method. In addition to analytical investigations of the underlying assumptions, we use both real and simulated genotypes to test empirically the method's accuracy. The results reveal that, when used as a means by which to identify individuals as members of a population sample, the specificity is low in several circumstances. We find that the misclassifications stem from violations of assumptions that are crucial to the technique yet hard to control in practice, and we explore the feasibility of several methods to improve the sensitivity. Additionally, we find that the specificity may still be lower than expected even in ideal circumstances. However, despite the metric's inadequacies for identifying the presence of an individual in a sample, our results suggest potential avenues for future research on tuning this method to problems of ancestry inference or disease prediction. By revealing both the strengths and limitations of the proposed method, we hope to elucidate situations in which this distance metric may be used in an appropriate manner. We also discuss the implications of our findings in forensics applications and in the protection of GWAS participant privacy.

q-bio.GN