SearcharxivSearch

arXiv subjects

Gad Getz

Publications and source records attributed to Gad Getz.

11 recordsLinked to original sources

An exact tau-leaping method

The Gillespie algorithm and its extensions are commonly used for the simulation of chemical reaction networks. A limitation of these algorithms is that they have to process and update the system after every reaction, requiring significant computation. Another class of algorithms, based on the tau-leaping method, is able to simulate multiple reactions at a time at the cost of decreased accuracy. We present a new algorithm for the exact simulation of chemical reaction networks that is capable of sampling multiple reactions at a time via a first-order approximation similarly to the tau-leaping methods. We prove that the algorithm has an improved runtime complexity compared to existing methods for the exact simulation of chemical reaction networks, and present an efficient and easy to use implementation that outperforms existing methods in practice.

q-bio.MN

Towards quantum-enabled cell-centric therapeutics

In recent years, there has been tremendous progress in the development of quantum computing hardware, algorithms and services leading to the expectation that in the near future quantum computers will be capable of performing simulations for natural science applications, operations research, and machine learning at scales mostly inaccessible to classical computers. Whereas the impact of quantum computing has already started to be recognized in fields such as cryptanalysis, natural science simulations, and optimization among others, very little is known about the full potential of quantum computing simulations and machine learning in the realm of healthcare and life science (HCLS). Herein, we discuss the transformational changes we expect from the use of quantum computation for HCLS research, more specifically in the field of cell-centric therapeutics. Moreover, we identify and elaborate open problems in cell engineering, tissue modeling, perturbation modeling, and bio-topology while discussing candidate quantum algorithms for research on these topics and their potential advantages over classical computational approaches.

quant-ph

An exact method to compute a $p$-value for the beyond-pairwise correlations among cancer gene mutations

The increasing observation of mutual exclusivity correlations among cancer gene mutations is a key component for identifying driver events or pathways in cancer genome analysis. Here we report a rigorous statistical method to compute an exact $p$-value for the beyond-pairwise mutual exclusivity or co-occurrence relationships among cancer gene mutations by enumerating a null distribution of overlapping mutations across more than two genes. The validity and the advantage of our method is explicitly demonstrated in both cancer gene mutations and simulation data through the comparison to the permutation test.

q-bio.QM

High-order chromatin architecture determines the landscape of chromosomal alterations in cancer

The rapid growth of cancer genome structural information provides an opportunity for a better understanding of the mutational mechanisms of genomic alterations in cancer and the forces of selection that act upon them. Here we test the evidence for two major forces, spatial chromosome structure and purifying (or negative) selection, that shape the landscape of somatic copy-number alterations (SCNAs) in cancer1. Using a maximum likelihood framework we compare SCNA maps and three-dimensional genome architecture as determined by genome-wide chromosome conformation capture (HiC) and described by the proposed fractal-globule (FG) model2. This analysis provides evidence that the distribution of chromosomal alterations in cancer is spatially related to three-dimensional genomic architecture and additionally suggests that purifying selection as well as positive selection shapes the landscape of SCNAs during somatic evolution of cancer cells.

q-bio.GN

Semi-Supervised Learning -- A Statistical Physics Approach

We present a novel approach to semi-supervised learning which is based on statistical physics. Most of the former work in the field of semi-supervised learning classifies the points by minimizing a certain energy function, which corresponds to a minimal k-way cut solution. In contrast to these methods, we estimate the distribution of classifications, instead of the sole minimal k-way cut, which yields more accurate and robust results. Our approach may be applied to all energy functions used for semi-supervised learning. The method is based on sampling using a Multicanonical Markov chain Monte-Carlo algorithm, and has a straightforward probabilistic interpretation, which allows for soft assignments of points to classes, and also to cope with yet unseen class types. The suggested approach is demonstrated on a toy data set and on two real-life data sets of gene expression.

cs.LG

Gene expression analysis reveals a strong signature of an interferon induced pathway in childhood lymphoblastic leukemia as well as in breast and ovarian cancer

On the basis of epidemiological studies, infection was suggested to play a role in the etiology of human cancer. While for some cancers such a role was indeed demonstrated, there is no direct biological support for the role of viral pathogens in the pathogenesis of childhood leukemia. Using a novel bioinformatic tool, that alternates between clustering and standard statistical methods of analysis, we performed a "double blind" search of published gene expression data of subjects with different childhood ALL subtypes, looking for unanticipated partitions of patients, induced by unexpected groups of genes with correlated expression. We discovered a group of about thirty genes, related to the interferon response pathway, whose expression levels divide the ALL samples into two subgroups; high in 50, low in 285 patients. Leukemic subclasses prevalent in early childhood (the age most susceptible to infection) are over-represented in the high expression subgroup. Similar partitions, induced by the same genes, were found also in breast and ovarian cancer but not in lung cancer, prostate cancer and lymphoma. About 40% of breast cancer samples expressed the "interferon- related" signature. It is of interested that several studies demonstrated MMTV-like sequences in about 40% of breast cancer samples. Our discovery of an unanticipated strong signature of an interferon induced pathway provides molecular support for a role for either inflammation or viral infection in the pathogenesis of childhood leukemia as well as breast and ovarian cancer.

q-bio.GN

Design Principle of Gene Expression Used by Human Stem Cells; Implication for Pluripotency

Human embryonic stem cells (ESC) are undifferentiated and are endowed with the capacities of self renewal and pluripotential differentiation. Adult stem cells renew their own tissue, but whether they can trans-differentiate to other tissues is still controversial. To understand the genetic program that underlies the pluripotency of stem cells, we compared the transcription profile of ESC with that of progenitor/stem cells of human hematopoietic and keratinocytic origins, along with their mature cells to be viewed as snapshots along tissue differentiation. ESC gene profile show higher complexity with significantly more highly expressed genes than adult cells. We hypothesize that ESC use a strategy of expressing genes that represent various differentiation pathways and selection of only a few for continuous expression upon differentiation to a particular target. Such a strategy may be necessary for the pluripotency of ESC. The progenitors of either hematopoietic or keratinocytic cells also follow the same design principle. Using advanced clustering, we show that many of the ESC expressed genes are turned off in the progenitors/stem cells followed by a further downregulation in adult tissues. Concomitantly, genes specific to the target tissue are upregulated towards matured cells of skin or blood.

q-bio.QM

FSSP to SCOP and CATH (F2CS) Prediction Server

Summary: The F2CS server provides access to the software, F2CS2.00, that implements an automated prediction method of SCOP and CATH classifications of proteins, based on their FSSP Z-scores (Getz et al., 2002), Availability: Free, at http://www.weizmann.ac.il/physics/complex/compphys/f2cs/. Contact: eytan.domany@weizmann.ac.il Supplementary information: The site contains links to additional figures and tables.

q-bio.BM

Outcome signature genes in breast cancer: is there a unique set?

Motivation: Predicting the metastatic potential of primary malignant tissues has direct bearing on choice of therapy. Several microarray studies yielded gene sets whose expression profiles successfully predicted survival (Ramaswamy et al 2003; Sorlie et al 2001; van't Veer et al 2003). Nevertheless, the overlap between these gene sets is almost zero. Such small overlaps were observed also in other complex diseases (Lossos et al 2003; Miklos and Maleszka 2004), and the variables that could account for the differences had evoked a wide interest. One of the main open questions in this context is whether the disparity can be attributed only to trivial reasons such as different technologies, different patients and different types of analysis. Results: To answer this question we concentrated on one single breast cancer dataset, and analyzed it by one single method, the one which was used by van't Veer et al to produce a set of outcome predictive genes. We showed that in fact the resulting set of genes is not unique; it is strongly influenced by the subset of patients used for gene selection. Many equally predictive lists could have been produced from the same analysis. Three main properties of the data explain this sensitivity: (a) many genes are correlated with survival; (b) the differences between these correlations are small; (c) the correlations fluctuate strongly when measured over different subsets of patients. A possible biological explanation for these properties is discussed.

q-bio.QM

Coupled Two-Way Clustering Analysis of Breast Cancer and Colon Cancer Gene Expression Data

We present and review Coupled Two Way Clustering, a method designed to mine gene expression data. The method identifies submatrices of the total expression matrix, whose clustering analysis reveals partitions of samples (and genes) into biologically relevant classes. We demonstrate, on data from colon and breast cancer, that we are able to identify partitions that elude standard clustering analysis.

physics.bio-ph

Automated assignment of SCOP and CATH protein structure classification from FSSP scores

We present an automated procedure to assign CATH and SCOP classifications to proteins whose FSSP score is available. CATH classification is assigned down to the topology level and SCOP classification to the fold level. As the FSSP database is updated weekly, this method makes it possible to update also CATH and SCOP with the same frequency. Our predictions have a nearly perfect success rate when ambiguous cases are discarded. These ambiguous cases are intrinsic in any protein structure classification, which relies on structural information alone. Hence, we introduce the notion of ``twilight zone for structure classification''. We further suggest that in order to resolve these ambiguous cases other criteria of classification, based also on information about sequence and function, must be used.

cond-mat.soft