SearcharxivSearch

arXiv subjects

Eytan Domany

Publications and source records attributed to Eytan Domany.

At least 19 recordsLinked to original sources

Michael E. Fisher - teacher, mentor, colleague and friend: a (very) personal account

The only rational way of educating is to be an example. If one cant help it, a warning example. Albert Einstein. I had the good fortune and privilege of having Michael Fisher as my teacher, supervisor, mentor and friend. During my years as a scientist, teacher and supervisor of about one hundred students and post docs I found myself innumerable times realizing that I am following or at least trying to follow Michaels example. These pages are my attempt to convey recollections of my association with Michael, focusing on how he served as an example for me.

cond-mat.stat-mech

Temperature Integration: an efficient procedure for calculation of free energy differences

We propose a method, Temperature Integration, which allows an efficient calculation of free energy differences between two systems of interest, with the same degrees of freedom, which may have rough energy landscapes. The method is based on calculating, for each single system, the difference between the values of lnZ at two temperatures, using a Parallel Tempering procedure. If our two systems of interest have the same phase space volume, they have the same values of lnZ at high-T, and we can obtain the free energy difference between them, using the two single-system calculations described above. If the phase space volume of a system is known, our method can be used to calculate its absolute (versus relative) free energy as well. We apply our method and demonstrate its efficiency on a toy model of hard rods on a 1-dimensional ring.

cond-mat.stat-mech

On the Number of Samples Needed to Learn the Correct Structure of a Bayesian Network

Bayesian Networks (BNs) are useful tools giving a natural and compact representation of joint probability distributions. In many applications one needs to learn a Bayesian Network (BN) from data. In this context, it is important to understand the number of samples needed in order to guarantee a successful learning. Previous work have studied BNs sample complexity, yet it mainly focused on the requirement that the learned distribution will be close to the original distribution which generated the data. In this work, we study a different aspect of the learning, namely the number of samples needed in order to learn the correct structure of the network. We give both asymptotic results, valid in the large sample limit, and experimental results, demonstrating the learning behavior for feasible sample sizes. We show that structure learning is a more difficult task, compared to approximating the correct distribution, in the sense that it requires a much larger number of samples, regardless of the computational power available for the learner.

cs.LG

Ranking Under Uncertainty

Ranking objects is a simple and natural procedure for organizing data. It is often performed by assigning a quality score to each object according to its relevance to the problem at hand. Ranking is widely used for object selection, when resources are limited and it is necessary to select a subset of most relevant objects for further processing. In real world situations, the object's scores are often calculated from noisy measurements, casting doubt on the ranking reliability. We introduce an analytical method for assessing the influence of noise levels on the ranking reliability. We use two similarity measures for reliability evaluation, Top-K-List overlap and Kendall's tau measure, and show that the former is much more sensitive to noise than the latter. We apply our method to gene selection in a series of microarray experiments of several cancer types. The results indicate that the reliability of the lists obtained from these experiments is very poor, and that experiment sizes which are necessary for attaining reasonably stable Top-K-Lists are much larger than those currently available. Simulations support our analytical results.

cs.AI

FDR control with adaptive procedures and FDR monotonicity

The steep rise in availability and usage of high-throughput technologies in biology brought with it a clear need for methods to control the False Discovery Rate (FDR) in multiple tests. Benjamini and Hochberg (BH) introduced in 1995 a simple procedure and proved that it provided a bound on the expected value, $\mathit{FDR}\leq q$. Since then, many authors tried to improve the BH bound, with one approach being designing adaptive procedures, which aim at estimating the number of true null hypothesis in order to get a better FDR bound. Our two main rigorous results are the following: (i) a theorem that provides a bound on the FDR for adaptive procedures that use any estimator for the number of true hypotheses ($m_0$), (ii) a theorem that proves a monotonicity property of general BH-like procedures, both for the case where the hypotheses are independent. We also propose two improved procedures for which we prove FDR control for the independent case, and demonstrate their advantages over several available bounds, on simulated data and on a large number of gene expression data sets. Both applications are simple and involve a similar amount of computation as the original BH procedure. We compare the performance of our proposed procedures with BH and other procedures and find that in most cases we get more power for the same level of statistical significance.

stat.ME

Combining chromosomal arm status and significantly aberrant genomic locations reveals new cancer subtypes

Many types of tumors exhibit chromosomal losses or gains, as well as local amplifications and deletions. Within any given tumor type, sample specific amplifications and deletionsare also observed. Typically, a region that is aberrant in more tumors,or whose copy number change is stronger, would be considered as a more promising candidate to be biologically relevant to cancer. We sought for an intuitive method to define such aberrations and prioritize them. We define V, the volume associated with an aberration, as the product of three factors: a. fraction of patients with the aberration, b. the aberrations length and c. its amplitude. Our algorithm compares the values of V derived from real data to a null distribution obtained by permutations, and yields the statistical significance, p value, of the measured value of V. We detected genetic locations that were significantly aberrant and combined them with chromosomal arm status to create a succint fingerprint of the tumor genome. This genomic fingerprint is used to visualize the tumors, highlighting events that are co ocurring or mutually exclusive. We allpy the method on three different public array CGH datasets of Medulloblastoma and Neuroblastoma, and demonstrate its ability to detect chromosomal regions that were known to be altered in the tested cancer types, as well as to suggest new genomic locations to be tested. We identified a potential new subtype of Medulloblastoma, which is analogous to Neuroblastoma type 1.

q-bio.GN

Positional distribution of human transcription factor binding sites

We developed a method for estimating the positional distribution of transcription fac-tor (TF) binding sites using ChIP-chip data, and applied it to recently published experiments on binding sites of nine TFs; OCT4, SOX2, NANOG, HNF1A, HNF4A, HNF6, FOXA2, USF1 and CREB1. The data were obtained from a genome-wide cov-erage of promoter regions from 8kb upstream of the Transcription Start Site (TSS) to 2kb downstream. The number of target genes of each TF ranges from few hundred to several thousand. We found that for each of the nine TFs the estimated binding site distribution is closely approximated by a mixture of two components: a narrow peak, localized within 300 base pairs upstream of the TSS, and a distribution of almost uni-form density within the tested region. Using Gene Ontology and Enrichment analysis, we were able to associate (for each of the TFs studied) the target genes of both types of binding with known biological processes. Most GO terms were enriched either among the proximal targets or among those with a uniform distribution of binding sites. For example, the three stemness-related TFs have several hundred target genes that belong to "development" and "morphogenesis" whose binding sites belong to the uniform dis-tribution.

q-bio.MN

Wide-Scale Analysis of Human Functional Transcription Factor Binding Reveals a Strong Bias towards the Transcription Start Site

We introduce a novel method to screen the promoters of a set of genes with shared biological function, against a precompiled library of motifs, and find those motifs which are statistically over-represented in the gene set. The gene sets were obtained from the functional Gene Ontology (GO) classification; for each set and motif we optimized the sequence similarity score threshold, independently for every location window (measured with respect to the TSS), taking into account the location dependent nucleotide heterogeneity along the promoters of the target genes. We performed a high throughput analysis, searching the promoters (from 200bp downstream to 1000bp upstream the TSS), of more than 8000 human and 23,000 mouse genes, for 134 functional Gene Ontology classes and for 412 known DNA motifs. When combined with binding site and location conservation between human and mouse, the method identifies with high probability functional binding sites that regulate groups of biologically related genes. We found many location-sensitive functional binding events and showed that they clustered close to the TSS. Our method and findings were put to several experimental tests. By allowing a "flexible" threshold and combining our functional class and location specific search method with conservation between human and mouse, we are able to identify reliably functional TF binding sites. This is an essential step towards constructing regulatory networks and elucidating the design principles that govern transcriptional regulation of expression. The promoter region proximal to the TSS appears to be of central importance for regulation of transcription in human and mouse, just as it is in bacteria and yeast.

q-bio.MN

Non-trivial link overlap distribution in three-dimensional Ising spin glasses

We investigate the distributions of the link overlap, P(Q), in 3-dimensional Ising spin glasses. We use clustering methodology to identify a set of pairs of states from different Gibbs states, and calculate its contribution to P(Q). We find that the distribution over this set does not become trivial as the system size increases.

cond-mat.dis-nn

Semi-Supervised Learning -- A Statistical Physics Approach

We present a novel approach to semi-supervised learning which is based on statistical physics. Most of the former work in the field of semi-supervised learning classifies the points by minimizing a certain energy function, which corresponds to a minimal k-way cut solution. In contrast to these methods, we estimate the distribution of classifications, instead of the sole minimal k-way cut, which yields more accurate and robust results. Our approach may be applied to all energy functions used for semi-supervised learning. The method is based on sampling using a Multicanonical Markov chain Monte-Carlo algorithm, and has a straightforward probabilistic interpretation, which allows for soft assignments of points to classes, and also to cope with yet unseen class types. The suggested approach is demonstrated on a toy data set and on two real-life data sets of gene expression.

cs.LG

Identification of a Proliferation Gene Cluster Associated with HPV E6/E7 Expression Level and Viral DNA Load in Invasive Cervical Carcinoma

Specific HPV DNA sequences are associated with more than 90% of invasive carcinomas of the uterine cervix. Viral E6 and E7 oncogenes are key mediators in cell transformation by disrupting TP53 and RB pathways. To investigate molecular mechanisms involved in the progression of invasive cervical carcinoma, we performed a gene expression study on cases selected according to viral and clinical parameters. Using Coupled Two-Way Clustering and Sorting Points Into Neighbourhoods methods, we identified a Cervical Cancer Proliferation Cluster composed of 163 highly correlated transcripts, many of which corresponded to E2F pathway genes controlling cell proliferation, whereas no primary TP53 targets were present in this cluster. The average expression level of the genes of this cluster was higher in tumours with an early relapse than in tumours with a favourable course (P=0.026). Moreover, we found that E6/E7 mRNA expression level was positively correlated with the expression level of the cluster genes and with viral DNA load. These findings suggest that HPV E6/E7 expression level plays a key role in the progression of invasive carcinoma of the uterine cervix via the deregulation of cellular genes controlling tumour cell proliferation. HPV expression level may thus correspond to a biological marker useful for prognosis assessment and specific therapy of the disease.

q-bio.TO

The promoters of human cell cycle genes integrate signals from two tumor suppressive pathways during cellular transformation

Deciphering regulatory events that drive malignant transformation represents a major challenge for systems biology. Here we analyzed genome-wide transcription profiling of an in-vitro transformation process. We focused on a cluster of genes whose expression levels increased as a function of p53 and p16INK4A tumor suppressors inactivation. This cluster predominantly consists of cell cycle genes and constitutes a signature of a diversity of cancers. By linking expression profiles of the genes in the cluster with the dynamic behavior of p53 and p16INK4A, we identified a promoter architecture that integrates signals from the two tumor suppressive channels and that maps their activity onto distinct levels of expression of the cell cycle genes, which in turn, correspond to different cellular proliferation rates. Taking components of the mitotic spindle as an example, we experimentally verified our predictions that p53-mediated transcriptional repression of several of these novel targets is dependent on the activities of p21, NFY and E2F. Our study demonstrates how a well-controlled transformation process allows linking between gene expression, promoter architecture and activity of upstream signaling molecules.

q-bio.MN

Induction in myeloid leukemic cells of genes that are expressed in different normal tissues

Using DNA microarray and cluster analysis of expressed genes in a cloned line (M1-t-p53) of myeloid leukemic cells, we have analyzed the expression of genes that are preferentially expressed in different normal tissues. Clustering of 547 highly expressed genes in these leukemic cells showed 38 genes preferentially expressed in normal hematopoietic tissues and 122 other genes preferentially expressed in different normal non-hematopoietic tissues including neuronal tissues, muscle, liver and testis. We have also analyzed the genes whose expression in the leukemic cells changed after activation of wild-type p53 and treatment with the cytokine interleukin 6 (IL-6) or the calcium mobilizer thapsigargin (TG). Out of 620 such genes in the leukemic cells that were differentially expressed in normal tissues, clustering showed 80 genes that were preferentially expressed in hematopoietic tissues and 132 genes in different normal non-hematopietic tissues that also included neuronal tissues, muscle, liver and testis. Activation of p53 and treatment with IL-6 or TG induced different changes in the genes preferentially expressed in these normal tissues. These myeloid leukemic cells thus express genes that are expressed in normal non-hematopoietic tissues, and various treatments can reprogram these cells to induce other such non-hematopoietic genes. The results indicate that these leukemic cells share with normal hematopoietic stem cells the plasticity of differentiation to different cell types. It is suggested that this reprogramming to induce in malignant cells genes that are expressed in different normal tissues may be of clinical value in therapy.

q-bio.TO

Human cancers over express genes that are specific to a variety of normal human tissues

We have analyzed gene expression data from 3 different kinds of samples: normal human tissues, human cancer cell lines and leukemic cells from lymphoid and myeloid leukemia pediatric patients. We have searched for genes that are over expressed in human cancer and also show specific patterns of tissue-dependent expression in normal tissues. Using the expression data of the normal tissues we identified 4346 genes with a high variability of expression, and clustered these genes according to their relative expression level. Of 91 stable clusters obtained, 24 clusters included genes preferentially expressed either only in hematopoietic tissues or in hematopoietic and 1-2 other tissues; 28 clusters included genes preferentially expressed in various non-hematopoietic tissues such as neuronal, testis, liver, kidney, muscle, lung, pancreas and placenta. Analysis of the expression levels of these 2 groups of genes in the human cancer cell lines and leukemias, identified genes that were highly expressed in cancer cells but not in their normal counterparts, and were thus over expressed in the cancers. The different cancer cell lines and leukemias varied in the number and identity of these over expressed genes. The results indicate that many genes that are over expressed in human cancer cells are specific to a variety of normal tissues, including normal tissues other than those from which the cancer originated. It is suggested that this general property of cancer cells plays a major role in determining the behavior of the cancers, including their metastatic potential.

q-bio.TO

Gene expression analysis reveals a strong signature of an interferon induced pathway in childhood lymphoblastic leukemia as well as in breast and ovarian cancer

On the basis of epidemiological studies, infection was suggested to play a role in the etiology of human cancer. While for some cancers such a role was indeed demonstrated, there is no direct biological support for the role of viral pathogens in the pathogenesis of childhood leukemia. Using a novel bioinformatic tool, that alternates between clustering and standard statistical methods of analysis, we performed a "double blind" search of published gene expression data of subjects with different childhood ALL subtypes, looking for unanticipated partitions of patients, induced by unexpected groups of genes with correlated expression. We discovered a group of about thirty genes, related to the interferon response pathway, whose expression levels divide the ALL samples into two subgroups; high in 50, low in 285 patients. Leukemic subclasses prevalent in early childhood (the age most susceptible to infection) are over-represented in the high expression subgroup. Similar partitions, induced by the same genes, were found also in breast and ovarian cancer but not in lung cancer, prostate cancer and lymphoma. About 40% of breast cancer samples expressed the "interferon- related" signature. It is of interested that several studies demonstrated MMTV-like sequences in about 40% of breast cancer samples. Our discovery of an unanticipated strong signature of an interferon induced pathway provides molecular support for a role for either inflammation or viral infection in the pathogenesis of childhood leukemia as well as breast and ovarian cancer.

q-bio.GN

From finite-system entropy to entropy rate for a Hidden Markov Process

A recent result presented the expansion for the entropy rate of a Hidden Markov Process (HMP) as a power series in the noise variable $\eps$. The coefficients of the expansion around the noiseless ($\eps = 0$) limit were calculated up to 11th order, using a conjecture that relates the entropy rate of a HMP to the entropy of a process of finite length (which is calculated analytically). In this communication we generalize and prove the validity of the conjecture, and discuss the theoretical and practical consequences of our new theorem.

cs.IT

Taylor series expansions for the entropy rate of Hidden Markov Processes

Finding the entropy rate of Hidden Markov Processes is an active research topic, of both theoretical and practical importance. A recently used approach is studying the asymptotic behavior of the entropy rate in various regimes. In this paper we generalize and prove a previous conjecture relating the entropy rate to entropies of finite systems. Building on our new theorems, we establish series expansions for the entropy rate in two different regimes. We also study the radius of convergence of the two series expansions.

cs.IT