SearcharxivSearch

arXiv subjects

Gideon Rechavi

Publications and source records attributed to Gideon Rechavi.

6 recordsLinked to original sources

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

cs.AI

Gene expression analysis reveals a strong signature of an interferon induced pathway in childhood lymphoblastic leukemia as well as in breast and ovarian cancer

On the basis of epidemiological studies, infection was suggested to play a role in the etiology of human cancer. While for some cancers such a role was indeed demonstrated, there is no direct biological support for the role of viral pathogens in the pathogenesis of childhood leukemia. Using a novel bioinformatic tool, that alternates between clustering and standard statistical methods of analysis, we performed a "double blind" search of published gene expression data of subjects with different childhood ALL subtypes, looking for unanticipated partitions of patients, induced by unexpected groups of genes with correlated expression. We discovered a group of about thirty genes, related to the interferon response pathway, whose expression levels divide the ALL samples into two subgroups; high in 50, low in 285 patients. Leukemic subclasses prevalent in early childhood (the age most susceptible to infection) are over-represented in the high expression subgroup. Similar partitions, induced by the same genes, were found also in breast and ovarian cancer but not in lung cancer, prostate cancer and lymphoma. About 40% of breast cancer samples expressed the "interferon- related" signature. It is of interested that several studies demonstrated MMTV-like sequences in about 40% of breast cancer samples. Our discovery of an unanticipated strong signature of an interferon induced pathway provides molecular support for a role for either inflammation or viral infection in the pathogenesis of childhood leukemia as well as breast and ovarian cancer.

q-bio.GN

Evolutionarily conserved human targets of adenosine to inosine RNA editing

A-to-I RNA editing by ADARs is a post-transcriptional mechanism for expanding the proteomic repertoire. Genetic recoding by editing was so far observed for only a few mammalian RNAs that are predominantly expressed in nervous tissues. However, as these editing targets fail to explain the broad and severe phenotypes of ADAR1 knockout mice, additional targets for editing by ADARs were always expected. Using comparative genomics and expressed sequence analysis, we identified and experimentally verified four additional candidate human substrates for ADAR-mediated editing: FLNA, BLCAP, CYFIP2 and IGFBP7. Additionally, editing of three of these substrates was verified in the mouse while two of them were validated in chicken. Interestingly, none of these substrates encodes a receptor protein but two of them are strongly expressed in the CNS and seem important for proper nervous system function. The editing pattern observed suggests that some of the affected proteins might have altered physiological properties leaving the possibility that they can be related to the phenotypes of ADAR1 knockout mice.

q-bio.GN

Is abundant A-to-I RNA editing primate-specific?

A-To-I RNA editing is common to all eukaryotes, associated with various neurological functions. Recently, A-to-I editing was found to occur abundantly in the human transcriptome. Here we show that the frequency of A-to-I editing in humans is at least an order of magnitude higher as that of mouse, rat, chicken or fly. The extraordinary frequency of RNA editing in human is explained by the dominance of the primate-specific Alu element in the human transcriptome, which increases the number of double-stranded RNA substrates.

q-bio.GN

Systematic identification of abundant A-to-I editing sites in the human transcriptome

RNA editing by members of the double-stranded RNA-specific ADAR family leads to site-specific conversion of adenosine to inosine (A-to-I) in precursor messenger RNAs. Editing by ADARs is believed to occur in all metazoa, and is essential for mammalian development. Currently, only a limited number of human ADAR substrates are known, while indirect evidence suggests a substantial fraction of all pre-mRNAs being affected. Here we describe a computational search for ADAR editing sites in the human transcriptome, using millions of available expressed sequences. 12,723 A-to-I editing sites were mapped in 1,637 different genes, with an estimated accuracy of 95%, raising the number of known editing sites by two orders of magnitude. We experimentally validated our method by verifying the occurrence of editing in 26 novel substrates. A-to-I editing in humans primarily occurs in non-coding regions of the RNA, typically in Alu repeats. Analysis of the large set of editing sites indicates the role of editing in controlling dsRNA stability.

q-bio.GN

Design Principle of Gene Expression Used by Human Stem Cells; Implication for Pluripotency

Human embryonic stem cells (ESC) are undifferentiated and are endowed with the capacities of self renewal and pluripotential differentiation. Adult stem cells renew their own tissue, but whether they can trans-differentiate to other tissues is still controversial. To understand the genetic program that underlies the pluripotency of stem cells, we compared the transcription profile of ESC with that of progenitor/stem cells of human hematopoietic and keratinocytic origins, along with their mature cells to be viewed as snapshots along tissue differentiation. ESC gene profile show higher complexity with significantly more highly expressed genes than adult cells. We hypothesize that ESC use a strategy of expressing genes that represent various differentiation pathways and selection of only a few for continuous expression upon differentiation to a particular target. Such a strategy may be necessary for the pluripotency of ESC. The progenitors of either hematopoietic or keratinocytic cells also follow the same design principle. Using advanced clustering, we show that many of the ESC expressed genes are turned off in the progenitors/stem cells followed by a further downregulation in adult tissues. Concomitantly, genes specific to the target tissue are upregulated towards matured cells of skin or blood.

q-bio.QM