SearcharxivSearch

arXiv subjects

Anyou Wang

Publications and source records attributed to Anyou Wang.

7 recordsLinked to original sources

Integrating Fréchet distance and AI reveals the evolutionary trajectory and origin of SARS-CoV-2

A genome, composed of a precisely ordered sequence of four nucleotides (ATCG), encompasses a multitude of specific genome features like AAA motif. Mutations occurring within a genome disrupt the sequential order and composition of these features, thereby influencing the evolutionary trajectories and yielding variants. The evolutionary relatedness between a variant and its ancestor can be estimated by assessing evolutionary distances across a spectrum of genome features. This study develops a novel, alignment-free algorithm that considers both the sequential order and composition of genome features, enabling computation of the Fréchet distance (Fr) across multiple genome features to quantify the evolutionary status of a variant. Integrating this algorithm with an artificial recurrent neural network (RNN) reveals the quantitative evolutionary trajectory and origin of SARS-CoV-2, a puzzle unsolved by alignment-based phylogenetics. The RNN generates the evolutionary trajectory from Fr data at two levels: genome sequence mutations and organism variants. At the genome sequence level, SARS-CoV-2 evolutionarily shortens its genome to enhance its infectious capacity. Mutating signature features, such as TTA and GCT, increases its infectious potential and drives its evolution. At the organism level, variants mutating a single biomarker possess low infectious potential. However, mutating multiple markers dramatically increases their infectious capacity, propelling the COVID-19 pandemic. SARS-CoV-2 likely originates from mink coronavirus variants, with its origin trajectory traced as follows: mink, cat, tiger, mouse, hamster, dog, lion, gorilla, leopard, bat, and pangolin. Together, mutating multiple signature features and biomarkers delineates the evolutionary trajectory of mink-origin SARS-CoV-2, leading to the COVID-19 pandemic. Full text and detailed on https://combai.org/ai/covidgenome/

q-bio.OT

Noncoding RNAs evolutionarily extend animal lifespan

The mechanisms underlying lifespan evolution in organisms have long been mysterious. However, recent studies have demonstrated that organisms evolutionarily gain noncoding RNAs (ncRNAs) that carry endogenous profound functions in higher organisms, including lifespan. This study unveils ncRNAs as crucial drivers driving animal lifespan evolution. Species in the animal kingdom evolutionarily increase their ncRNA length in their genomes, coinciding with trimming mitochondrial genome length. This leads to lower energy consumption and ultimately lifespan extension. Notably, during lifespan extension, species exhibit a gradual acquisition of long-life ncRNA motifs while concurrently losing short-life motifs. These longevity-associated ncRNA motifs, such as GGTGCG, are particularly active in key tissues, including the endometrium, ovary, testis, and cerebral cortex. The activation of ncRNAs in the ovary and endometrium offers insights into why women generally exhibit longer lifespans than men. This groundbreaking discovery reveals the pivotal role of ncRNAs in driving lifespan evolution and provides a fundamental foundation for the study of longevity and aging.

q-bio.PE

Noncoding RNAs and deep learning neural network discriminate multi-cancer types

Detecting cancers at early stages can dramatically reduce mortality rates. Therefore, practical cancer screening at the population level is needed. Here, we develop a comprehensive detection system to classify all common cancer types. By integrating artificial intelligence deep learning neural network and noncoding RNA biomarkers selected from massive data, our system can accurately detect cancer vs healthy object with 96.3% of AUC of ROC (Area Under Curve of a Receiver Operating Characteristic curve). Intriguinely, with no more than 6 biomarkers, our approach can easily discriminate any individual cancer type vs normal with 99% to 100% AUC. Furthermore, a comprehensive marker panel can simultaneously multi-classify all common cancers with a stable 78% of accuracy at heterological cancerous tissues and conditions. This provides a valuable framework for large scale cancer screening. The AI models and plots of results were available in https://combai.org/ai/cancerdetection/

q-bio.MN

Noncoding RNAs serve as the deadliest regulators for cancer

Cancer is one of the leading causes of human death. Many efforts have made to understand its mechanism and have further identified many proteins and DNA sequence variations as suspected targets for therapy. However, drugs targeting these targets have low success rates, suggesting the basic mechanism still remains unclear. Here, we develop a computational software combining Cox proportional-hazards model and stability-selection to unearth an overlooked, yet the most important cancer drivers hidden in massive data from The Cancer Genome Atlas (TCGA), including 11,574 RNAseq samples and clinic data. Generally, noncoding RNAs primarily regulate cancer deaths and work as the deadliest cancer inducers and repressors, in contrast to proteins as conventionally thought. Especially, processed-pseudogenes serve as the primary cancer inducers, while lincRNA and antisense RNAs dominate the repressors. Strikingly, noncoding RNAs serves as the universal strongest regulators for all cancer types although personal clinic variables such as alcohol and smoking significantly alter cancer genome. Furthermore, noncoding RNAs also work as central hubs in cancer regulatory network and as biomarkers to discriminate cancer types. Therefore, noncoding RNAs overall serve as the deadliest cancer regulators, which refreshes the basic concept of cancer mechanism and builds a novel basis for cancer research and therapy. Biological functions of pseudogenes have rarely been recognized. Here we reveal them as the most important cancer drivers for all cancer types from big data, breaking a wall to explore their biological potentials.

q-bio.GN

Systematically Dissecting the Global Mechanism of miRNA Functions in Pluripotent Stem Cells

MicroRNAs (miRNAs) critically modulate stem cell properties like pluripotency, but the fundamental mechanism remains largely unknown. This study systematically analyzes multiple-omics data and builds a systems physical network including genome-wide interactions between miRNAs and their targets to reveal the systems mechanism of miRNA functions in mouse pluripotent stem cells. Globally, miRNAs directly repress the pluripotent core factors during differentiation state. Surprisingly, during pluripotent state, the top important miRNAs do not directly regulate the pluripotent core factors as thought, but they only directly target the pluripotent signal pathways and directly repress developmental processes. Furthermore, at pluripotent state miRNAs predominately repress DNA methyltransferases, the core enzymes for DNA methylation. The decreasing methylation repressed by miRNAs in turn activates the top miRNAs and pluripotent core factors, creating an active circuit system to modulate pluripotency. MiRNAs vary their functions with different stem cell states. While miRNAs directly repress pluripotent core factors to facilitate the differentiation during cell differentiation, they also help stem cells to maintain pluripotency by activating pluripotent cores through directly repressing DNA methylation systems and primarily inhibiting development.

q-bio.MN

A quantitative system for discriminating induced pluripotent stem cells, embryonic stem cells and somatic cells

Embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs) derived from somatic cells (SCs) provide promising resources for regenerative medicine and medical research, leading to a daily identification of new cell lines. However, an efficient system to discriminate the cell lines is lacking. Here, we developed a quantitative system to discriminate the three cell types, iPSCs, ESCs and SCs. The system contains DNA-methylation biomarkers and mathematical models, including an artificial neural network and support vector machines. All biomarkers were unbiasedly selected by calculating an eigengene score derived from analysis of genome-wide DNA methylations. With 30 biomarkers, or even with as few as 3 top biomarkers, this system can discriminate SCs from ESCs and iPSCs with almost 100% accuracy, and with approximately 100 biomarkers, the system can distinguish ESCs from iPSCs with an accuracy of 95%. This robust system performs precisely with raw data without normalization as well as with converted data in which the continuous methylation levels are accounted. Strikingly, this system can even accurately predict new samples generated from different microarray platforms and the next-generation sequencing. The subtypes of cells, such as female and male iPSCs and fetal and adult SCs, can also be discriminated with this system. Thus, this quantitative system works as a novel general and accurate framework for discriminating the three cell types, iPSCs, ESCs, and SCs and this strategy supports the notion that DNA-methylation generally varies among the three cell types.

q-bio.QM

A Systemic Receptor Network Triggered by Human cytomegalovirus Entry

Virus entry is a multistep process that triggers a variety of cellular pathways interconnecting into a complex network, yet the molecular complexity of this network remains largely unsolved. Here, by employing systems biology approach, we reveal a systemic virus-entry network initiated by human cytomegalovirus (HCMV), a widespread opportunistic pathogen. This network contains all known interactions and functional modules (i.e. groups of proteins) coordinately responding to HCMV entry. The number of both genes and functional modules activated in this network dramatically declines shortly, within 25 min post-infection. While modules annotated as receptor system, ion transport, and immune response are continuously activated during the entire process of HCMV entry, those for cell adhesion and skeletal movement are specifically activated during viral early attachment, and those for immune response during virus entry. HCMV entry requires a complex receptor network involving different cellular components, comprising not only cell surface receptors, but also pathway components in signal transduction, skeletal development, immune response, endocytosis, ion transport, macromolecule metabolism and chromatin remodeling. Interestingly, genes that function in chromatin remodeling are the most abundant in this receptor system, suggesting that global modulation of transcriptions is one of the most important events in HCMV entry. Results of in silico knock out further reveal that this entire receptor network is primarily controlled by multiple elements, such as EGFR (Epidermal Growth Factor) and SLC10A1 (sodium/bile acid cotransporter family, member 1). Thus, our results demonstrate that a complex systemic network, in which components coordinating efficiently in time and space contributes to virus entry.

q-bio.MN