Searcharxiv⌕ Search

arXiv subjects

Philip E. Bourne

Publications and source records attributed to Philip E. Bourne.

At least 19 recordsLinked to original sources

Programmable Virtual Humans Toward Human Physiologically-Based Drug Discovery

Artificial intelligence (AI) has sparked immense interest in drug discovery, but most current approaches only digitize existing high-throughput experiments. They remain constrained by conventional pipelines. As a result, they do not address the fundamental challenges of predicting drug effects in humans. Similarly, biomedical digital twins, largely grounded in real-world data and mechanistic models, are tailored for late-phase drug development and lack the resolution to model molecular interactions or their systemic consequences, limiting their impact in early-stage discovery. This disconnect between early discovery and late development is one of the main drivers of high failure rates in drug discovery. The true promise of AI lies not in augmenting current experiments but in enabling virtual experiments that are impossible in the real world: testing novel compounds directly in silico in the human body. Recent advances in AI, high-throughput perturbation assays, and single-cell and spatial omics across species now make it possible to construct programmable virtual humans: dynamic, multiscale models that simulate drug actions from molecular to phenotypic levels. By bridging the translational gap, programmable virtual humans offer a transformative path to optimize therapeutic efficacy and safety earlier than ever before. This perspective introduces the concept of programmable virtual humans, explores their roles in a new paradigm of drug discovery centered on human physiology, and outlines key opportunities, challenges, and roadmaps for their realization.

cs.CY↗

The Biological Data Sustainability Paradox

Biological data in digital form has become a, if not the, driving force behind innovations in biology, medicine, and the environment. No study and no model would be complete without access to digital data (including text) collected by others and available in public repositories. With this ascent in the fundamental importance of data for reproducible scientific progress has come a troubling paradox.

q-bio.OT↗

MAD-FC: A Fold Change Visualization with Readability, Proportionality, and Symmetry

We propose a fold change visualization that demonstrates a combination of properties from log and linear plots of fold change. A useful fold change visualization can exhibit: (1) readability, where fold change values are recoverable from datapoint position; (2) proportionality, where fold change values of the same direction are proportionally distant from the point of no change; (3) symmetry, where positive and negative fold changes are equidistant to the point of no change; and (4) high dynamic range, where datapoint values are discernable across orders of magnitude. A linear visualization has readability and partial proportionality but lacks high dynamic range and symmetry (because negative direction fold changes are bound between [0, 1] while positive are between [1, $\infty$]). Log plots of fold change have partial readability, high dynamic range, and symmetry, but lack proportionality because of the log transform. We outline a new transform and visualization, named mirrored axis distortion of fold change (MAD-FC), that extends a linear visualization of fold change data to exhibit readability, proportionality, and symmetry (but still has the limited dynamic range of linear plots). We illustrate the use of MAD-FC with biomedical data using various fold change charts. We argue that MAD-FC plots may be a more useful visualization than log or linear plots for applications that require a limited dynamic range (approximately $\pm$2 orders of magnitude or $\pm$8 units in log2 space).

cs.GR↗

Contra-Analysis for Determining Negligible Effect Size in Scientific Research

Scientific experiments study interventions that show evidence of an effect size that is meaningfully large, negligibly small, or inconclusively broad. Previously, we proposed contra-analysis as a decision-making process to help determine which interventions have a meaningfully large effect by using contra plots to compare effect size across broadly related experiments. Here, we extend the use of contra plots to determine which results have evidence of negligible (near-zero) effect size. Determining if an effect size is negligible is important for eliminating alternative scientific explanations and identifying approximate independence between an intervention and the variable measured. We illustrate that contra plots can score negligible effect size across studies, inform the selection of a threshold for negligible effect based on broadly related results, and determine which results have evidence of negligible effect with a hypothesis test. No other data visualization can carry out all three of these tasks for analyzing negligible effect size. We demonstrate this analysis technique on real data from biomedical research. This new application of contra plots can differentiate statistically insignificant results with high strength (narrow and near-zero interval estimate of effect size) from those with low strength (broad interval estimate of effect size). Such a designation could help resolve the File Drawer problem in science, where statistically insignificant results are underreported because their interpretation is ambiguous and nonstandard. With our proposed procedure, results designated with negligible effect will be considered strong and publishable evidence of near-zero effect size.

stat.ME↗

Advance in Reversible Covalent Kinase Inhibitors

Reversible covalent kinase inhibitors (RCKIs) are a class of novel kinase inhibitors attracting increasing attention because they simultaneously show the selectivity of covalent kinase inhibitors, yet avoid permanent protein-modification-induced adverse effects. Over the last decade, RCKIs have been reported to target different kinases, including atypical kinases. Currently, three RCKIs are undergoing clinical trials to treat specific diseases, for example, Pemphigus, an autoimmune disorder. In this perspective, first, RCKIs are systematically summarized, including characteristics of electrophilic groups, chemical scaffolds, nucleophilic residues, and binding modes. Second, we provide insights into privileged electrophiles, the distribution of nucleophiles and hence effective design strategies for RCKIs. Finally, we provide a brief perspective on future design strategies for RCKIs, including those that target proteins other than kinases.

q-bio.MN↗

Contra-Analysis: Prioritizing Meaningful Effect Size in Scientific Research

At every phase of scientific research, scientists must decide how to allocate limited resources to pursue the research inquiries with the greatest potential. This prioritization dictates which controlled interventions are studied, awarded funding, published, reproduced with repeated experiments, investigated in related contexts, and translated for societal use. There are many factors that influence this decision-making, but interventions with larger effect size are often favored because they exert the greatest influence on the system studied. To inform these decisions, scientists must compare effect size across studies with dissimilar experiment designs to identify the interventions with the largest effect. These studies are often only loosely related in nature, using experiments with a combination of different populations, conditions, timepoints, measurement techniques, and experiment models that measure the same phenomenon with a continuous variable. We name this assessment contra-analysis and propose to use credible intervals of the relative difference in means to compare effect size across studies in a meritocracy between competing interventions. We propose a data visualization, the contra plot, that allows scientists to score and rank effect size between studies that measure the same phenomenon, aid in determining an appropriate threshold for meaningful effect, and perform hypothesis tests to determine which interventions have meaningful effect size. We illustrate the use of contra plots with real biomedical research data. Contra-analysis promotes a practical interpretation of effect size and facilitates the prioritization of scientific research.

stat.ME↗

Playing catch-up in building an open research commons

On August 2, 2021 a group of concerned scientists and US funding agency and federal government officials met for an informal discussion to explore the value and need for a well-coordinated US Open Research Commons (ORC); an interoperable collection of data and compute resources within both the public and private sectors which are easy to use and accessible to all.

cs.DL↗

The Most Difference in Means: A Statistic for the Strength of Null and Near-Zero Results

Statistical insignificance does not suggest the absence of effect, yet scientists must often use null results as evidence of negligible (near-zero) effect size to falsify scientific hypotheses. Doing so must assess a result's null strength, defined as the evidence for a negligible effect size. Such an assessment would differentiate strong null results that suggest a negligible effect size from weak null results that suggest a broad range of potential effect sizes. We propose the most difference in means ($δ_M$) as a two-sample statistic that can both quantify null strength and perform a hypothesis test for negligible effect size. To facilitate consensus when interpreting results, our statistic allows scientists to conclude that a result has negligible effect size using different thresholds with no recalculation required. To assist with selecting a threshold, $δ_M$ can also compare null strength between related results. Both $δ_M$ and the relative form of $δ_M$ outperform other candidate statistics in comparing null strength. We compile broadly related results and use the relative $δ_M$ to compare null strength across different treatments, measurement methods, and experiment models. Reporting the relative $δ_M$ may provide a technical solution to the file drawer problem by encouraging the publication of null and near-zero results.

stat.ME↗

The Least Difference in Means: A Statistic for Effect Size Strength and Practical Significance

With limited resources, scientific inquiries must be prioritized for further study, funding, and translation based on their practical significance: whether the effect size is large enough to be meaningful in the real world. Doing so must evaluate a result's effect strength, defined as a conservative assessment of practical significance. We propose the least difference in means ($δ_L$) as a two-sample statistic that can quantify effect strength and perform a hypothesis test to determine if a result has a meaningful effect size. To facilitate consensus, $δ_L$ allows scientists to compare effect strength between related results and choose different thresholds for hypothesis testing without recalculation. Both $δ_L$ and the relative $δ_L$ outperform other candidate statistics in identifying results with higher effect strength. We use real data to demonstrate how the relative $δ_L$ compares effect strength across broadly related experiments. The relative $δ_L$ can prioritize research based on the strength of their results.

stat.ME↗

Why it takes a village to manage and share data

Implementation plans for the National Institutes of Health policy for data management and sharing, which takes effect in 2023, provide an opportunity to reflect on the stakeholders, infrastructures, practice, economics, and sustainability of data sharing. Responsibility for fulfilling data sharing requirements tends to fall on principal investigators, whereas it takes a village of stakeholders to construct, manage, and sustain the necessary knowledge infrastructure for disseminating data products. Individual scientists have mixed incentives, and many disincentives to share data, all of which vary by research domain, methods, resources, and other factors. Motivations and investments for data sharing also vary widely among academic institutional stakeholders such as university leadership, research computing, libraries, and individual schools and departments. Stakeholder concerns are interdependent along many dimensions, seven of which are explored: what data to share; context and credit; discovery; methods and training; intellectual property; data science programs; and international tensions. Data sharing is not a simple matter of individual practice, but one of infrastructure, institutions, and economics. Governments, funding agencies, and international science organizations all will need to invest in commons approaches for data sharing to develop into a sustainable international ecosystem.

cs.DL↗

Exploration of Dark Chemical Genomics Space via Portal Learning: Applied to Targeting the Undruggable Genome and COVID-19 Anti-Infective Polypharmacology

Advances in biomedicine are largely fueled by exploring uncharted territories of human biology. Machine learning can both enable and accelerate discovery, but faces a fundamental hurdle when applied to unseen data with distributions that differ from previously observed ones -- a common dilemma in scientific inquiry. We have developed a new deep learning framework, called {\textit{Portal Learning}}, to explore dark chemical and biological space. Three key, novel components of our approach include: (i) end-to-end, step-wise transfer learning, in recognition of biology's sequence-structure-function paradigm, (ii) out-of-cluster meta-learning, and (iii) stress model selection. Portal Learning provides a practical solution to the out-of-distribution (OOD) problem in statistical machine learning. Here, we have implemented Portal Learning to predict chemical-protein interactions on a genome-wide scale. Systematic studies demonstrate that Portal Learning can effectively assign ligands to unexplored gene families (unknown functions), versus existing state-of-the-art methods, thereby allowing us to target previously "undruggable" proteins and design novel polypharmacological agents for disrupting interactions between SARS-CoV-2 and human proteins. Portal Learning is general-purpose and can be further applied to other areas of scientific inquiry.

q-bio.QM↗

Using the structural kinome to systematize kinase drug discovery

Kinase-targeted drug design is challenging. It requires designing inhibitors that can bind to specific kinases when all kinase catalytic domains share a common folding scaffold that binds ATP. Thus, obtaining the desired selectivity, given the whole human kinome, is a fundamental task during early-stage drug discovery. This begins with deciphering the kinase-ligand characteristics, analyzing the structure-activity relationships, and prioritizing the desired drug molecules across the whole kinome. Currently, there are more than 300 kinases with released PDB structures, which provides a substantial structural basis to gain these necessary insights. Here, we review in silico structure-based methods - notably, a function-site interaction fingerprint approach used in exploring the complete human kinome. In silico methods can be explored synergistically with multiple cell-based or protein-based assay platforms such as KINOMEscan. We conclude with new drug discovery opportunities associated with kinase signaling networks and using machine/deep learning techniques broadly referred to as structural biomedical data science.

q-bio.BM↗

A Birds-eye (Re)View of Acid-suppression Drugs, COVID-19, and the Highly Variable Literature

We consider the recent surge of information on the potential benefits of acid-suppression drugs in the context of COVID-19, with an eye on the variability (and confusion) across the reported findings--at least as regards the popular antacid famotidine. The inconsistencies reflect contradictory conclusions from independent clinical-based studies that took roughly similar approaches, in terms of experimental design (retrospective, cohort-based, etc.) and statistical analyses (propensity-score matching and stratification, etc.). The confusion has significant ramifications in choosing therapeutic interventions: e.g., do potential benefits of famotidine indicate its use in a particular COVID-19 case? Beyond this pressing therapeutic issue, conflicting information on famotidine must be resolved before its integration in ontological and knowledge graph-based frameworks, which in turn are useful in drug repurposing efforts. To begin systematically structuring the rapidly accumulating information, in the hopes of clarifying and reconciling the discrepancies, we consider the contradictory information along three proposed 'axes': (1) a context-of-disease axis, (2) a degree-of-[therapeutic]-benefit axis, and (3) a mechanism-of-action axis. We suspect that incongruencies in how these axes have been (implicitly) treated in past studies has led to the contradictory indications for famotidine and COVID-19. We also trace the evolution of information on acid-suppression agents as regards the transmission, severity, and mortality of COVID-19, given the many literature reports that have accumulated. By grouping the studies conceptually and thematically, we identify three eras in the progression of our understanding of famotidine and COVID-19. Harmonizing these findings is a key goal for both clinical standards-of-care (COVID and beyond) as well as ontological and knowledge graph-based approaches.

q-bio.TO↗

Deep Learning of Protein Structural Classes: Any Evidence for an 'Urfold'?

Recent computational advances in the accurate prediction of protein three-dimensional (3D) structures from amino acid sequences now present a unique opportunity to decipher the interrelationships between proteins. This task entails--but is not equivalent to--a problem of 3D structure comparison and classification. Historically, protein domain classification has been a largely manual and subjective activity, relying upon various heuristics. Databases such as CATH represent significant steps towards a more systematic (and automatable) approach, yet there still remains much room for the development of more scalable and quantitative classification methods, grounded in machine learning. We suspect that re-examining these relationships via a Deep Learning (DL) approach may entail a large-scale restructuring of classification schemes, improved with respect to the interpretability of distant relationships between proteins. Here, we describe our training of DL models on protein domain structures (and their associated physicochemical properties) in order to evaluate classification properties at CATH's "homologous superfamily" (SF) level. To achieve this, we have devised and applied an extension of image-classification methods and image segmentation techniques, utilizing a convolutional autoencoder model architecture. Our DL architecture allows models to learn structural features that, in a sense, 'define' different homologous SFs. We evaluate and quantify pairwise 'distances' between SFs by building one model per SF and comparing the loss functions of the models. Hierarchical clustering on these distance matrices provides a new view of protein interrelationships--a view that extends beyond simple structural/geometric similarity, and towards the realm of structure/function properties.

q-bio.BM↗

Machine Learning for Classification of Protein Helix Capping Motifs

The biological function of a protein stems from its 3-dimensional structure, which is thermodynamically determined by the energetics of interatomic forces between its amino acid building blocks (the order of amino acids, known as the sequence, defines a protein). Given the costs (time, money, human resources) of determining protein structures via experimental means such as X-ray crystallography, can we better describe and compare protein 3D structures in a robust and efficient manner, so as to gain meaningful biological insights? We begin by considering a relatively simple problem, limiting ourselves to just protein secondary structural elements. Historically, many computational methods have been devised to classify amino acid residues in a protein chain into one of several discrete secondary structures, of which the most well-characterized are the geometrically regular $α$-helix and $β$-sheet; irregular structural patterns, such as 'turns' and 'loops', are less understood. Here, we present a study of Deep Learning techniques to classify the loop-like end cap structures which delimit $α$-helices. Previous work used highly empirical and heuristic methods to manually classify helix capping motifs. Instead, we use structural data directly--including (i) backbone torsion angles computed from 3D structures, (ii) macromolecular feature sets (e.g., physicochemical properties), and (iii) helix cap classification data (from CAPS-DB)--as the ground truth to train a bidirectional long short-term memory (BiLSTM) model to classify helix cap residues. We tried different network architectures and scanned hyperparameters in order to train and assess several models; we also trained a Support Vector Classifier (SVC) to use as a baseline. Ultimately, we achieved 85% class-balanced accuracy with a deep BiLSTM model.

q-bio.BM↗

Structural insights into characterizing binding sites in EGFR kinase mutants

Over the last two decades epidermal growth factor receptor (EGFR) kinase has become an important target to treat non-small cell lung cancer (NSCLC). Currently, three generations of EGFR kinase-targeted small molecule drugs have been FDA approved. They nominally produce a response at the start of treatment and lead to a substantial survival benefit for patients. However, long-term treatment results in acquired drug resistance and further vulnerability to NSCLC. Therefore, novel EGFR kinase inhibitors that specially overcome acquired mutations are urgently needed. To this end, we carried out a comprehensive study of different EGFR kinase mutants using a structural systems pharmacology strategy. Our analysis shows that both wild-type and mutated structures exhibit multiple conformational states that have not been observed in solved crystal structures. We show that this conformational flexibility accommodates diverse types of ligands with multiple types of binding modes. These results provide insights for designing a new-generation of EGFR kinase inhibitor that combats acquired drug-resistant mutations through a multi-conformation-based drug design strategy.

q-bio.MN↗

Overview of Current Type I/II Kinase Inhibitors

Research on kinase-targeting drugs has made great strides over the last 30 years and is attracting greater attention for the treatment of yet more kinase-related diseases. Currently, 42 kinase drugs have been approved by the FDA, most of which (39) are Type I/II inhibitors. Notwithstanding these advances, it is desirable to target additional kinases for drug development as more than 200 diseases, particularly cancers, are directly associated with aberrant kinase regulation and signaling. Here, we review the extant Type I/II drugs systematically to obtain insights into the binding pocket characteristics, the associated features of Type I/II drugs, and the mechanism of action to facilitate future kinase drug design and discovery. We conclude by summarizing the main successes and limitations of targeting kinase for the development of drugs.

q-bio.MN↗

Structural biology meets data science: Does anything change?

Data science has emerged from the proliferation of digital data, coupled with advances in algorithms, software and hardware (e.g., GPU computing). Innovations in structural biology have been driven by similar factors, spurring us to ask: can these two fields impact one another in deep and hitherto unforeseen ways? We posit that the answer is yes. New biological knowledge lies in the relationships between sequence, structure, function and disease, all of which play out on the stage of evolution, and data science enables us to elucidate these relationships at scale. Here, we consider the above question from the five key pillars of data science: acquisition, engineering, analytics, visualization and policy, with an emphasis on machine learning as the premier analytics approach.

q-bio.QM↗