SearcharxivSearch

arXiv subjects

David A. Winkler

Publications and source records attributed to David A. Winkler.

6 recordsLinked to original sources

EFI: A Toolbox for Feature Importance Fusion and Interpretation in Python

This paper presents an open-source Python toolbox called Ensemble Feature Importance (EFI) to provide machine learning (ML) researchers, domain experts, and decision makers with robust and accurate feature importance quantification and more reliable mechanistic interpretation of feature importance for prediction problems using fuzzy sets. The toolkit was developed to address uncertainties in feature importance quantification and lack of trustworthy feature importance interpretation due to the diverse availability of machine learning algorithms, feature importance calculation methods, and dataset dependencies. EFI merges results from multiple machine learning models with different feature importance calculation approaches using data bootstrapping and decision fusion techniques, such as mean, majority voting and fuzzy logic. The main attributes of the EFI toolbox are: (i) automatic optimisation of ML algorithms, (ii) automatic computation of a set of feature importance coefficients from optimised ML algorithms and feature importance calculation techniques, (iii) automatic aggregation of importance coefficients using multiple decision fusion techniques, and (iv) fuzzy membership functions that show the importance of each feature to the prediction task. The key modules and functions of the toolbox are described, and a simple example of their application is presented using the popular Iris dataset.

cs.LG

Mechanistic Interpretation of Machine Learning Inference: A Fuzzy Feature Importance Fusion Approach

With the widespread use of machine learning to support decision-making, it is increasingly important to verify and understand the reasons why a particular output is produced. Although post-training feature importance approaches assist this interpretation, there is an overall lack of consensus regarding how feature importance should be quantified, making explanations of model predictions unreliable. In addition, many of these explanations depend on the specific machine learning approach employed and on the subset of data used when calculating feature importance. A possible solution to improve the reliability of explanations is to combine results from multiple feature importance quantifiers from different machine learning approaches coupled with re-sampling. Current state-of-the-art ensemble feature importance fusion uses crisp techniques to fuse results from different approaches. There is, however, significant loss of information as these approaches are not context-aware and reduce several quantifiers to a single crisp output. More importantly, their representation of 'importance' as coefficients is misleading and incomprehensible to end-users and decision makers. Here we show how the use of fuzzy data fusion methods can overcome some of the important limitations of crisp fusion methods.

cs.LG

Computationally repurposed drugs and natural products against RNA dependent RNA polymerase as potential COVID-19 therapies

For fast development of COVID-19, it is only feasible to use drugs (off label use) or approved natural products that are already registered or been assessed for safety in previous human trials. These agents can be quickly assessed in COVID-19 patients, as their safety and pharmacokinetics should already be well understood. Computational methods offer promise for rapidly screening such products for potential SARS-CoV-2 activity by predicting and ranking the affinities of these compounds for specific virus protein targets. The RNA-dependent RNA polymerase (RdRP) is a promising target for SARS-CoV-2 drug development given it has no human homologs making RdRP inhibitors potentially safer, with fewer off-target effects that drugs targeting other viral proteins. We combined robust Vina docking on RdRP with molecular dynamic (MD) simulation of the top 80 identified drug candidates to yield a list of the most promising RdRP inhibitors. Literature reviews revealed that many of the predicted inhibitors had been shown to have activity in in vitro assays or had been predicted by other groups to have activity. The novel hits revealed by our screen can now be conveniently tested for activity in RdRP inhibition assays and if conformed testing for antiviral activity invitro before being tested in human trials

q-bio.BM

In silico comparison of spike protein-ACE2 binding affinities across species; significance for the possible origin of the SARS-CoV-2 virus

The devastating impact of the COVID-19 pandemic caused by SARS coronavirus 2 (SARS CoV 2) has raised important questions about viral origin, mechanisms of zoonotic transfer to humans, whether companion or commercial animals can act as reservoirs for infection, and why there are large variations in SARS-CoV-2 susceptibilities across animal species. Powerful in silico modelling methods can rapidly generate information on newly emerged pathogens to aid countermeasure development and predict future behaviours. Here we report an in silico structural homology modelling, protein-protein docking, and molecular dynamics simulation study of the key infection initiating interaction between the spike protein of SARS-Cov-2 and its target, angiotensin converting enzyme 2 (ACE2) from multiple species. Human ACE2 has the strongest binding interaction, significantly greater than for any species proposed as source of the virus. Binding to pangolin ACE2 was the second strongest, possibly due to the SARS-CoV-2 spike receptor binding domain (RBD) being identical to pangolin CoV spike RDB. Except for snake, pangolin and bat for which permissiveness has not been tested, all those species in the upper half of the affinity range (human, monkey, hamster, dog, ferret) have been shown to be at least moderately permissive to SARS-CoV-2 infection, supporting a correlation between binding affinity and permissiveness. Our data indicates that the earliest isolates of SARS-CoV-2 were surprisingly well adapted to human ACE2, potentially explaining its rapid transmission.

q-bio.BM

Computational screening of repurposed drugs and natural products against SARS-Cov-2 main protease (Mpro) as potential COVID-19 therapies

There remains an urgent need to identify existing drugs that might be suitable for treating patients suffering from COVID-19 infection. Drugs rarely act at a single molecular target, with off target effects often being responsible for undesirable side effects and sometimes, beneficial synergy between targets for a specific illness. Off target activities have also led to blockbuster drugs in some cases, e.g. Viagra for erectile dysfunction and Minoxidil for male pattern hair loss. Drugs already in use or in clinical trials plus approved natural products constitute a rich resource for discovery of therapeutic agents that can be repurposed for existing and new conditions, based on the rationale that they have already been assessed for safety in man. A key question then is how to rapidly and efficiently screen such compounds for activity against new pandemic pathogens such as COVID-19. Here we show how a fast and robust computational process can be used to screen large libraries of drugs and natural compounds to identify those that may inhibit the main protease of SARS-Cov-2 (3CL pro, Mpro). We show how the resulting shortlist of candidates with strongest binding affinities is highly enriched in compounds that have been independently identified as potential antivirals against COVID-19. The top candidates also include a substantial number of drugs and natural products not previously identified as having potential COVID-19 activity, thereby providing additional targets for experimental validation. This in silico screening pipeline may also be useful for repurposing of existing drugs and discovery of new drug candidates against other medically important pathogens and for use in future pandemics.

q-bio.BM

Impressive computational acceleration by using machine learning for 2-dimensional super-lubricant materials discovery

The screening of novel materials is an important topic in the field of materials science. Although traditional computational modeling, especially first-principles approaches, is a very useful and accurate tool to predict the properties of novel materials, it still demands extensive and expensive state-of-the-art computational resources. Additionally, they can be often extremely time consuming. We describe a time and resource-efficient machine learning approach to create a large dataset of structural properties of van der Waals layered structures. In particular, we focus on the interlayer energy and the elastic constant of layered materials composed of two different 2-dimensional (2D) structures, that are important for novel solid lubricant and super-lubricant materials. We show that machine learning models can recapitulate results of computationally expansive approaches (i.e. density functional theory) with high accuracy.

physics.comp-ph