SearcharxivSearch

arXiv subjects

Gueorgui Mihaylov

Publications and source records attributed to Gueorgui Mihaylov.

8 recordsLinked to original sources

Machine learning model leveraging SMILES-derived NMR spectroscopy data to predict dopamine D1 receptor antagonists: a prospective framework for forecasting the impact of engineered nanoparticles on the functionalities of small biomolecules

The article proposes a conceptual approach for evaluating the impact of engineered nanoparticles (NPs) on the functionality of small biomolecules. The developed machine learning (ML) model is based on in-silico 13C NMR spectroscopy chemical shifts derived by the SMILES notations on small biomolecules. The rationale behind this approach is that 13C NMR provide information about the atom environment of the carbon atoms. Thus, decomposing the small biomolecules into their fundamental 13C NMR spectral data, and performing classification based on the count and position of chemical peaks, establishes a baseline for evaluating the impact of NPs on the functionality of small biomolecules, even if the ML model is not based on nano data. The approach mitigates not only the scarcity of nano-bio data but also hold potential for building of NP`s portfolio by utilising data collected from various in vitro, in situ, in vivo, and organ-on-a-chip environments across multiple timeframes. Such a framework enables predictive modeling based on these multi-environmental datasets, facilitating a deeper understanding of NP behaviour. The methodology was demonstrated using data from bioassay focused on human dopamine D1 receptor antagonists provided by PubChem. The model was train with 26,766 samples and test on 5,466 samples, achieving Accuracy of 70.8%, Precision of 74.3%, recall of 63.6%, F1-score of 68.5% and ROC of 70.8% were achieved by the Support Vector classifier, with an Area Under the Curve (AUC) of 76% and Matthews Correlation Coefficient, MCC=0.4204. A secondary, non-NP-related ML model was developed to complement the study case. It uses PubChem compound and substance identifiers (CIDs and SIDs) to predict whether pre-designed small biomolecules have the potential to be human dopamine D1 receptor antagonists.

q-bio.OT

Machine Learning - driven insights for predicting the impact of nanoparticles on the functionality of biomolecules, Illustrated by the case of DNA Damage-Inducible Transcript 3 (CHOP) inhibitors

This study introduces a pioneering machine learning (ML)-based approach for predicting the impact of nanoparticle (NP) carriers on the functionality of attached small biomolecules. It was hypothesised that NP interactions induce measurable perturbations in the atomic environment of the small biomolecules, which are reliably captured by chemical shifts in 13C and 1H NMR spectroscopy. Ten datasets were generated by combining 13C, 1H NMR spectroscopy data, derived from SMILES notations and molecular features provided by PubChem. The resulting datasets were used to train predictive models via traditional ML algorithms (Scikit-learn) and Deep Neural Network DNN (PyTorch). The methodology was demonstrated through a quantitative high-throughput screening (qHTS) focused on DNA Damage-Inducible Transcript 3 (CHOP) inhibitors. The optimal ML performance was achieved by the Random Forest Classifier, which was trained on 19,184 samples and tested on 4,000, resulting in 81.1% accuracy, 83.4% precision, 77.7% recall, 80.4% F1-score, 81.1% ROC, and a five-fold cross-validation score of 0.821. Complementing the main study, two computational approaches were developed to enhance CHOP inhibitor prediction. The first identifies the most desirable/undesirable functional groups for CHOP inhibition. The second, a CID_SID ML model, achieved 90.1% accuracy in predicting whether compounds designed for other purposes possess CHOP inhibition potential.

q-bio.QM

In Silico Functional Profiling of Engineered Small Molecules: A Machine Learning Approach Leveraging PubChem Identifiers (CID_SID ML model)

The article introduces a concept for a time- and cost-effective methodological framework leveraging machine learning (ML) models for both early-stage drug development and clinical trial support. The rationale for this approach is the inherent scalability and speed enabled by using pre-calculated data embedded in existing PubChem identifiers (CID and SID), thereby eliminating the computationally intensive step of on-the-fly molecular descriptor generation. The approach was effectively demonstrated across four diverse bioassays: antagonists of the human D3 dopamine receptor, Rab9 promoter activators, small-molecule inhibitors of CHOP, and antagonists of the human M1 muscarinic receptor. A comparison, based on Matthews correlation coefficient (MCC), was conducted between the CID_SID ML model, the MORGAN2-based ML model, and the RDKit-transformed SMILES model for these four case studies, revealing that no method is universally superior in terms of performance. Furthermore, the CID_SID model averaged a rapid execution time of only 3.3 seconds; the ML models relying on explicit structural descriptors, such as MORGAN2 and RDKit-transformed SMILES, demonstrated high computational costs, with processing times averaging 106.0 and 109.6 seconds, respectively. While negligible for a single ML model, these times would cause a significant difference in computational resource consumption when scaled across a framework involving over a million buildings. Moreover, the CID_SID ML model achieved strong average performance metrics: Accuracy of 83.52%, Precision of 89.62%, Recall of 75.65%, F1-Score of 81.93% and ROC of 83.53%.

q-bio.QM

IUPAC-Induced Computational Approaches for Identifying Boosters of Small Biomolecule Functionality: A Case Study of Human Tyrosyl-DNA Phosphodiesterase 1 (TDP1) Inhibitors

This paper introduces several proof-of-concept (PoC) computational methods intended to offer biochemical researchers straightforward, time- and cost-effective strategies to accelerate their work. While Machine Learning (ML) models were developed, the study's central purpose was to explore approaches for the identification of desirable functional groups/fragments in small biomolecules regarding a specific functionality, which, in this case, was human tyrosyl-DNA phosphodiesterase 1 (TDP1) inhibition. This was achieved primarily by tokenising IUPAC names to generate features. Additionally, the applicability of the CID_SID ML model for predicting TDP1 activity was developed and explored. Since these computational approaches were not experimentally validated due to a lack of appropriate laboratory facilities, they are presented as open proposals for further laboratory investigation.

q-bio.QM

Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists

The study expands the application of scikit-learn-based machine learning (ML) to the prediction of small biomolecule functionalities based on Carbon 13 isotope (13C) NMR spectroscopy data derived from Simplified Molecular Input Line Entry System (SMILES) notations. The methodology previously demonstrated by predicting dopamine D1 receptor antagonists was upgraded with the addition of new molecular features derived from the PubChem database. The enhanced ML model obtained 75.8% Accuracy, 84.2% Precision, 63.6% Recall, 72.5% F1-score and 75.8 % ROC, when is trained on 25,532 samples and tested on 5,466 samples. To evaluate the applicability of the methodology for a variety of case studies, a comparison was conducted between the prediction capabilities of the ML models based on the human dopamine D1 receptor antagonists and on the neuronal Transthyretin (TTR) transcription activators. Since the TTR bioassay did not contain the required number of samples for comparison, the results were obtained hypothetically. Gradient Boosting classifier was the optimal model for TTR transcription activators, achieving hypothetical 67.4% Accuracy, 74.0% Precision, 53.5% Recall, 62.1% F1-score, 67.4 % ROC, if it could be trained with 25,532 samples and tested with 5,466 samples. In addition to the main study, to the attention of those interested in neuronal TTR, the CID_SID ML model has been developed to predict whether a compound, initially designed for another purpose, possesses TTR transcription activation capabilities. This ML model was based solely on its PubChem CID and SID and achieved 81.5% Accuracy, 94.6% Precision, 66.8% Recall, 78.3% F1-score, 81.5 % ROC.

q-bio.QM

Algorithms for Shipping Container Delivery Scheduling

Motivated by distribution problems arising in the supply chain of Haleon, we investigate a discrete optimization problem that we call the "container delivery scheduling problem". The problem models a supplier dispatching ordered products with shipping containers from manufacturing sites to distribution centers, where orders are collected by the buyers at agreed due times. The supplier may expedite or delay item deliveries to reduce transshipment costs at the price of increasing inventory costs, as measured by the number of containers and distribution center storage/backlog costs, respectively. The goal is to compute a delivery schedule attaining good trade-offs between the two. This container delivery scheduling problem is a temporal variant of classic bin packing problems, where the item sizes are not fixed, but depend on the item due times and delivery times. An approach for solving the problem should specify a batching policy for container consolidation and a scheduling policy for deciding when each container should be delivered. Based on the available item due times, we develop algorithms with sequential and nested batching policies as well as on-time and delay-tolerant scheduling policies. We elaborate on the problem's hardness and substantiate the proposed algorithms with positive and negative approximation bounds, including the derivation of an algorithm achieving an asymptotically tight 2-approximation ratio.

math.OC

Improved Galactic Foreground Removal for B-Modes Detection with Clustering Methods

Characterizing the sub-mm Galactic emission has become increasingly critical especially in identifying and removing its polarized contribution from the one emitted by the Cosmic Microwave Background (CMB). In this work, we present a parametric foreground removal performed onto sub-patches identified in the celestial sphere by means of spectral clustering. Our approach takes into account efficiently both the geometrical affinity and the similarity induced by the measurements and the accompanying errors. The optimal partition is then used to parametrically separate the Galactic emission encoding thermal dust and synchrotron from the CMB one applied on two nominal observations of forthcoming experiments from the ground and from the space. Performing the parametric fit singularly on each of the clustering derived regions results in an overall improvement: both controlling the bias and the uncertainties in the CMB $B-$mode recovered maps. We finally apply this technique using the map of the number of clouds along the line of sight, $\mathcal{N}_c$, as estimated from HI emission data and perform parametric fitting onto patches derived by clustering on this map. We show that adopting the $\mathcal{N}_c$ map as a tracer for the patches related to the thermal dust emission, results in reducing the $B-$mode residuals post-component separation. The code is made publicly available.

astro-ph.CO

High-Dimensional Changepoint Detection via a Geometrically Inspired Mapping

High-dimensional changepoint analysis is a growing area of research and has applications in a wide range of fields. The aim is to accurately and efficiently detect changepoints in time series data when both the number of time points and dimensions grow large. Existing methods typically aggregate or project the data to a smaller number of dimensions; usually one. We present a high-dimensional changepoint detection method that takes inspiration from geometry to map a high-dimensional time series to two dimensions. We show theoretically and through simulation that if the input series is Gaussian then the mappings preserve the Gaussianity of the data. Applying univariate changepoint detection methods to both mapped series allows the detection of changepoints that correspond to changes in the mean and variance of the original time series. We demonstrate that this approach outperforms the current state-of-the-art multivariate changepoint methods in terms of accuracy of detected changepoints and computational efficiency. We conclude with applications from genetics and finance.

stat.ME