Searcharxiv⌕ Search

arXiv subjects

Fatemeh Mahmoudi

Publications and source records attributed to Fatemeh Mahmoudi.

5 recordsLinked to original sources

An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors

Blood-brain barrier (BBB) permeability is a critical determinant in the development of central nervous system therapeutics because it directly influences the ability of drug candidates to reach their target sites within the brain. In this study, an explainable machine learning framework was developed to predict BBB permeability using molecular descriptors generated from the MoleculeNet BBBP dataset with the RDKit cheminformatics toolkit. Fifteen physicochemical descriptors extracted from 2,039 compounds were used to train four supervised machine learning algorithms, including Logistic Regression, Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost). Hyperparameter optimization was performed using GridSearchCV, while model interpretability was investigated using SHapley Additive exPlanations (SHAP). Among the evaluated models, the optimized XGBoost classifier achieved the best predictive performance, with an accuracy of 88.97%, a precision of 88.92%, a recall of 97.76%, an F1-score of 93.13%, and a ROC-AUC of 0.9282. Stratified five-fold cross-validation further demonstrated the robustness of the proposed model, yielding a mean ROC-AUC of 0.8982 +/- 0.0130. Feature importance and SHAP analyses consistently identified TPSA, HBD, and LogP as the most influential molecular descriptors governing BBB permeability prediction. Overall, the proposed framework provides an accurate, interpretable, and computationally efficient approach for BBB permeability prediction and may serve as a valuable tool for the early-stage screening of CNS drug candidates.

cs.LG↗

A Machine Learning Framework for Predicting Glass-Forming Ability in Ternary Alloy Systems

Predicting the glass-forming ability (GFA) of chemical compositions remains a fundamental challenge in materials science, especially for oxide glasses with broad compositional diversity. Traditional empirical and thermodynamic approaches often fail to capture the complex, nonlinear factors governing vitrification. In this study, we applied two ensemble machine learning algorithms-Random Forest (RF) and Extreme Gradient Boosting (XGB)-to the glass_ternary_hipt dataset to predict the GFA of ternary oxide glasses directly from composition-derived descriptors. Both models achieved excellent predictive accuracy (R^2 > 0.92, MAE < 0.04), confirming that GFA is learnable from compositional features alone. Feature importance analysis revealed that electronegativity variance, atomic size mismatch, and valence electron descriptors are the most influential factors, while cohesive energy and ionic radius provided secondary contributions. These chemically interpretable features align with established theories of glass formation, thereby bridging predictive performance with physical understanding. The novelty of this work lies in systematically extending ML-based predictive modeling to ternary oxide glasses, a class less studied compared to metallic and binary systems. Our results demonstrate that ensemble learning not only enables accurate GFA prediction but also provides actionable insights for designing new glass compositions with enhanced stability.

cond-mat.mtrl-sci↗

Variable Selection with Broken Adaptive Ridge Regression for Interval-Censored Competing Risks Data

Competing risks data refer to situations where the occurrence of one event pre- cludes the possibility of other events happening, resulting in multiple mutually exclusive events. This data type is commonly encountered in medical research and clinical trials, exploring the interplay between different events and informing decision-making in fields such as healthcare and epidemiology. We develop a penal- ized variable selection procedure to handle such complex data in an interval-censored setting. We consider a broad class of semiparametric transformation regression mod- els, including popular models such as proportional and non-proportional hazards models. To promote sparsity and select variables specific to each event, we employ the broken adaptive ridge (BAR) penalty. This approach allows us to simultane- ously select important risk factors and estimate their effects for each event under investigation. We establish the oracle property of the BAR procedure and evaluate its performance through simulation studies. The proposed method is applied to a real-life HIV cohort dataset, further validating its applicability in practice.

stat.ME↗

Variable selection in the joint frailty model of recurrent and terminal events using Broken Adaptive Ridge regression

We introduce a novel method to simultaneously perform variable selection and estimation in the joint frailty model of recurrent and terminal events using the Broken Adaptive Ridge Regression penalty. The BAR penalty can be summarized as an iteratively reweighted squared $L_2$-penalized regression, which approximates the $L_0$-regularization method. Our method allows for the number of covariates to diverge with the sample size. Under certain regularity conditions, we prove that the BAR estimator implemented under the model framework is consistent and asymptotically normally distributed, which are known as the oracle properties in the variable selection literature. In our simulation studies, we compare our proposed method to the Minimum Information Criterion (MIC) method. We apply our method on the Medical Information Mart for Intensive Care (MIMIC-III) database, with the aim of investigating which variables affect the risks of repeated ICU admissions and death during ICU stay.

stat.ME↗

Penalized Variable Selection with Broken Adaptive Ridge Regression for Semi-competing Risks Data

Semi-competing risks data arise when both non-terminal and terminal events are considered in a model. Such data with multiple events of interest are frequently encountered in medical research and clinical trials. In this framework, terminal event can censor the non-terminal event but not vice versa. It is known that variable selection is practical in identifying significant risk factors in high-dimensional data. While some recent works on penalized variable selection deal with these competing risks separately without incorporating possible correlation between them, we perform variable selection in an illness-death model using shared frailty where semiparametric hazard regression models are used to model the effect of covariates. We propose a broken adaptive ridge (BAR) penalty to encourage sparsity and conduct extensive simulation studies to compare its performance with other popular methods. We perform variable selection in an event specific manner so that the potential risk factors and covariates effects can be estimated and selected, simultaneously corresponding to each event in the study. The grouping effect, as well as the oracle property of the proposed BAR procedure are investigated using simulation studies. The proposed method is then applied to real-life data arising from a Colon Cancer study.

stat.ME↗