Searcharxiv⌕ Search

arXiv subjects

Abdul-Nasah Soale

Publications and source records attributed to Abdul-Nasah Soale.

11 recordsLinked to original sources

Robust Dual-Regularized Variable Selection under Outlier Contamination

Real data often contain unusual observations that can exert disproportionate effects on variable selection, especially in complex predictor settings. We propose a two-stage {\it sparse median outer product of gradients (smOPG)} method for variable selection in single index models with outlier contamination. We first estimate sparse local gradients via \(\ell_1\)-penalized local median regression and then recover the active predictor set from a rank-one sparse approximation of the resulting gradient matrix using regularized singular value decomposition. The combination of median regression and local weighting provides robustness to both response outliers and leverage points. Extensive simulations across varying dimensions and contamination mechanisms demonstrate the favorable variable selection performance of smOPG relative to existing methods. Applications to air pollution and genomic data demonstrate practical utility, while theory establishes active-set recovery without requiring selection consistency of individual local regressions.

stat.ME↗

Near-Optimal Nitrogen Recommendations for Precision Agriculture via Sequential Screening and Hierarchical Refinement

Nitrogen fertilizer management plays a central role in balancing agricultural productivity and environmental sustainability, yet identifying optimal application strategies remains difficult because treatment responses vary substantially across locations and many fertilizer choices are statistically indistinguishable near the optimum. This paper develops a hierarchical refinement procedure, built on sequential screening, for fertilizer recommendation in multi-site experiments that explicitly accounts for spatial heterogeneity while prioritizing parsimonious, decision-oriented selection. Rather than targeting a single estimated best treatment, the proposed method first conducts sequential screening at a higher aggregation level to eliminate clearly inferior fertilizer choices and then refines recommendations locally among the surviving candidates. We study the asymptotic properties of the proposed estimators and show that it provides screening-safety guaranteed recommendations. The efficacy of the new approach is investigated through a multi-state, multi-year corn nitrogen trial. The results show that no single fertilizer regime is uniformly optimal within a state; instead, each state is associated with multiple recommended choices, and the most common recommendation typically covers only about one-third to one-half of decision units, underscoring substantial within-state heterogeneity. Representative site-level comparisons further demonstrate that the proposed method often yields lower total nitrogen recommendations than state-level or hindsight benchmarks while maintaining competitive agronomic performance.

stat.ME↗

A Distance Covariance-based Estimator

This paper proposes an estimator that relaxes the conventional relevance condition in instrumental variable (IV) analyses. The method allows endogenous covariates to be weakly correlated, uncorrelated, or even mean-independent -- though not independent -- of the instruments, enabling the use of the maximal set of relevant instruments in a given application. Identification is attainable without exclusion restrictions and without finite-moment assumptions on the disturbance term. Under either of two non-nested exogeneity conditions, combined with mild regularity conditions, the parameter of interest is identified. The estimator is shown to be consistent and asymptotically normal, and the relaxed relevance condition required for identification is testable.

econ.EM↗

Adaptive Influence Diagnostics in High-Dimensional Regression

An adaptive Cook's distance (ACD) for diagnosing influential observations in high-dimensional single-index models with multicollinearity and outlier contamination is proposed. ACD is a model-free technique built on sparse local linear gradients to temper leverage effects. In simulations spanning low- and high-dimensional design settings with strong correlation, ACD based on LASSO (ACD-LASSO) and SCAD (ACD-SCAD) penalties reduced masking and swamping relative to classical Cook's distance and local influence as well as the DF-Model and Case-Weight adjusted solution for LASSO. Trimming points flagged by ACD stabilizes variable selection while preserving core signals. Applications to two datasets--the 1960 US cities pollution study and a high-dimensional riboflavin genomics experiment show consistent gains in selection stability and interpretability.

stat.ME↗

Clustered Covariate Regression

High covariate dimensionality is increasingly occurrent in model estimation, and existing techniques to address this issue typically require sparsity or discrete heterogeneity of the \emph{unobservable} parameter vector. However, neither restriction may be supported by economic theory in some empirical contexts, leading to severe bias and misleading inference. The clustering-based grouped parameter estimator (GPE) introduced in this paper drops both restrictions and maintains the natural one that the parameter support be bounded. GPE exhibits robust large sample properties under standard conditions and accommodates both sparse and non-sparse parameters whose support can be bounded away from zero. Extensive Monte Carlo simulations demonstrate the excellent performance of GPE in terms of bias reduction and size control compared to competing estimators. An empirical application of GPE to estimating price and income elasticities of demand for gasoline highlights its practical utility.

econ.EM↗

On metric choice in dimension reduction for Fréchet regression

Fréchet regression is becoming a mainstay in modern data analysis for analyzing non-traditional data types belonging to general metric spaces. This novel regression method is especially useful in the analysis of complex health data such as continuous monitoring and imaging data. Fréchet regression utilizes the pairwise distances between the random objects, which makes the choice of metric crucial in the estimation. In this paper, existing dimension reduction methods for Fréchet regression are reviewed, and the effect of metric choice on the estimation of the dimension reduction subspace is explored for the regression between random responses and Euclidean predictors. Extensive numerical studies illustrate how different metrics affect the central and central mean space estimators. Two real applications involving analysis of brain connectivity networks of subjects with and without Parkinson's disease and an analysis of the distributions of glycaemia based on continuous glucose monitoring data are provided, to demonstrate how metric choice can influence findings in real applications.

stat.ME↗

Detecting influential observations in single-index Fréchet regression

Regression with random data objects is becoming increasingly common in modern data analysis. Unfortunately, this novel regression method is not immune to the trouble caused by unusual observations. A metric Cook's distance extending the original Cook's distances of Cook (1977) to regression between metric-valued response objects and Euclidean predictors is proposed. The performance of the metric Cook's distance is demonstrated in regression across four different response spaces in an extensive experimental study. Two real data applications involving the analyses of distributions of COVID-19 transmission in the State of Texas and the analyses of the structural brain connectivity networks are provided to illustrate the utility of the proposed method in practice.

stat.CO↗

Sufficient dimension reduction for regression with metric space-valued responses

Data visualization and dimension reduction for regression between a general metric space-valued response and Euclidean predictors is proposed. Current Fréchét dimension reduction methods require that the response metric space be continuously embeddable into a Hilbert space, which imposes restriction on the type of metric and kernel choice. We relax this assumption by proposing a Euclidean embedding technique which avoids the use of kernels. Under this framework, classical dimension reduction methods such as ordinary least squares and sliced inverse regression are extended. An extensive simulation experiment demonstrates the superior performance of the proposed method on synthetic data compared to existing methods where applicable. The real data analysis of factors influencing the distribution of COVID-19 transmission in the U.S. and the association between BMI and structural brain connectivity of healthy individuals are also investigated.

stat.ME↗

A selective review of sufficient dimension reduction for multivariate response regression

We review sufficient dimension reduction (SDR) estimators with multivariate response in this paper. A wide range of SDR methods are characterized as inverse regression SDR estimators or forward regression SDR estimators. The inverse regression family include pooled marginal estimators, projective resampling estimators, and distance-based estimators. Ordinary least squares, partial least squares, and semiparametric SDR estimators, on the other hand, are discussed as estimators from the forward regression family.

stat.ME↗

On expectile-assisted inverse regression estimation for sufficient dimension reduction

Moment-based sufficient dimension reduction methods such as sliced inverse regression may not work well in the presence of heteroscedasticity. We propose to first estimate the expectiles through kernel expectile regression, and then carry out dimension reduction based on random projections of the regression expectiles. Several popular inverse regression methods in the literature are extended under this general framework. The proposed expectile-assisted methods outperform existing moment-based dimension reduction methods in both numerical studies and an analysis of the Big Mac data.

stat.CO↗

On sufficient dimension reduction via principal asymmetric least squares

In this paper, we introduce principal asymmetric least squares (PALS) as a unified framework for linear and nonlinear sufficient dimension reduction. Classical methods such as sliced inverse regression (Li, 1991) and principal support vector machines (Li, Artemiou and Li, 2011) may not perform well in the presence of heteroscedasticity, while our proposal addresses this limitation by synthesizing different expectile levels. Through extensive numerical studies, we demonstrate the superior performance of PALS in terms of both computation time and estimation accuracy. For the asymptotic analysis of PALS for linear sufficient dimension reduction, we develop new tools to compute the derivative of an expectation of a non-Lipschitz function.

math.ST↗