SearcharxivSearch

arXiv subjects

Natasha K. Martin

Publications and source records attributed to Natasha K. Martin.

3 recordsLinked to original sources

Regression Not-to-the-Mean: An Oddity of Regression, Illustrated with the Risk of Overdose Deaths

Recent works in econometrics have shown that there can be issues with applying a constant treatment effect model in longitudinal settings with staggered treatment and heterogeneous treatment effects. We focus on the issue that the estimated constant treatment effect may be a weighted average, with some negative weights, of treatment effects that are heterogeneous across treatment durations. When this issue arises, the estimated constant treatment effect and estimated heterogeneous treatment effects may result in conflicting results. Through the example of estimating the effect of drug-induced homicide (DIH) prosecutions reported by media on unintentional drug-overdose deaths in the United States, we illustrate how the negative weighting issue can lead to conflicting results in practice. Moreover, although research has shown that the negative weight issue may arise in linear regression models, we show this issue may also arise in logistic regression models. Using a linear link, we estimated a constant treatment effect risk ratio of 0.977 (95% CI:(0.866, 1.101)) and an average risk ratio of 0.728 (range: 0.507-0.979) over different treatment durations. Using a logistic link, we estimated a constant treatment risk ratio effect of 1.064 (95% CI: (0.972, 1.165)) and an average risk ratio of 0.739 (range: 0.538-1.008) over different treatment durations. Under both models, the estimated constant treatment effect is either smaller in magnitude or has a different sign than almost all estimated heterogeneous treatment effects, suggesting a negative weighting issue is present. Our results suggest additional care is needed when applying constant treatment effect models in longitudinal settings.

stat.AP

CCMnet: A Software Package for Network Generation with Congruence Class Models

We introduce CCMnet, an R package designed to generate network ensembles that accurately reflect the uncertainty inherent in empirical data. While traditional network modeling often results in ensembles with fixed property values or model-determined levels of variability, CCMnet enables a continuous spectrum of variability for network properties, including edge counts, degree distribution, and mixing patterns. By defining probability distributions directly over congruence classes of networks, the package allows researchers to specify the uncertainty in network properties across the generated ensemble to match a specific sampling design or empirical distribution. Furthermore, this formulation provides a principled framework that encompasses several classic models (e.g., Erdős--Rényi model, stochastic block models, and certain exponential random graph models) that implicitly share this structural basis, while offering the flexibility to specify arbitrary, even non-parametric, distributions for network properties. CCMnet implements a Markov chain Monte Carlo (MCMC) framework to sample from these models. The utility of the package is illustrated by generating posterior predictive network ensembles representing school friendship networks.

stat.CO

Early warning of Mpox outbreaks in U.S. jurisdictions using Lasso Vector Autoregression models with cross-jurisdictional lags

Mpox is an orthopoxvirus that infects humans and animals and is transmitted primarily through close physical contact. The episodic and spatially heterogeneous dynamics of Mpox transmission underscores the need for timely, area-specific forecasts to support targeted public health responses in the U.S. We develop a Vector Autoregression model with Lasso regularization (VAR-Lasso) to generate rolling two-week-ahead forecasts of weekly Mpox cases for eight high-incidence U.S. jurisdictions using national surveillance data from the Centers for Disease Control and Prevention (CDC). The VAR-Lasso model identifies significant long-lag, cross-jurisdictional predictors. For a case study in San Diego County (SDC), these statistical predictors align with phylogenetic analysis that traces a 2023 cluster in SDC to an outbreak in Illinois six months earlier. As the need for public health action is often greatest when incidence is increasing, our performance evaluation focuses on positive-slope weighted error metrics. Forecast performance of the VAR-Lasso model is compared to a uni-variate Auto-Regressive (AR) Lasso model and a naive moving-average estimate. The models are compared using slope-weighted Root Mean Squared Error (RMSE), slope-weighted Mean Absolute Error (MAE), and slope-weighted bias. Across all observations, the VAR-Lasso model reduces slope-weighted RMSE, MAE, and bias by 12%, 7%, and 66% relative to the AR model, and by 16%, 13%, and 76% relative to the naive benchmark. Our findings highlight the value of sparse multivariate time-series models that leverage cross-jurisdictional case data for early forecasting of Mpox outbreaks. Such forecasting can aid health departments in proactively providing timely resources and messaging to mitigate the risks of a future outbreak.

stat.AP