SearcharxivSearch

arXiv subjects

Jonathan Levy

Publications and source records attributed to Jonathan Levy.

7 recordsLinked to original sources

Vaccine Search Patterns Provide Insights into Vaccination Intent

Despite ample supply of COVID-19 vaccines, the proportion of fully vaccinated individuals remains suboptimal across much of the US. Rapid vaccination of additional people will prevent new infections among both the unvaccinated and the vaccinated, thus saving lives. With the rapid rollout of vaccination efforts this year, the internet has become a dominant source of information about COVID-19 vaccines, their safety and efficacy, and their availability. We sought to evaluate whether trends in internet searches related to COVID-19 vaccination - as reflected by Google's Vaccine Search Insights (VSI) index - could be used as a marker of population-level interest in receiving a vaccination. We found that between January and August of 2021: 1) Google's weekly VSI index was associated with the number of new vaccinations administered in the subsequent three weeks, and 2) the average VSI index in earlier months was strongly correlated (up to r = 0.89) with vaccination rates many months later. Given these results, we illustrate an approach by which data on search interest may be combined with other available data to inform local public health outreach and vaccination efforts. These results suggest that the VSI index may be useful as a leading indicator of population-level interest in or intent to obtain a COVID-19 vaccine, especially early in the vaccine deployment efforts. These results may be relevant to current efforts to administer COVID-19 vaccines to unvaccinated individuals, to newly eligible children, and to those eligible to receive a booster shot. More broadly, these results highlight the opportunities for anonymized and aggregated internet search data, available in near real-time, to inform the response to public health emergencies.

cs.SI

Transporting stochastic direct and indirect effects to new populations

Transported mediation effects may contribute to understanding how and why interventions may work differently when applied to new populations. However, we are not aware of any estimators for such effects. Thus, we propose several different estimators of transported stochastic direct and indirect effects: an inverse-probability of treatment stabilized weighted estimator, a doubly robust estimator that solves the estimating equation, and a doubly robust substitution estimator in the targeted minimum loss-based framework. We demonstrate their finite sample properties in a simulation study.

stat.ME

Tutorial: Deriving The Efficient Influence Curve for Large Models

This paper aims to provide a tutorial for upper level undergraduate and graduate students in statistics, biostatistics and epidemiology on deriving influence functions for non-parametric and semi-parametric models. The author will build on previously known efficiency theory and provide a useful identity and formulaic technique only relying on the basics of integration which, are self-contained in this tutorial and can be used in most any setting one might encounter in practice. The paper provides many examples of such derivations for well-known influence functions as well as for new parameters of interest. The influence function remains a central object for constructing efficient estimators for large models, such as the one-step estimator and the targeted maximum likelihood estimator. We will not touch upon these estimators at all but readers familiar with these estimators might find this tutorial of particular use.

math.ST

Kernel Smoothing of the Treatment Effect CDF

The strata-specific treatment effect or so-called blip for a randomly drawn strata of confounders defines a random variable and a corresponding cumulative distribution function. However, the CDF is not pathwise differentiable, necessitating a kernel smoothing approach to estimate it at a given point or perhaps many points. Assuming the CDF is continuous, we derive the efficient influence curve of the kernel smoothed version of the blip CDF and a CV-TMLE estimator. The estimator is asymptotically efficient under two conditions, one of which involves a second order remainder term which, in this case, shows us that knowledge of the treatment mechanism does not guarantee a consistent estimate. The remainder term also teaches us exactly how well we need to estimate the nuisance parameters (outcome model and treatment mechanism) to guarantee asymptotic efficiency. Through simulations we verify theoretical properties of the estimator and show the importance of machine learning over conventional regression approaches to fitting the nuisance parameters. We also derive the bias and variance of the estimator, the orders of which are analogous to a kernel density estimator. This estimator opens up the possibility of developing methodology for optimal choice of the kernel and bandwidth to form confidence bounds for the CDF itself.

stat.ME

An Easy Implementation of CV-TMLE

In the world of targeted learning, cross-validated targeted maximum likelihood estimators, CV-TMLE [Zheng:2010aa], has a distinct advantage over TMLE [Laan:2006aa] in that one less condition is required of CV-TMLE in order to achieve asymptotic efficiency in the nonparametric or semiparametric settings. CV-TMLE as originally formulated, consists of averaging usually 10 (for 10-fold cross-validation) parameter estimates, each of which is performed on a validation set separate from where the initial fit was trained. The targeting step is usually performed as a pooled regression over all validation folds but in each fold, we separately evaluate any means as well as the parameter estimate. One nice thing about CV-TMLE, is that we average 10 plug-in estimates so the plug-in quality of preserving the natural parameter bounds is respected. Our adjustment of this procedure also preserves the plug-in characteristic as well as avoids the donsker condtion. The advantage of our procedure is the implementation of the targeting is identical to that of a regular TMLE, once all the validation set initial predictions have been formed. In short, we stack the validation set predictions and pretend as if we have a regular TMLE, which is not necessarily quite a plug-in estimator on each fold but overall will perform asymptotically the same and might have some slight advantage, a subject for future research. In the case of average treatment effect, treatment specific mean and mean outcome under a stochastic intervention, the procedure overlaps exactly with the originally formulated CV-TMLE with a pooled regression for the targeting.

stat.ME

A Fundamental Measure of Treatment Effect Heterogeneity

We offer a non-parametric plug-in estimator for an important measure of treatment effect variability and provide minimum conditions under which the estimator is asymptotically efficient. The stratum specific treatment effect function or so-called blip function, is the average treatment effect for a randomly drawn stratum of confounders. The mean of the blip function is the average treatment effect (ATE), whereas the variance of the blip function (VTE), the main subject of this paper, measures overall clinical effect heterogeneity, perhaps providing a strong impetus to refine treatment based on the confounders. VTE is also an important measure for assessing reliability of the treatment for an individual. The CV-TMLE provides simultaneous plug-in estimates and inference for both ATE and VTE, guaranteeing asymptotic efficiency under one less condition than for TMLE. This condition is difficult to guarantee a priori, particularly when using highly adaptive machine learning that we need to employ in order to eliminate bias. Even in defiance of this condition, CV-TMLE sampling distributions maintain normality, not guaranteed for TMLE, and have a lower mean squared error than their TMLE counterparts. In addition to verifying the theoretical properties of TMLE and CV-TMLE through simulations, we point out some of the challenges in estimating VTE, which lacks double robustness and might be unavoidably biased if the true VTE is small and sample size insufficient. We will provide an application of the estimator on a data set for treatment of acute trauma patients.

stat.ME

Canonical Least Favorable Submodels:A New TMLE Procedure for Multidimensional Parameters

This paper is a fundamental addition to the world of targeted maximum likelihood estimation (TMLE) (or likewise, targeted minimum loss estimation) for simultaneous estimation of multi-dimensional parameters of interest. TMLE, as part of the targeted learning framework, offers a crucial step in constructing efficient plug-in estimators for nonparametric or semiparametric models. The so-called targeting step of targeted learning, involves fluctuating the initial fit of the model in a way that maximally adjusts the plug-in estimate per change in the log likelihood. Previously for multidimensional parameters of interest, iterative TMLE's were constructed using locally least favorable submodels as defined in van der Laan and Gruber, 2016, which are indexed by a multidimensional fluctuation parameter. In this paper we define a canonical least favorable submodel in terms of a single dimensional epsilon for a $d$-dimensional parameter of interest. One can view the clfm as the iterative analog to the one-step TMLE as constructed in van der Laan and Gruber, 2016. It is currently implemented in several software packages we provide in the last section. Using a single epsilon for the targeting step in TMLE could be useful for high dimensional parameters, where using a fluctuation parameter of the same dimension as the parameter of interest could suffer the consequences of curse of dimensionality. The clfm also enables placing the so-called clever covariate denominator as an inverse weight in an offset intercept model. It has been shown that such weighting mitigates the effect of large inverse weights sometimes caused by near positivity violations.

stat.ME