SearcharxivSearch

arXiv subjects

Lucas Godoy Garraza

Publications and source records attributed to Lucas Godoy Garraza.

4 recordsLinked to original sources

Principal Stratification with Bayesian Additive Regression Trees for Count-Valued Intermediate Variables: Estimating the Effect of Fertility on Women's Employment

Estimating the causal effect of fertility on women's employment is challenging because fertility and labour-market decisions are jointly determined. Instrumental-variable strategies are widely used, but credible instruments are rare and their validity often depends on covariates. Two-stage least squares, the dominant implementation, does not flexibly accommodate covariate-dependent instrument validity and is poorly suited to count-valued treatment and effect heterogeneity more broadly. We extend an existing framework that combines principal stratification with Bayesian Additive Regression Trees (BART) to settings with count-valued intermediate variables, such as number of children. The approach defines a causal estimand that respects the count structure of the intermediate variable, while BART enables flexible, covariate-dependent modelling of principal strata and outcomes, accommodating covariate-dependent instrument validity and producing estimates of effect heterogeneity. Simulations demonstrate that our approach outperforms conventional and flexible instrumental-variable estimators under nonlinear confounding and heterogeneous treatment effects. Applied to Demographic and Health Survey data from Nigeria, Senegal, and Kenya, the approach suggests a negative average effect in Nigeria but no clear average effect elsewhere, with employment penalties concentrated among younger, less-educated women, heterogeneity that standard approaches would obscure. The approach is available in the R package PrinceBART.

stat.ME

Combining BART and Principal Stratification to estimate the effect of intermediate on primary outcomes with application to estimating the effect of family planning on employment in sub-Saharan Africa

There is interest in learning about the causal effect of family planning (FP) on empowerment related outcomes. Experimental data related to this question are available from trials in which FP programs increase access to FP. While program assignment is unconfounded, FP uptake and subsequent empowerment may share common causes. We use principal stratification to estimate the causal effect of an intermediate FP outcome on a primary outcome of interest, among women affected by a FP program. Within strata defined by the potential reaction to the program, FP uptake is unconfounded. To minimize the need for parametric assumptions, we propose to use Bayesian Additive Regression Trees (BART) for modeling stratum membership and outcomes of interest. We refer to the combined approach as Prince BART. We evaluate Prince BART through a simulation study and use it to assess the causal effect of modern contraceptive use on employment in six cities in Nigeria, based on quasi-experimental data from a FP program trial during the first half of the 2010s. We show that findings differ between Prince BART and alternative modeling approaches based on parametric assumptions.

stat.ME

Combining BART and Principal Stratification to estimate the effect of intermediate variables on primary outcomes with application to estimating the effect of family planning on employment in Nigeria and Senegal

There is interest in learning about the causal effects of modern contraceptive use on empowerment outcomes. Data on this question often come from family planning (FP) programs that increase access to FP and facilitate contraceptive use among some women, rather than directly assigning use. Women whose contraceptive behavior changes because of these programs ("compliers") may differ from target populations in ways that alter the consequences of contraceptive use for empowerment outcomes. We propose a two-step approach. First, we use principal stratification and Bayesian Additive Regression Trees (BART) to estimate the effect of modern contraceptive use among compliers in the study population, treating the FP program as an instrument rather than as the treatment of interest. Second, we generalize these complier-specific effects to a broader population by averaging conditional effects over the covariate distribution in the target population, with uncertainty in that distribution quantified via a Bayesian bootstrap applied to external complex survey data. We examine performance in simulation designs previously used to evaluate IV estimators. We then apply the approach to employment among urban women in Nigeria and Senegal, finding strong and heterogeneous effects of contraceptive use. Sensitivity analyses suggest robustness to violations of assumptions for internal and external validity.

stat.ME

Adaptive Selection of the Optimal Strategy to Improve Precision and Power in Randomized Trials

Benkeser et al. demonstrate how adjustment for baseline covariates in randomized trials can meaningfully improve precision for a variety of outcome types. Their findings build on a long history, starting in 1932 with R.A. Fisher and including more recent endorsements by the U.S. Food and Drug Administration and the European Medicines Agency. Here, we address an important practical consideration: *how* to select the adjustment approach -- which variables and in which form -- to maximize precision, while maintaining Type-I error control. Balzer et al. previously proposed *Adaptive Prespecification* within TMLE to flexibly and automatically select, from a prespecified set, the approach that maximizes empirical efficiency in small trials (N$<$40). To avoid overfitting with few randomized units, selection was previously limited to working generalized linear models, adjusting for a single covariate. Now, we tailor Adaptive Prespecification to trials with many randomized units. Using $V$-fold cross-validation and the estimated influence curve-squared as the loss function, we select from an expanded set of candidates, including modern machine learning methods adjusting for multiple covariates. As assessed in simulations exploring a variety of data generating processes, our approach maintains Type-I error control (under the null) and offers substantial gains in precision -- equivalent to 20-43\% reductions in sample size for the same statistical power. When applied to real data from ACTG Study 175, we also see meaningful efficiency improvements overall and within subgroups.

stat.ME