SearcharxivSearch

arXiv subjects

Ariel Linden

Publications and source records attributed to Ariel Linden.

8 recordsLinked to original sources

Cohen's f or Mean Standardized Differences? Assessing Covariate Balance with Multivalued Treatments

Assessing covariate balance across more than two treatment groups has no established omnibus standard: the prevailing practice averages, or takes the maximum of, pairwise standardized mean differences (SMD), while Cohen's f - the classical generalization of Cohen's d to more than two groups - offers an alternative grounded in an established effect-size framework, but the two have not been formally compared. We extend both to arbitrary weighted and covariate-adjusted models via a community-contributed Stata command, esizereg, and validate them in a simulation study of three, four, and six treatment groups under correctly specified and misspecified generalized-propensity-score weighting, correlating each statistic against downstream treatment-effect bias. Cohen's f tracks estimation bias comparably to mean absolute SMD, both pooled (r = 0.93) and within weighting arm (r approximately 0.79); maximum absolute SMD is generally weakest. This ranking was unchanged under a deliberately adversarial nonlinear/interaction outcome model. Cohen's f and the SMD-based statistics are not numerically comparable: we derive an exact representation of f as a size-weighted quadratic function of the pairwise SMDs and prove the minimum attainable ratio between f and mean absolute SMD under equal group weights, so conventional SMD thresholds should not apply directly to f. We recommend reporting f alongside its per-level and pairwise decomposition.

stat.ME

Weighted k-Sample Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling Tests for Assessing Covariate Balance

Weighted distributional tests for covariate balance are currently limited to two-group comparisons. We extend the Kolmogorov-Smirnov, Anderson-Darling, and Cramer-von Mises tests to an arbitrary number of weighted groups k >= 2, using existing k-sample generalizations and a shared permutation-inference procedure; each statistic reduces exactly to its two-group counterpart at k = 2. An accompanying post-hoc pairwise procedure with four multiple-comparison adjustments localizes which groups differ following an omnibus rejection. In a four-scenario simulation study at k = 3, Type I error remained close to nominal, and the comparative advantages established for two groups were preserved: Kolmogorov-Smirnov was most powerful against a centrally located discrepancy, Anderson-Darling against a tail-located discrepancy, and Anderson-Darling and Cramer-von Mises performed comparably against a diffuse discrepancy. Omnibus power was nonetheless uniformly lower than in the matching two-group setting - not because the underlying discrepancy is diluted, but because adding groups enlarges the null distribution itself, so the post-hoc procedure, whose pairwise statistics carry no such penalty, is often the more sensitive tool whenever a specific group's imbalance is suspected. The methods are implemented in the Stata commands kstest, adtest, and cvmtest.

stat.ME

Weighted Extensions of the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling Tests for Assessing Covariate Balance

Assessing covariate balance is a core diagnostic step in causal inference, but commonly used summary measures can miss meaningful distributional differences they are not designed to detect. Distributional goodness-of-fit tests, including the Kolmogorov-Smirnov (KS), Anderson-Darling (AD), and Cramer-von Mises (CVM) tests, offer a more complete comparison but have previously been available only for unweighted data. We extend all three to accommodate case weights of any origin, using a shared label-permutation inference procedure that requires no assumption about how the weights were generated. In a four-scenario simulation study, all three weighted tests controlled Type I error close to nominal across sample sizes from 1,000 to 4,000 under substantial weight variability. Each test's known unweighted comparative advantage was preserved under weighting for two of three discrepancy types: KS was most powerful against a centrally located discrepancy, and AD was overwhelmingly most powerful against a tail-located discrepancy, while AD and CVM performed comparably against a diffuse discrepancy, both outperforming KS. These findings support AD as a reasonable general-purpose default for routine covariate balance assessment, while KS retains an advantage when a centrally concentrated imbalance is specifically suspected. The methods are implemented in the Stata commands kstest, adtest, and cvmtest.

stat.ME

Beta Regression with Autoregressive Errors for Interrupted Time Series Analysis of Proportion and Rate Outcomes: A Simulation Study

Interrupted time series analyses (ITSA) of proportion and rate outcomes are frequently estimated using ordinary least squares regression despite the bounded nature of these outcomes. When methods appropriate for bounded outcomes are used, the standard approach is a quasi-likelihood generalized linear model (GLM) with heteroskedasticity- and autocorrelation-consistent (HAC) standard errors. However, no existing estimator jointly models the beta-distributed conditional density and autoregressive (AR) error structure. We introduce betark, a Stata implementation of a joint conditional maximum likelihood estimator for beta regression with AR(k) errors based on a recursive substitution that yields a closed-form conditional beta likelihood with autoregressive dependence of arbitrary order. Unlike two-stage approaches, betark jointly estimates the mean, precision, and AR(k) coefficients in a single likelihood, so reported standard errors directly account for autocorrelation without separate correction. A Monte Carlo study compared betark with a quasi-binomial GLM using Newey-West HAC standard errors across AR(1)-AR(3) processes, three series lengths, and four effect sizes in a single-group ITSA design. Both methods were essentially unbiased, but betark produced better-calibrated inference than GLM+HAC in most scenarios, with the largest gains under highly persistent autocorrelation, where GLM+HAC Type I error exceeded 60% for short series. Misspecifying the AR order by one lag and varying the starting mean and pre-intervention trend had only modest effects on performance. However, betark's own Type I error remained elevated under highly persistent AR(3) processes even for the longest series examined.

stat.ME

Extending Prais-Winsten Regression to Panel Data with Higher-Order Autoregressive Errors: A Simulation Study

We extend the Prais-Winsten AR(k) generalized least squares (GLS) transformation to panel data within the Beck-Katz panel-corrected standard error (PCSE) framework and implement the method in the community-contributed Stata package xtpraisk. As the panel extension of Prais-Winsten, xtpraisk is the natural comparator to xtscc, the panel extension of Newey-West and implementation of the Driscoll-Kraay estimator. We conduct a Monte Carlo simulation to validate the statistical properties of xtpraisk and compare its finite-sample performance with xtscc. The simulation spans autoregressive orders 1-3, three autocorrelation scenarios, three panel sizes, six series lengths, and five effect sizes, with 2,000 replications per condition. Across all conditions, xtpraisk achieved higher power than xtscc while maintaining near-nominal Type I error rates, confidence interval coverage, and standard error calibration. In contrast, xtscc exhibited systematic standard error underestimation and inflated Type I error at short series lengths, with both deficiencies worsening as autoregressive order increased. Both estimators were essentially unbiased. Misspecification of the autoregressive order did not degrade xtpraisk's inferential performance, and cross-panel correlation and panel size had negligible effects on the relative performance of either estimator. The results indicate that xtpraisk is preferable when both statistical efficiency and valid inference are priorities, particularly under persistent higher-order autocorrelation and short to moderate series lengths.

stat.ME

A Goodness-of-Fit Test for Mixed-Effects Logistic Regression

Mixed-effects logistic regression is widely used for binary outcomes in hierarchical data, yet formal goodness-of-fit tests remain limited to random-intercept models and do not address sparse cluster settings. We extend a grouping-based Wald test to mixed-effects logistic models with random slopes. The procedure groups observations by predicted probabilities within clusters, augments the model with pooled group indicators, and tests their joint significance using a Wald statistic. To accommodate small clusters, we introduce a data-driven rule for selecting the number of groups, G=min(10,n_min), where n_min is the smallest cluster size, ensuring feasible estimation. Simulation studies across 24 null scenarios show that the test maintains nominal Type I error in three-level random slope models, including at smaller sample sizes than previously studied. The test exhibits increasing power to detect fixed-effects misspecification: power against omitted nonlinearity rises from 0.07 to 1.00 across effect sizes, and power against omitted interactions reaches 0.87. As expected, the test has no power to detect omission of a clustering level, reflecting its focus on residual structure in predicted probabilities. In sparse balanced designs, fixing G=10 leads to complete test failure, whereas the data-driven rule performs reliably. The method is implemented in the Stata program mlm_gof.

stat.ME

Multiple-group (Controlled) Interrupted Time Series Analysis with Higher-Order Autoregressive Errors: A Simulation Study Comparing Newey-West and Prais-Winsten Methods

Previous comparisons of ordinary least squares with Newey-West standard errors (OLS-NW) and Prais-Winsten (PW) regression in multiple-group interrupted time series analysis have been limited to first-order autoregressive (AR[1]) errors because PW estimation for higher-order AR[k] processes was previously unavailable. We conducted the first systematic evaluation of OLS-NW and PW under AR[2] and AR[3] error structures using Monte Carlo simulation. Simulations examined mild positive, oscillatory, and high persistent autocorrelation across varying series lengths and effect sizes. OLS-NW generally showed higher apparent power but substantially inflated Type I error and poor coverage, particularly under persistent autocorrelation, where inferential performance worsened with increasing AR order and series length. PW maintained substantially better inferential calibration across nearly all conditions. Both methods were approximately unbiased.

stat.AP

Improving causal inference in interrupted time series analysis: the triple difference design

Background: Interrupted time series analysis (ITSA) is widely used to evaluate health policy and intervention effects. While multiple-group ITSA (MG-ITSA) improves causal inference by incorporating a control group, residual confounding from unmeasured time-varying factors may remain. The triple-difference interrupted time series (DDD-ITSA) design extends this approach by adding a second control group to further isolate treatment effects, but it remains underutilized and lacks formal guidance. Methods: We formalize the DDD-ITSA framework, specify the regression model, define key parameters for estimating level and trend effects, and clarify interpretation of the triple-difference estimand. We illustrate the approach using a worked example evaluating California's Proposition 99 cigarette tax and its impact on per-capita cigarette sales. Results: In the example, all groups were balanced on pre-intervention level and trend. The triple-difference estimand indicated a statistically significant annual reduction of -1.76 per-capita cigarette packs in California relative to the secondary control (P = 0.020; 95 percent CI: -3.24, -0.28), consistent with results from the primary comparison. Differences between control groups were not significant. Conclusions: DDD-ITSA strengthens causal inference when two-group comparisons may be confounded by leveraging an additional control group to remove remaining biases and assess heterogeneity. Implementation is facilitated by updates to the itsa Stata package. Careful attention to control selection, baseline balance, and autocorrelation remains essential.

stat.AP