SearcharxivSearch

arXiv subjects

P. M. Aronow

Publications and source records attributed to P. M. Aronow.

At least 19 recordsLinked to original sources

Undocumented Behavior in the gsynth R package and its Consequences for Three Published Studies

Prior to the version 1.3.1 update on CRAN in December 2025, gsynth, a popular R package for estimating Interactive Fixed Effects (IFE) models, could drastically and systematically underestimate standard errors. This underestimation would occur when two estimation options (inference = "parametric", and EM = TRUE) were used together, in which case the package would apply a parametric bootstrap procedure to Gobillon and Magnac (2016)'s IFE-EM estimator. The package ceased supporting this combination in December 2025, and the latest documentation now describes the parametric bootstrap as not suitable for use with the IFE-EM estimator due to a theoretical incompatibility. Our focus is an implementation error we identified in the pre-1.3.1 versions of gsynth: the parametric bootstrap used when EM = TRUE did not match the algorithm proposed in Xu (2017), using in-sample residuals instead of out-of-sample errors. We show that this implementation error can cause underestimation by orders of magnitude. We conduct an empirical Monte Carlo study using randomly assigned placebo treatments on a series of state-level panel datasets, and show that gsynth could yield high false positive rates in realistic settings. We identify three papers published in the American Political Science Review that are affected by this behavior. Reanalyzing the relevant sections of these papers, we show that (i) correcting the implementation error renders most findings insignificant, and (ii) using Xu (2017)'s Generalized Synthetic Control method in place of IFE-EM renders every finding insignificant.

stat.ME

Overlap structure of mixed even $p$-spin models

We identify the limiting overlap structure of mixed even $p$-spin models, including the Sherrington-Kirkpatrick model, at every fixed deterministic external field. In particular, we show that the disorder-averaged distribution of the overlap, or its absolute value at zero field, converges to the Parisi measure. We identify the full limiting array using the overlap array of the associated Ruelle probability cascade. We also establish the corresponding Ghirlanda-Guerra identities and determine the limiting law of the quenched overlap distribution. For the zero-field Sherrington-Kirkpatrick model, we additionally prove temperature chaos at every pair of distinct nonnegative inverse temperatures.

math.PR

Nearly-minimax variance estimation under rough random design

We identify the minimax exponent for constant conditional variance estimation under rough random design. The unknown design density is bounded above and away from zero, with no smoothness assumption, and the conditional error laws may depend on the covariates and have uniformly bounded fourth moments. For an $s$-Hölder regression function with $s>1$ in dimension $d>4s$, the minimax root-mean-square risk lies between $cn^{-2(s+1)/(d+4)}e^{-C\sqrt{\log n}}$ and $Cn^{-2(s+1)/(d+4)}$. These bounds show that the rate proposed by Robins is not uniformly attainable over this model class. For $0 1$ and $d\le4s$, it is $n^{-1/2}$.

math.ST

A positive resolution of the gap-entropy conjecture

We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in $[0,1]$, and a unique optimal arm. For each suboptimal arm $i$, let $Δ_i=μ_*-μ_i$ be its gap from the optimal mean, and write $H=\sum_{i\ne *}Δ_i^{-2}$. Let $p_r$ be the fraction of $H$ contributed by arms with $2^{-(r+1)}<Δ_i\le2^{-r}$, and let $\mathrm{Ent}(I)=\sum_{r:p_r>0} p_r\log(1/p_r)$. Among all algorithms that identify the optimal arm with probability at least $1-δ$ on every Gaussian instance, the optimal expected number of samples on a given instance, averaged over all permutations of the arm labels, is within absolute constant factors of $H(\log(1/δ)+\mathrm{Ent}(I))$. Moreover, there is an algorithm, independent of the instance, whose expected number of samples is bounded by a constant multiple of this quantity plus $g^{-2}\log\log(e^e/g)$, where $g=\min_{i\ne *}Δ_i$ is the gap to the closest competitor.

cs.LG

Quantitative Parisi formulas and fluctuations in the Sherrington-Kirkpatrick model

We establish quantitative Parisi formulas and fluctuation bounds for the Sherrington-Kirkpatrick model at zero external field. In particular, using that the Parisi measure is supported on an interval, we show that for every fixed inverse temperature greater than one, the variance of the logarithmic partition function lies between orders $N^{4/15}$ and $N^{7/15}$. The same exponents hold for the ground-state variance. We also obtain upper and lower bounds on the finite-size bias of the mean free energy and ground-state energy, together with two-sided upper-tail estimates with exponent $6/5$ at both positive and zero temperature.

math.PR

Optimal central limit theorem for bounded random variables in high dimensions

Let $W=n^{-1/2}\sum_{i=1}^n X_i$, where the $X_i$ are independent centered random vectors in ${\mathbb R}^p$ with $|X_{ij}|\le B$ almost surely. Suppose that $\text{Cov}(W)$ has unit diagonal and smallest eigenvalue at least $b^2>0$. We prove that the distance between $W$ and a Gaussian vector with the same covariance, uniformly over axis-aligned rectangles, is at most $C\min\{1,b^{-2}Bn^{-1/2}\log^{3/2}(ep)\}$. For fixed $b$, the dependence on summand size and dimension matches known lower bounds in growing-dimensional regimes. The proof combines a concentration estimate near rectangle boundaries with a carefully chosen Gaussian comparison.

math.PR

Replica symmetry breaking at arbitrary depth

We study finite-step replica symmetry breaking (RSB) in mixed even $p$-spin Ising spin glasses at positive temperature and zero external field. For every integer $k\geq1$, we construct finite polynomial mixtures whose Parisi measures are exactly $k$-RSB, with each such phase persisting on an open set of coefficients. More generally, for any admissible finite polynomial covariance $Ξ$ with a nondegenerate $r$-RSB Parisi measure, adding $λq^p$ for sufficiently large even $p$ produces, as $λ$ varies near a critical value, a unique $r$-RSB-to-$(r+1)$-RSB transition. Since such covariances exist at every finite RSB depth, this yields transitions of arbitrary depth. The main idea of the proof is that for large $p$, the perturbation $λq^p$ is negligible for overlaps bounded away from $1$, while its effect is concentrated at overlaps very close to $1$.

math.PR

On the Foundations of the Design-Based Approach

The design-based paradigm may be adopted in causal inference and survey sampling when we assume Rubin's stable unit treatment value assumption (SUTVA) or impose similar frameworks. While often taken for granted, such assumptions entail strong claims about the data generating process. We develop an alternative design-based approach: we first invoke a generalized, non-parametric model that allows for unrestricted forms of interference, such as spillover. We define an associated set of inferential targets and discuss their interpretation under SUTVA and a weaker assumption that we call the No Unmodeled Revealable Variation Assumption (NURVA). We then reconstruct the standard paradigm, reconsidering SUTVA at the end rather than assuming it at the beginning. Despite its similarity to SUTVA, we demonstrate the practical limitations of NURVA alone for identifying substantively interesting quantities. In so doing, we provide clarity on the nature and importance of SUTVA for applied research.

stat.ME

Adaptive Confidence Sets for Binary Regression without Design Smoothness

We study honest adaptive confidence sets for the regression function in random-design binary regression under $L^2(dx)$ loss. Assuming only known bounds $0<c\leq g\leq C<\infty$ on the unknown design density, we construct asymptotically honest, rate-adaptive confidence sets without requiring $g$ to be smooth. Full adaptation is possible when the range of regression-function smoothness spans at most a factor of two. Over wider smoothness ranges, adaptation is achieved on the usual separated classes at the corresponding testing rates $n^{-2s/(4s+d)}$. A lower bound under the uniform design shows that these separation rates are rate-optimal. This answers a question raised by Mukherjee and Sen (2018).

math.ST

Maximum cut and maximum bisection in random regular graphs

We prove that, for every fixed degree $d$ and a uniformly random simple $d$-regular graph $G_{n,d}$ on an even number $n>d$ of vertices, $$\mathbb{E}\operatorname{MaxCut}(G_{n,d})-\mathbb{E}\operatorname{MaxBis}(G_{n,d})=o(n).$$ In other words, requiring the two sides of a cut to have exactly the same size changes the expected optimal cut by only a sublinear number of edges. We prove this first for the configuration model by comparing cuts of each possible cardinality on an $n$-vertex graph with bisections of a related graph on $2n$ vertices. An important technical tool is Huang's interpolation theorem (Huang, 2018); to apply it, we establish a structural property of the change in the optimal cut when a single edge is added. This comparison, together with concentration estimates and Huang's convergence theorem for maximum bisection, shows that maximum cut and maximum bisection have the same limiting density. Conditioning the configuration model on being simple then gives the result for uniformly random simple regular graphs.

math.CO

Minimax unbiased estimation for finite populations with bounded outcomes

We study design-unbiased estimation of the finite-population total $\sum_{i=1}^N y_i$ when each outcome satisfies known bounds $y_i\in[a_i,b_i]$. For any sampling design with inclusion probabilities $π_i>0$, we prove a sharp lower bound on the worst-case squared error over the rectangular parameter space. This bound is attained if and only if the unit inclusion indicators are pairwise independent, in which case the minimax estimator is the midpoint-differenced Horvitz-Thompson estimator $\sum_{i=1}^N m_i+\sum_{i\in S}(y_i-m_i)/π_i$, with $m_i=(a_i+b_i)/{2}$. We then solve the joint design-and-estimation problem under the constraint $\sum_i π_i\le n$. We find that a minimax strategy samples units independently with probabilities $π_i^\ast=\min(1,c (b_i-a_i))$ where $c>0$ is chosen so that $\sum_i π_i^\ast=n$, and uses the midpoint-differenced estimator. This extends Gabler (1990)'s linear minimax result to the full class of design-unbiased estimators. We also show that the estimator is admissible among unbiased estimators and affine equivariant.

math.ST

Using LLMs to Directly Guess Conditional Expectations Can Improve Efficiency in Causal Estimation

We propose a simple yet effective use of LLM-powered AI tools to improve causal estimation. In double machine learning, the accuracy of causal estimates of the effect of a treatment on an outcome in the presence of a high-dimensional confounder depends on the performance of estimators of conditional expectation functions. We show that predictions made by generative models trained on historical data can be used to improve the performance of these estimators relative to approaches that solely rely on adjusting for embeddings extracted from these models. We argue that the historical knowledge and reasoning capacities associated with these generative models can help overcome curse-of-dimensionality problems in causal inference problems. We consider a case study using a small dataset of online jewelry auctions, and demonstrate that inclusion of LLM-generated guesses as predictors can improve efficiency in estimation.

cs.LG

Fast computation of exact confidence intervals for randomized experiments with binary outcomes

Given a randomized experiment with binary outcomes, exact confidence intervals for the average causal effect of the treatment can be computed through a series of permutation tests. This approach requires minimal assumptions and is valid for all sample sizes, as it does not rely on large-sample approximations such as those implied by the central limit theorem. We show that these confidence intervals can be found in $O(n \log n)$ permutation tests in the case of balanced designs, where the treatment and control groups have equal sizes, and $O(n^2)$ permutation tests in the general case. Prior to this work, the most efficient known constructions required $O(n^2)$ such tests in the balanced case [Li and Ding, 2016], and $O(n^4)$ tests in the general case [Rigdon and Hudgens, 2015]. Our results thus facilitate exact inference as a viable option for randomized experiments far larger than those accessible by previous methods. We also generalize our construction to produce confidence intervals for other causal estimands, including the relative risk ratio and odds ratio, yielding similar computational gains.

stat.ME

Randomization-based confidence sets for the local average treatment effect

We consider the problem of generating confidence sets in randomized experiments with noncompliance. We show that a refinement of a randomization-based procedure proposed by Imbens and Rosenbaum (2005) has desirable properties. Namely, we show that using a studentized Anderson--Rubin-type statistic as a test statistic yields confidence sets that are finite-sample exact under treatment effect homogeneity, and remain asymptotically valid for the Local Average Treatment Effect when the treatment effect is heterogeneous. We provide a uniform analysis of this procedure and efficient algorithms to construct the confidence set.

math.ST

Design-Based Inference for Spatial Experiments under Unknown Interference

We consider design-based causal inference for spatial experiments in which treatments may have effects that bleed out and feed back in complex ways. Such spatial spillover effects violate the standard ``no interference'' assumption for standard causal inference methods. The complexity of spatial spillover effects also raises the risk of misspecification and bias in model-based analyses. We offer an approach for robust inference in such settings without having to specify a parametric outcome model. We define a spatial ``average marginalized effect'' (AME) that characterizes how, in expectation, units of observation that are a specified distance from an intervention location are affected by treatment at that location, averaging over effects emanating from other intervention nodes. We show that randomization is sufficient for non-parametric identification of the AME even if the nature of interference is unknown. Under mild restrictions on the extent of interference, we establish asymptotic distributions of estimators and provide methods for both sample-theoretic and randomization-based inference. We show conditions under which the AME recovers a structural effect. We illustrate our approach with a simulation study. Then we re-analyze a randomized field experiment and a quasi-experiment on forest conservation, showing how our approach offers robust inference on policy-relevant spillover effects.

stat.ME

Operationalizing Counterfactual Metrics: Incentives, Ranking, and Information Asymmetry

From the social sciences to machine learning, it has been well documented that metrics to be optimized are not always aligned with social welfare. In healthcare, Dranove et al. (2003) showed that publishing surgery mortality metrics actually harmed the welfare of sicker patients by increasing provider selection behavior. We analyze the incentive misalignments that arise from such average treated outcome metrics, and show that the incentives driving treatment decisions would align with maximizing total patient welfare if the metrics (i) accounted for counterfactual untreated outcomes and (ii) considered total welfare instead of averaging over treated patients. Operationalizing this, we show how counterfactual metrics can be modified to behave reasonably in patient-facing ranking systems. Extending to realistic settings when providers observe more about patients than the regulatory agencies do, we bound the decay in performance by the degree of information asymmetry between principal and agent. In doing so, our model connects principal-agent information asymmetry with unobserved heterogeneity in causal inference.

cs.LG

Dyadic Clustering in International Relations

Quantitative empirical inquiry in international relations often relies on dyadic data. Standard analytic techniques do not account for the fact that dyads are not generally independent of one another. That is, when dyads share a constituent member (e.g., a common country), they may be statistically dependent, or "clustered." Recent work has developed dyadic clustering robust standard errors (DCRSEs) that account for this dependence. Using these DCRSEs, we reanalyzed all empirical articles published in International Organization between January 2014 and January 2020 that feature dyadic data. We find that published standard errors for key explanatory variables are, on average, approximately half as large as DCRSEs, suggesting that dyadic clustering is leading researchers to severely underestimate uncertainty. However, most (67% of) statistically significant findings remain statistically significant when using DCRSEs. We conclude that accounting for dyadic clustering is both important and feasible, and offer software in R and Stata to facilitate use of DCRSEs in future research.

stat.ME

On the reliability of published findings using the regression discontinuity design in political science

The regression discontinuity (RD) design offers identification of causal effects under weak assumptions, earning it a position as a standard method in modern political science research. But identification does not necessarily imply that causal effects can be estimated accurately with limited data. In this paper, we highlight that estimation under the RD design involves serious statistical challenges and investigate how these challenges manifest themselves in the empirical literature in political science. We collect all RD-based findings published in top political science journals in the period 2009-2018. The distribution of published results exhibits pathological features; estimates tend to bunch just above the conventional level of statistical significance. A reanalysis of all studies with available data suggests that researcher discretion is not a major driver of these features. However, researchers tend to use inappropriate methods for inference, rendering standard errors artificially small. A retrospective power analysis reveals that most of these studies were underpowered to detect all but large effects. The issues we uncover, combined with well-documented selection pressures in academic publishing, cause concern that many published findings using the RD design may be exaggerated.

stat.ME