SearcharxivSearch

arXiv subjects

Stefan Wager

Publications and source records attributed to Stefan Wager.

At least 19 recordsLinked to original sources

What Would it Cost to End Extreme Poverty?

We study poverty minimization via direct transfers, framing this as a statistical learning problem while retaining the information constraints faced by real-world programs. Using nationally representative household consumption surveys from 34 countries that together account for 76% of the world's poor, we estimate that reducing the poverty rate to 1% (from a baseline of 13%) would cost $211 B nominal per year. This is 4.0 times the corresponding reduction in the aggregate poverty gap, but only 19% of the cost of universal basic income. Extrapolated globally, the results imply a cost of 0.28% of global GDP to (approximately) end extreme poverty.

econ.GN

What is the Long-Term Value of Reliability?

We describe Chronos LTV, a system to measure the long-term impact of delays and other service defects on key business metrics. We use Markov decision processes to model customer interactions over time, and formalize our target estimand as the marginal policy effect with respect to moving the average delay rate. Given this setup, we show that we can identify long-term effects under a sequential unconfoundedness assumption where delays are as good as random given observed order characteristics; and can estimate these effects using a simple covariate-balancing algorithm.

stat.ME

Unlocking Deep Demand Flexibility via Dynamic Signals

The rapid proliferation of distributed energy resources (DERs) and the electrification of residential loads offer significant potential for grid flexibility but pose stability challenges under static pricing regimes. Specifically, high levels of automation under static Time-of-Use (TOU) tariffs often induce ``device synchronization,'' where simultaneous responses from home energy management systems (HEMS) create artificial demand peaks that threaten grid stability. This paper proposes a privacy-preserving, one-way dynamic signaling framework to unlock deep demand flexibility from HEMS. We utilize a feedback-based learning algorithm that updates day-ahead price profiles based on aggregate substation demand and environmental contexts, effectively closing the loop between utility objectives and aggregated edge behaviors. The framework is rigorously validated using high-fidelity simulations on an 84-bus distribution network populated with hundreds of HEMS controlling diverse devices, including HVAC, PV, batteries, and flexible loads. Results demonstrate that the proposed mechanism achieves substantial reductions in both peak demand and total load variation. Extensive analyses across diverse climates and scalable deployments confirm the framework's robustness, indicating that dynamic pricing acts as a force multiplier for DERs, with peak shaving potential increasing significantly under high renewable penetration scenarios.

math.OC

Neyman Jackknife: Design-Based Variance Estimation for Causal Inference under Interference

We propose a framework, the Neyman Jackknife, for conservative variance estimation in finite-population causal inference under interference. Our approach provides a general, flexible blueprint that enables conservative variance estimation whenever we are able to recompute our target estimator with some treatment assignments omitted. In classical settings, our approach recovers estimators closely related to the Neyman estimator under SUTVA and the Newey-West HAC variance estimator for time series. Numerical experiments suggest that our general-purpose framework yields variance estimators that can match or even surpass the performance of baselines that were purpose-built for specific applications.

stat.ME

Estimating Dynamic Marginal Policy Effects under Sequential Unconfoundedness

We develop methods for estimating how infinitesimal policy changes affect long-term outcomes in dynamic systems. We show that dynamic marginal policy effects (MPEs) can be identified via tractable reduced-form expressions, and can be estimated under a general sequential unconfoundedness assumption. We also propose a doubly robust estimator for dynamic MPEs. Our approach does not require observing full dynamic state information (as is typically assumed for off-policy evaluation in Markov decision processes), and does not incur an exponential curse of horizon (as is typical in non-Markovian off-policy evaluation). We demonstrate practicality and robustness of our approach in a number of simulations, including one motivated by a dynamic pricing application where people use past prices to form a reference level for current prices.

stat.ME

Nonparametric Regression Discontinuity Designs with Survival Outcomes

Quasi-experimental evaluations are central for generating real-world causal evidence and complementing insights from randomized trials. The regression discontinuity design (RDD) is a quasi-experimental design that can be used to estimate the causal effect of treatments that are assigned based on a running variable crossing a threshold. Such threshold-based rules are ubiquitous in healthcare, where predictive and prognostic biomarkers frequently guide treatment decisions. However, standard RD estimators rely on complete outcome data, an assumption often violated in time-to-event analyses where censoring arises from loss to follow-up. To address this issue, we propose a nonparametric approach that leverages doubly robust censoring corrections and can be paired with existing RD estimators. Our approach can handle multiple survival endpoints, long follow-up times, and covariate-dependent variation in survival and censoring. We discuss the relevance of our approach across multiple areas of applications and demonstrate its usefulness through simulations and the prostate component of the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial where our new approach offers several advantages, including higher efficiency and robustness to misspecification. We have also developed an open-source software package, $\texttt{rdsurvival}$, for the $\texttt{R}$ language.

stat.ML

Sequentially-Rerandomized Switchback Experiments

Large-scale online platforms and marketplace systems often evaluate new policies through experiments that randomize treatment across operational units (e.g., geographies, regions, or clusters) over many time periods. In these settings, standard A/B testing can be inefficient or unreliable due to a limited number of units, substantial cross-unit heterogeneity, non-stationarity, and potential carryover across periods. We propose Sequentially-Rerandomized Switchback Experiments (SRSB), a new experimental design that helps mitigate these challenges. SRSB re-randomizes treatment at each time period such as to enforce balance on pre-specified prognostic variables constructed from past observations. In the absence of carryover, SRSB improves precision by leveraging temporal dependence through balancing lagged outcomes and covariates; we develop finite-sample randomization inference under a sharp null as well as asymptotic inference as the number of periods grows. We then extend SRSB to settings with first-order carryover and introduce a blocked SRSB variant that rerandomizes within strata defined by the previous treatment to form stable and comparable "stay" groups. Extensive simulations demonstrate the practical gains and robustness of SRSB relative to standard switchback designs.

stat.ME

Treatment effect estimation under convergent network interference

Under network interference, a unit's observed outcome depends on the treatment assignment of its neighboring units in an exposure graph. Existing design-based asymptotic theory typically considers local interference by restricting neighborhood sizes in the exposure graph. Such methods do not apply to dense exposure graphs, so prior work has often adopted a superpopulation approach instead, imposing regularity through random-graph models. In this paper, we introduce a notion of convergence for a sequence of finite populations under anonymous interference. Building on the graph limit framework of Lov\'{a}sz and Szegedy, we show that large-scale geometry of the exposure graph can provide a source of regularity beyond sparsity assumptions or random-graph modeling. Under Bernoulli assignment, our convergence notion yields asymptotic normality of standard estimators for the average direct effect, even on dense, non-random exposure graphs. As a special case, graphon-based random-graph models studied in prior work generate finite populations that converge in our sense. Under these models, graph randomness generates exposure graphs with stable large-scale geometry, while first-order uncertainty in average direct effect estimation is driven by treatment assignment.

math.ST

Non-parametric Causal Inference in Dynamic Thresholding Designs

We consider causal inference in dynamic settings where treatment is assigned by thresholding a state variable that can change over time. There is a large literature on regression-discontinuity methods building on the fact that, in the static setting, treatment assignment via threshold crossing induces a quasi-experimental design that enables pragmatic causal inference. But dynamic settings involve challenges not present in the static setting, e.g., past treatments may affect current state and thus future treatments, and so existing regression-discontinuity methods do not apply. Here, we show that dynamic thresholding designs identify a marginal policy effect that nests the classical regression-discontinuity parameter in the static setting; and propose a tailored local linear regression estimator that is consistent for this marginal policy effect. We demonstrate our approach using an experiment that emulates real-world optimization of thresholds for continuous glucose monitoring using data generated from an FDA-approved simulator.

stat.ME

Off-Policy Evaluation in Markov Decision Processes under Weak Distributional Overlap

Doubly robust methods hold considerable promise for off-policy evaluation in Markov decision processes (MDPs) under sequential ignorability: They have been shown to converge as $1/\sqrt{T}$ with the horizon $T$, to be statistically efficient in large samples, and to allow for modular implementation where preliminary estimation tasks can be executed using standard reinforcement learning techniques. Existing results, however, make heavy use of a strong distributional overlap assumption whereby the stationary distributions of the target policy and the data-collection policy are within a bounded factor of each other -- and this assumption is typically only credible when the state space of the MDP is bounded. In this paper, we re-visit the task of off-policy evaluation in MDPs under a weaker notion of distributional overlap, and introduce a class of truncated doubly robust (TDR) estimators which we find to perform well in this setting. When the distribution ratio of the target and data-collection policies is square-integrable (but not necessarily bounded), our approach recovers the large-sample behavior previously established under strong distributional overlap. When this ratio is not square-integrable, TDR is still consistent but with a slower-than-$1/\sqrt{T}$-rate; furthermore, this rate of convergence is minimax over a class of MDPs defined only using mixing conditions. We validate our approach numerically and find that, in our experiments, appropriate truncation plays a major role in enabling accurate off-policy evaluation when strong distributional overlap does not hold.

stat.ML

Data Fusion for High-Resolution Estimation

High-resolution estimates of population health indicators are critical for precision public health. We propose a method for high-resolution estimation that fuses distinct data sources: an unbiased, low-resolution data source (e.g. aggregated administrative data) and a potentially biased, high-resolution data source (e.g. individual-level online survey responses). We assume that the potentially biased, high-resolution data source is generated from the population under a model of sampling bias where observables can have arbitrary impact on the probability of response but the difference in the log probabilities of response between units with the same observables is linear in the difference between sufficient statistics of their observables and outcomes. Our data fusion method learns a distribution that is closest (in the sense of KL divergence) to the online survey distribution and consistent with the aggregated administrative data and our model of sampling bias. This approach significantly reduces bias in high-resolution estimates compared to baselines that rely on a single data source alone on a testbed that includes repeated measurements of three indicators measured by both the (online) Household Pulse Survey and ground-truth data sources at two geographic resolutions over the same time period.

stat.ME

Optimal Targeting in Dynamic Systems

Modern treatment targeting methods often rely on estimating a conditional average treatment effect (CATE) using machine learning tools. While effective in identifying who benefits from treatment on the individual level, these approaches typically overlook system-level dynamics that may arise when treatments induce strain on shared capacity. We study the problem of targeting in Markovian systems, where treatment decisions must be made one at a time as units arrive, and early decisions can impact later outcomes through delayed or limited access to resources. We show that optimal policies in such settings compare CATE-like quantities to state-specific thresholds, where each threshold reflects the expected cumulative impact on the system of treating an additional individual in the given state. We propose an algorithm that augments standard CATE estimation with state-level value iteration to estimate these thresholds from observational data. Theoretical results establish consistency and convergence guarantees, and empirical studies demonstrate that our method improves long-run outcomes considerably relative to individual-level CATE targeting rules and generic offline reinforcement learning algorithms.

stat.ME

Admissibility of Completely Randomized Trials: A Large-Deviation Approach

When an experimenter has the option of running an adaptive trial, is it admissible to ignore this option and run a non-adaptive trial instead? We provide a negative answer to this question in the best-arm identification problem, where the experimenter aims to allocate measurement efforts judiciously to confidently deploy the most effective treatment arm. We find that, whenever there are at least three treatment arms, there exist simple adaptive designs that universally and strictly dominate non-adaptive completely randomized trials. This dominance is characterized by a notion called efficiency exponent, which quantifies a design's statistical efficiency when the experimental sample is large. Our analysis focuses on the class of batched arm elimination designs, which progressively eliminate underperforming arms at pre-specified batch intervals. We characterize simple sufficient conditions under which these designs universally and strictly dominate completely randomized trials. These results resolve the second open problem posed in Qin [2022].

stat.ML

Noise-Induced Randomization in Regression Discontinuity Designs

Regression discontinuity designs assess causal effects in settings where treatment is determined by whether an observed running variable crosses a pre-specified threshold. Here we propose a new approach to identification, estimation, and inference in regression discontinuity designs that uses knowledge about exogenous noise (e.g., measurement error) in the running variable. In our strategy, we weight treated and control units to balance a latent variable of which the running variable is a noisy measure. Our approach is driven by effective randomization provided by the noise in the running variable, and complements standard formal analyses that appeal to continuity arguments while ignoring the stochastic nature of the assignment mechanism.

stat.ME

Policy Learning with Competing Agents

Decision makers often aim to learn a treatment assignment policy under a capacity constraint on the number of agents that they can treat. When agents can respond strategically to such policies, competition arises, complicating estimation of the optimal policy. In this paper, we study capacity-constrained treatment assignment in the presence of such interference. We consider a dynamic model where the decision maker allocates treatments at each time step and heterogeneous agents myopically best respond to the previous treatment assignment policy. When the number of agents is large but finite, we show that the threshold for receiving treatment under a given policy converges to the policy's mean-field equilibrium threshold. Based on this result, we develop a consistent estimator for the policy gradient. In a semi-synthetic experiment with data from the National Education Longitudinal Study of 1988, we demonstrate that this estimator can be used for learning capacity-constrained policies in the presence of strategic behavior.

stat.ML

Optimal Mechanisms for Demand Response: An Indifference Set Approach

The time at which renewable (e.g., solar or wind) energy resources produce electricity cannot generally be controlled. In many settings, however, consumers have some flexibility in their energy consumption needs, and there is growing interest in demand-response programs that leverage this flexibility to shift energy consumption to better match renewable production -- thus enabling more efficient utilization of these resources. We study optimal demand response in a setting where consumers use home energy management systems (HEMS) to autonomously adjust their electricity consumption. Our core assumption is that HEMS operationalize flexibility by querying the consumer for their preferences and computing the ``indifference set'' of all energy consumption profiles that can be used to satisfy these preferences. Then, given an indifference set, HEMS can respond to grid signals while guaranteeing user-defined comfort and functionality; e.g., if a consumer sets a temperature range, a HEMS can precool and preheat to align with peak renewable production, thus improving efficiency without sacrificing comfort. We show that while price-based mechanisms are not generally optimal for demand response, they become asymptotically optimal in large markets under a mean-field limit. Furthermore, we show that optimal dynamic prices can be efficiently computed in large markets by only querying HEMS about their planned consumption under different price signals. Using an OpenDSS-powered grid simulation for Phoenix, Arizona, we demonstrate that our approach enables meaningful demand response without creating grid instability.

math.OC

PLRD: Partially Linear Regression Discontinuity Inference

Regression discontinuity designs have become one of the most popular research designs in empirical economics. We argue, however, that the widely used approaches to building confidence intervals in regression discontinuity designs often exhibit suboptimal behavior in practice. We propose a new estimator, the partially linear regression discontinuity (PLRD) estimator that, in set of a simulation studies carefully calibrated to twelve high-profile applications of regression discontinuity designs, has substantially lower estimation error than available comparison methods. Throughout our experiments, the confidence intervals built using PLRD are both valid and short. We also provide large-sample guarantees for PLRD. Our simulation study serves as a general template for how new econometric methods can be credibly evaluated relative to the existing alternatives by constructing simulation designs that generate synthetic data indistinguishable from the original data using the Wasserstein generative adversarial network methodology.

econ.EM

Treatment Effects in Market Equilibrium

Policy-relevant treatment effect estimation in a marketplace setting requires taking into account both the direct benefit of the treatment and any spillovers induced by changes to the market equilibrium. The standard way to address these challenges is to evaluate interventions via cluster-randomized experiments, where each cluster corresponds to an isolated market. This approach, however, cannot be used when we only have access to a single market (or a small number of markets). Here, we show how to identify and estimate policy-relevant treatment effects using a unit-level randomized trial run within a single large market. A standard Bernoulli-randomized trial allows consistent estimation of direct effects, and of treatment heterogeneity measures that can be used for welfare-improving targeting. Estimating spillovers - as well as providing confidence intervals for the direct effect - requires estimates of price elasticities, which we provide using an augmented experimental design. Our results rely on all spillovers being mediated via the (observed) prices of a finite number of traded goods, and the market power of any single unit decaying as the market gets large. We illustrate our results using a simulation calibrated to a conditional cash transfer experiment in the Philippines.

econ.EM