SearcharxivSearch

arXiv subjects

Taisuke Otsu

Publications and source records attributed to Taisuke Otsu.

At least 19 recordsLinked to original sources

Monotone Response for Random Objects

Monotone treatment response (MTR), monotone treatment selection (MTS), and monotone instrumental variable (MIV) assumptions are widely used to partially identify counterfactual mean outcomes, but existing analyses have focused almost exclusively on scalar outcomes. We develop a unified framework for partial identification with outcomes that take values in a general metric space under these monotonicity restrictions by embedding the metric space into an $L^2$ space and imposing coordinatewise monotonicity on the embedded functions. The proposed framework yields valid identified sets for Fr\'echet means in a broad class of random-object spaces and further delivers sharp identification results for distributional outcomes under the Wasserstein metric, interval-valued outcomes represented by support functions, and compositional outcomes under the Aitchison metric. We also establish a support-free characterization of the identified set under the joint MTR--MTS assumption. Numerical and empirical illustrations based on Job Corps earnings data and periodontal health distributions from the National Health and Nutrition Examination Survey demonstrate the empirical usefulness of the proposed framework.

econ.EM

Lee Bounds for Random Objects

In applied research, Lee (2009) bounds are widely applied to bound the average treatment effect in the presence of selection bias. This paper extends the methodology of Lee bounds to accommodate outcomes in a general metric space, such as compositional and distributional data. By exploiting a representation of the Fréchet mean of the potential outcome via embedding in an Euclidean or Hilbert space, we present a feasible characterization of the identified set of the causal effect of interest, and then propose its analog estimator and bootstrap confidence region. The proposed method is illustrated by numerical examples on compositional and distributional data.

econ.EM

Difference-in-Differences with Interval Data

Difference-in-differences (DID) is one of the most popular tools used to evaluate causal effects of policy interventions. This paper extends the DID methodology to accommodate interval outcomes, which are often encountered in empirical studies using survey or administrative data. We point out that a naive application or extension of the conventional parallel trends assumption may yield uninformative or counterintuitive results, and present a suitable identification strategy, called parallel shifts, which exhibits desirable properties. Practical attractiveness of the proposed method is illustrated by revisiting an influential minimum wage study by Card and Krueger (1994).

econ.EM

Regression adjustment in completely randomized experiments with many covariates

This paper investigates estimation and inference for average treatment effects in completely randomized experiments when researchers observe potentially many covariates. Within Neyman's (1923) design-based framework, allowing the number of covariates to grow more slowly than the sample size, we demonstrate that a cross-fitted regression adjustment estimator--adapted from Aronow and Middleton (2013)--exhibits more favorable asymptotic properties than existing alternatives, such as Lin's (2013) regression adjustment estimator and the bias-corrected estimator of Lei and Ding (2021). For inference, we derive the first- and second-order terms in the stochastic expansions of regression-adjusted estimators, analyze the higher-order behavior of existing inference procedures, and introduce a modified version of the HC3 standard error. The proposed methods extend naturally to stratified experiments with large strata. Simulation studies show that the cross-fitted estimator, in combination with the modified HC3, provides accurate point estimates and reliable size control across a wide range of data-generating processes.

econ.EM

Random sets from the perspective of metric statistics

Since the seminal work by Beresteanu and Molinari(2008), the random set theory and related inference methods have been widely applied in partially identified econometric models. Meanwhile, there is an emerging field in statistics for studying random objects in metric spaces, called metric statistics. This paper clarifies a relationship between two fundamental concepts in these literatures, the Aumann and Fréchet means, and presents some applications of metric statistics to econometric problems involving random sets.

math.ST

Empirical Likelihood for Random Forests and Ensembles

We develop an empirical likelihood (EL) framework for random forests and related ensemble methods, providing a likelihood-based approach to quantify their statistical uncertainty. Exploiting the incomplete $U$-statistic structure inherent in ensemble predictions, we construct an EL statistic that is asymptotically chi-squared when subsampling induced by incompleteness is not overly sparse. Under sparser subsampling regimes, the EL statistic tends to over-cover due to loss of pivotality; we therefore propose a modified EL that restores pivotality through a simple adjustment. Our method retains key properties of EL while remaining computationally efficient. Theory for honest random forests and simulations demonstrate that modified EL achieves accurate coverage and practical reliability relative to existing inference methods.

stat.ML

On Gaussian Approximation for M-Estimator

This study develops a non-asymptotic Gaussian approximation theory for distributions of M-estimators, which are defined as maximizers of empirical criterion functions. In existing mathematical statistics literature, numerous studies have focused on approximating the distributions of the M-estimators for statistical inference. In contrast to the existing approaches, which mainly focus on limiting behaviors, this study employs a non-asymptotic approach, establishes abstract Gaussian approximation results for maximizers of empirical criteria, and proposes a Gaussian multiplier bootstrap approximation method. Our developments can be considered as extensions of the seminal works (Chernozhukov, Chetverikov and Kato (2013, 2014, 2015)) on the approximation theory for distributions of suprema of empirical processes toward their maximizers. Through this work, we shed new lights on the statistical theory of M-estimators. Our theory covers not only regular estimators, such as the least absolute deviations, but also some non-regular cases where it is difficult to derive or to approximate numerically the limiting distributions such as non-Donsker classes and cube root estimators.

math.ST

Regression Discontinuity Designs for Functional Data and Random Objects in Geodesic Spaces

Regression discontinuity designs (RDDs) are widely used for causal inference in observational studies with cutoff-based treatment assignment, primarily for Euclidean outcomes. We propose the geodesic regression discontinuity design (GRDD), which extends RDDs to complex non-Euclidean outcomes, including networks, compositional data, functional data, and other random objects in geodesic metric spaces. Since algebraic operations are unavailable in such spaces, we define the causal effect at the cutoff as the geodesic connecting the local Fr\'echet means of untreated and treated outcomes, recovering the classical local average treatment effect in the scalar case. Estimation is conducted intrinsically via local Fr\'echet regression to preserve geometric validity and interpretability. For inference, we adopt an extrinsic approach by embedding the metric space into a Hilbert space, enabling tractable asymptotic analysis. We establish asymptotic normality and develop bootstrap-based procedures for hypothesis testing and confidence intervals for the treatment effect magnitude. We also propose a data-adaptive bandwidth selection method tailored to RDDs in metric spaces and study its empirical performance. Applications include compositional voting outcomes in UK elections and daily CO concentration curves after the Taipei metro introduction, and we extend the framework to fuzzy designs with imperfect compliance.

stat.ME

Geodesic Synthetic Control Methods for Random Objects and Functional Data

We introduce a geodesic synthetic control method for causal inference with panel outcomes that are random objects in a geodesic metric space. Examples include distributions, compositions, networks, symmetric positive-definite matrices, trees and functional data, among other data types that require a geometry beyond Euclidean vector spaces. The proposed method replaces Euclidean weighted averages by weighted Fr\'echet means and defines treatment effects as geodesics connecting untreated and treated potential outcomes. We develop a causal model with stochastic perturbations for object-valued untreated outcomes, establish consistency of the estimated weights for the perturbed population target, derive a bound on post-treatment prediction error, and give sufficient conditions for consistent recovery of the untreated counterfactual. For uncertainty quantification, we propose an intrinsic metric conformal calibration procedure that yields prediction regions for untreated counterfactual objects, confidence sets for geodesic treatment effects, confidence intervals for their magnitudes, and global no-effect tests. The method is illustrated through simulation studies for networks and symmetric positive-definite matrices, and through applications to employment composition changes following the 2011 Great East Japan Earthquake and the impact of abortion liberalization on fertility patterns in East Germany.

stat.ME

Geodesic Difference-in-Differences

Difference-in-differences (DID) is a widely used quasi-experimental design for causal inference, traditionally applied to scalar or Euclidean outcomes, while extensions to outcomes residing in non-Euclidean spaces remain limited. Existing methods for such outcomes have primarily focused on univariate distributions, leveraging linear operations in the space of quantile functions, but these approaches cannot be directly extended to outcomes in general metric spaces. In this paper, we propose geodesic DID, a novel DID framework for outcomes in uniquely geodesic metric spaces that admit a geodesic transport structure, including distributions, networks, and manifold-valued data. To address the absence of algebraic operations in these spaces, we use geodesics as proxies for differences and introduce the geodesic average treatment effect on the treated (ATT) as the causal estimand. We establish the identification of the geodesic ATT and derive the convergence rate of its sample versions, employing tools from metric geometry and empirical process theory. This framework is further extended to the case of staggered DID settings, allowing for multiple time periods and varying treatment timings. To illustrate the practical utility of geodesic DID, we analyze health impacts of the Soviet Union's collapse using age-at-death distributions and assess effects of U.S. electricity market liberalization on electricity generation compositions.

stat.ME

Isotonic propensity score matching

We propose a one-to-many matching estimator of the average treatment effect based on propensity scores estimated by isotonic regression. This approach is predicated on the assumption of monotonicity in the propensity score function, a condition that can be justified in many economic applications. We show that the nature of the isotonic estimator can help us to fix many problems of existing matching methods, including efficiency, choice of the number of matches, choice of tuning parameters, robustness to propensity score misspecification, and bootstrap validity. As a by-product, a uniformly consistent isotonic estimator is developed for our proposed matching method.

econ.EM

Multiway empirical likelihood

This paper develops a general methodology to conduct statistical inference for observations indexed by multiple sets of entities. We propose a novel multiway empirical likelihood statistic that converges to a chi-square distribution under the non-degenerate case, where corresponding Hoeffding type decomposition is dominated by linear terms. Our methodology is related to the notion of jackknife empirical likelihood but the leave-out pseudo values are constructed by leaving columns or rows. We further develop a modified version of our multiway empirical likelihood statistic, which converges to a chi-square distribution regardless of the degeneracy, and discover its desirable higher-order property compared to the t-ratio by the conventional Eicker-White type variance estimator. The proposed methodology is illustrated by several important statistical problems, such as bipartite network, generalized estimating equations, and three-way observations.

stat.ME

Geodesic Causal Inference

Adjusting for confounding and imbalance when establishing statistical relationships is an increasingly important task, and causal inference methods have emerged as the most popular tool to achieve this. Existing methodology has been developed primarily for outcomes that lie in Euclidean spaces. We introduce here a general framework for causal inference when outcomes reside in general geodesic metric spaces, where we draw on a novel geodesic calculus that facilitates scalar multiplication for geodesics and the quantification of treatment effects through the concept of geodesic average treatment effect. Using ideas from Fr\'echet regression, we obtain a doubly robust estimation of the geodesic average treatment effect and results on consistency and rates of convergence for the proposed estimators. We also develop an intrinsic uncertainty quantification framework for the treatment effect based on Fr\'echet objective functions. The proposed framework is illustrated through simulations and real data applications, including network-valued outcomes from New York City taxi trips to assess the impact of the COVID-19 pandemic, and compositional data on U.S. state-level energy sources to study the effect of coal mining.

stat.ME

Optimal testing in a class of nonregular models

This paper studies optimal hypothesis testing for nonregular econometric models with parameter-dependent support. We consider both one-sided and two-sided hypothesis testing and develop asymptotically uniformly most powerful tests based on a limit experiment. Our two-sided test becomes asymptotically uniformly most powerful without imposing further restrictions such as unbiasedness, and can be inverted to construct a confidence set for the nonregular parameter. Simulation results illustrate desirable finite-sample properties of the proposed tests.

math.ST

Regression Discontinuity Design with Potentially Many Covariates

This paper studies the case of possibly high-dimensional covariates in the regression discontinuity design (RDD) analysis. In particular, we propose estimation and inference methods for the RDD models with covariate selection which perform stably regardless of the number of covariates. The proposed methods combine the local approach using kernel weights with $\ell_{1}$-penalization to handle high-dimensional covariates. We provide theoretical and numerical results which illustrate the usefulness of the proposed methods. Theoretically, we present risk and coverage properties for our point estimation and inference methods, respectively. Under certain special case, the proposed estimator becomes more efficient than the conventional covariate adjusted estimator at the cost of an additional sparsity condition. Numerically, our simulation experiments and empirical example show the robust behaviors of the proposed methods to the number of covariates in terms of bias and variance for point estimation and coverage probability and interval length for inference.

econ.EM

Graph Neural Networks: Theory for Estimation with Application on Network Heterogeneity

This paper presents a novel application of graph neural networks for modeling and estimating network heterogeneity. Network heterogeneity is characterized by variations in unit's decisions or outcomes that depend not only on its own attributes but also on the conditions of its surrounding neighborhood. We delineate the convergence rate of the graph neural networks estimator, as well as its applicability in semiparametric causal inference with heterogeneous treatment effects. The finite-sample performance of our estimator is evaluated through Monte Carlo simulations. In an empirical setting related to microfinance program participation, we apply the new estimator to examine the average treatment effects and outcomes of counterfactual policies, and to propose an enhanced strategy for selecting the initial recipients of program information in social networks.

econ.EM

GLS under Monotone Heteroskedasticity

The generalized least square (GLS) is one of the most basic tools in regression analyses. A major issue in implementing the GLS is estimation of the conditional variance function of the error term, which typically requires a restrictive functional form assumption for parametric estimation or smoothing parameters for nonparametric estimation. In this paper, we propose an alternative approach to estimate the conditional variance function under nonparametric monotonicity constraints by utilizing the isotonic regression method. Our GLS estimator is shown to be asymptotically equivalent to the infeasible GLS estimator with knowledge of the conditional error variance, and involves only some tuning to trim boundary observations, not only for point estimation but also for interval estimation or hypothesis testing. Our analysis extends the scope of the isotonic regression method by showing that the isotonic estimates, possibly with generated variables, can be employed as first stage estimates to be plugged in for semiparametric objects. Simulation studies illustrate excellent finite sample performances of the proposed method. As an empirical example, we revisit Acemoglu and Restrepo's (2017) study on the relationship between an aging population and economic growth to illustrate how our GLS estimator effectively reduces estimation errors.

econ.EM

Conditional Likelihood Ratio Test with Many Weak Instruments

This paper extends validity of the conditional likelihood ratio (CLR) test developed by Moreira (2003) to instrumental variable regression models with unknown error variance and many weak instruments. In this setting, we argue that the conventional CLR test with estimated error variance loses exact similarity and is asymptotically invalid. We propose a modified critical value function for the likelihood ratio (LR) statistic with estimated error variance, and prove that this modified test achieves asymptotic validity under many weak instrument asymptotics. Our critical value function is constructed by representing the LR using four statistics, instead of two as in Moreira (2003). A simulation study illustrates the desirable properties of our test.

econ.EM