SearcharxivSearch

arXiv subjects

Daisuke Kurisu

Publications and source records attributed to Daisuke Kurisu.

At least 19 recordsLinked to original sources

Monotone Response for Random Objects

Monotone treatment response (MTR), monotone treatment selection (MTS), and monotone instrumental variable (MIV) assumptions are widely used to partially identify counterfactual mean outcomes, but existing analyses have focused almost exclusively on scalar outcomes. We develop a unified framework for partial identification with outcomes that take values in a general metric space under these monotonicity restrictions by embedding the metric space into an $L^2$ space and imposing coordinatewise monotonicity on the embedded functions. The proposed framework yields valid identified sets for Fr\'echet means in a broad class of random-object spaces and further delivers sharp identification results for distributional outcomes under the Wasserstein metric, interval-valued outcomes represented by support functions, and compositional outcomes under the Aitchison metric. We also establish a support-free characterization of the identified set under the joint MTR--MTS assumption. Numerical and empirical illustrations based on Job Corps earnings data and periodontal health distributions from the National Health and Nutrition Examination Survey demonstrate the empirical usefulness of the proposed framework.

econ.EM

Lee Bounds for Random Objects

In applied research, Lee (2009) bounds are widely applied to bound the average treatment effect in the presence of selection bias. This paper extends the methodology of Lee bounds to accommodate outcomes in a general metric space, such as compositional and distributional data. By exploiting a representation of the Fr\'echet mean of the potential outcome via embedding in an Euclidean or Hilbert space, we present a feasible characterization of the identified set of the causal effect of interest, and then propose its analog estimator and bootstrap confidence region. The proposed method is illustrated by numerical examples on compositional and distributional data.

econ.EM

Functional Synthetic Control Methods for Metric Space-Valued Outcomes

The synthetic control method (SCM) is a widely used tool for evaluating causal effects of policy changes in panel data settings. Recent studies have extended its framework to accommodate complex outcomes that take values in metric spaces, such as distributions, functions, networks, covariance matrices, and compositional data. However, due to the lack of linear structure in general metric spaces, theoretical guarantees for estimation and inference within these extended frameworks remain underdeveloped. In this study, we propose the functional synthetic control (FSC) method as an extension of the SCM for metric space-valued outcomes. To address challenges arising from the nonlinearlity of metric spaces, we leverage isometric embeddings into Hilbert spaces. Building on this approach, we develop the FSC and augmented FSC estimators for counterfactual outcomes, with the latter being a bias-corrected version of the former. We then derive their finite-sample error bounds to establish theoretical guarantees for estimation, and construct prediction sets based on these estimators to conduct inference on causal effects. We demonstrate the usefulness of the proposed framework through simulation studies and three empirical applications.

stat.ME

Difference-in-Differences with Interval Data

Difference-in-differences (DID) is one of the most popular tools used to evaluate causal effects of policy interventions. This paper extends the DID methodology to accommodate interval outcomes, which are often encountered in empirical studies using survey or administrative data. We point out that a naive application or extension of the conventional parallel trends assumption may yield uninformative or counterintuitive results, and present a suitable identification strategy, called parallel shifts, which exhibits desirable properties. Practical attractiveness of the proposed method is illustrated by revisiting an influential minimum wage study by Card and Krueger (1994).

econ.EM

Random sets from the perspective of metric statistics

Since the seminal work by Beresteanu and Molinari(2008), the random set theory and related inference methods have been widely applied in partially identified econometric models. Meanwhile, there is an emerging field in statistics for studying random objects in metric spaces, called metric statistics. This paper clarifies a relationship between two fundamental concepts in these literatures, the Aumann and Fr\'echet means, and presents some applications of metric statistics to econometric problems involving random sets.

math.ST

Regression Discontinuity Designs for Functional Data and Random Objects in Geodesic Spaces

Regression discontinuity designs (RDDs) are widely used for causal inference in observational studies with cutoff-based treatment assignment, primarily for Euclidean outcomes. We propose the geodesic regression discontinuity design (GRDD), which extends RDDs to complex non-Euclidean outcomes, including networks, compositional data, functional data, and other random objects in geodesic metric spaces. Since algebraic operations are unavailable in such spaces, we define the causal effect at the cutoff as the geodesic connecting the local Fr\'echet means of untreated and treated outcomes, recovering the classical local average treatment effect in the scalar case. Estimation is conducted intrinsically via local Fr\'echet regression to preserve geometric validity and interpretability. For inference, we adopt an extrinsic approach by embedding the metric space into a Hilbert space, enabling tractable asymptotic analysis. We establish asymptotic normality and develop bootstrap-based procedures for hypothesis testing and confidence intervals for the treatment effect magnitude. We also propose a data-adaptive bandwidth selection method tailored to RDDs in metric spaces and study its empirical performance. Applications include compositional voting outcomes in UK elections and daily CO concentration curves after the Taipei metro introduction, and we extend the framework to fuzzy designs with imperfect compliance.

stat.ME

Geodesic Synthetic Control Methods for Random Objects and Functional Data

We introduce a geodesic synthetic control method for causal inference with panel outcomes that are random objects in a geodesic metric space. Examples include distributions, compositions, networks, symmetric positive-definite matrices, trees and functional data, among other data types that require a geometry beyond Euclidean vector spaces. The proposed method replaces Euclidean weighted averages by weighted Fr\'echet means and defines treatment effects as geodesics connecting untreated and treated potential outcomes. We develop a causal model with stochastic perturbations for object-valued untreated outcomes, establish consistency of the estimated weights for the perturbed population target, derive a bound on post-treatment prediction error, and give sufficient conditions for consistent recovery of the untreated counterfactual. For uncertainty quantification, we propose an intrinsic metric conformal calibration procedure that yields prediction regions for untreated counterfactual objects, confidence sets for geodesic treatment effects, confidence intervals for their magnitudes, and global no-effect tests. The method is illustrated through simulation studies for networks and symmetric positive-definite matrices, and through applications to employment composition changes following the 2011 Great East Japan Earthquake and the impact of abortion liberalization on fertility patterns in East Germany.

stat.ME

Geodesic Difference-in-Differences

Difference-in-differences (DID) is a widely used quasi-experimental design for causal inference, traditionally applied to scalar or Euclidean outcomes, while extensions to outcomes residing in non-Euclidean spaces remain limited. Existing methods for such outcomes have primarily focused on univariate distributions, leveraging linear operations in the space of quantile functions, but these approaches cannot be directly extended to outcomes in general metric spaces. In this paper, we propose geodesic DID, a novel DID framework for outcomes in uniquely geodesic metric spaces that admit a geodesic transport structure, including distributions, networks, and manifold-valued data. To address the absence of algebraic operations in these spaces, we use geodesics as proxies for differences and introduce the geodesic average treatment effect on the treated (ATT) as the causal estimand. We establish the identification of the geodesic ATT and derive the convergence rate of its sample versions, employing tools from metric geometry and empirical process theory. This framework is further extended to the case of staggered DID settings, allowing for multiple time periods and varying treatment timings. To illustrate the practical utility of geodesic DID, we analyze health impacts of the Soviet Union's collapse using age-at-death distributions and assess effects of U.S. electricity market liberalization on electricity generation compositions.

stat.ME

Geodesic Causal Inference

Adjusting for confounding and imbalance when establishing statistical relationships is an increasingly important task, and causal inference methods have emerged as the most popular tool to achieve this. Existing methodology has been developed primarily for outcomes that lie in Euclidean spaces. We introduce here a general framework for causal inference when outcomes reside in general geodesic metric spaces, where we draw on a novel geodesic calculus that facilitates scalar multiplication for geodesics and the quantification of treatment effects through the concept of geodesic average treatment effect. Using ideas from Fr\'echet regression, we obtain a doubly robust estimation of the geodesic average treatment effect and results on consistency and rates of convergence for the proposed estimators. We also develop an intrinsic uncertainty quantification framework for the treatment effect based on Fr\'echet objective functions. The proposed framework is illustrated through simulations and real data applications, including network-valued outcomes from New York City taxi trips to assess the impact of the COVID-19 pandemic, and compositional data on U.S. state-level energy sources to study the effect of coal mining.

stat.ME

Adaptive deep learning for nonlinear time series models

In this paper, we develop a general theory for adaptive nonparametric estimation of the mean function of a non-stationary and nonlinear time series model using deep neural networks (DNNs). We first consider two types of DNN estimators, non-penalized and sparse-penalized DNN estimators, and establish their generalization error bounds for general non-stationary time series. We then derive minimax lower bounds for estimating mean functions belonging to a wide class of nonlinear autoregressive (AR) models that include nonlinear generalized additive AR, single index, and threshold AR models. Building upon the results, we show that the sparse-penalized DNN estimator is adaptive and attains the minimax optimal rates up to a poly-logarithmic factor for many nonlinear AR models. Through numerical simulations, we demonstrate the usefulness of the DNN methods for estimating nonlinear AR models with intrinsic low-dimensional structures and discontinuous or rough mean functions, which is consistent with our theory.

math.ST

Local-Polynomial Estimation for Multivariate Regression Discontinuity Designs

We study a multivariate regression discontinuity design in which treatment is assigned by crossing a boundary in the space of multiple running variables. We document that the existing bandwidth selector is suboptimal for a multivariate regression discontinuity design when the distance to a boundary point is used for its running variable, and introduce a multivariate local-linear estimator for multivariate regression discontinuity designs. Our estimator is asymptotically valid and can capture heterogeneous treatment effects over the boundary. We demonstrate that our estimator exhibits smaller root mean squared errors and often shorter confidence intervals in numerical simulations. We illustrate our estimator in our empirical applications of multivariate designs of a Colombian scholarship study and a U.S. House of representative voting study and demonstrate that our estimator reveals richer heterogeneous treatment effects with often shorter confidence intervals than the existing estimator.

econ.EM

Series ridge regression for spatial data on $\mathbb{R}^d$

This paper develops a general asymptotic theory of series estimators for spatial data collected at irregularly spaced locations within a sampling region $R_n \subset \mathbb{R}^d$. We employ a stochastic sampling design that can flexibly generate irregularly spaced sampling sites, encompassing both pure increasing and mixed increasing domain frameworks. Specifically, we focus on a spatial trend regression model and a nonparametric regression model with spatially dependent covariates. For these models, we investigate $L^2$-penalized series estimation of the trend and regression functions. We establish uniform and $L^2$ convergence rates and multivariate central limit theorems for general series estimators as main results. Additionally, we show that spline and wavelet series estimators achieve optimal uniform and $L^2$ convergence rates and propose methods for constructing confidence intervals for these estimators. Finally, we demonstrate that our dependence structure conditions on the underlying spatial processes cover a broad class of random fields, including L\'evy-driven continuous autoregressive and moving average random fields.

math.ST

Local polynomial trend regression for spatial data on $\mathbb{R}^d$

This paper develops a general asymptotic theory of local polynomial (LP) regression for spatial data observed at irregularly spaced locations in a sampling region $R_n \subset \mathbb{R}^d$. We adopt a stochastic sampling design that can generate irregularly spaced sampling sites in a flexible manner including both pure increasing and mixed increasing domain frameworks. We first introduce a nonparametric regression model for spatial data defined on $\mathbb{R}^d$ and then establish the asymptotic normality of LP estimators with general order $p \geq 1$. We also propose methods for constructing confidence intervals and establishing uniform convergence rates of LP estimators. Our dependence structure conditions on the underlying processes cover a wide class of random fields such as Lévy-driven continuous autoregressive moving average random fields. As an application of our main results, we discuss a two-sample testing problem for mean functions and their partial derivatives.

math.ST

Hierarchical Regression Discontinuity Design: Pursuing Subgroup Treatment Effects

Regression discontinuity design (RDD) is widely adopted for causal inference under intervention determined by a continuous variable. While one is interested in treatment effect heterogeneity by subgroups in many applications, RDD typically suffers from small subgroup-wise sample sizes, which makes the estimation results highly instable. To solve this issue, we introduce hierarchical RDD (HRDD), a hierarchical Bayes approach for pursuing treatment effect heterogeneity in RDD. A key feature of HRDD is to employ a pseudo-model based on a loss function to estimate subgroup-level parameters of treatment effects under RDD, and assign a hierarchical prior distribution to ''borrow strength'' from other subgroups. The posterior computation can be easily done by a simple Gibbs sampling, and the optimal bandwidth can be automatically selected by the Hyv\"{a}rinen scores for unnormalized models. We demonstrate the proposed HRDD through simulation and real data analysis, and show that HRDD provides much more stable point and interval estimation than separately applying the standard RDD method to each subgroup.

stat.ME

On the estimation of locally stationary functional time series

This study develops an asymptotic theory for estimating the time-varying characteristics of locally stationary functional time series (LSFTS). We investigate a kernel-based method to estimate the time-varying covariance operator and the time-varying mean function of an LSFTS. In particular, we derive the convergence rate of the kernel estimator of the covariance operator and associated eigenvalue and eigenfunctions and establish a central limit theorem for the kernel-based locally weighted sample mean. As applications of our results, we discuss methods for testing the equality of time-varying mean functions in two functional samples.

math.ST

Shrinkage Methods for Treatment Choice

This study examines the problem of determining whether to treat individuals based on observed covariates. The most common decision rule is the conditional empirical success (CES) rule proposed by Manski (2004), which assigns individuals to treatments that yield the best experimental outcomes conditional on the observed covariates. Conversely, using shrinkage estimators, which shrink unbiased but noisy preliminary estimates toward the average of these estimates, is a common approach in statistical estimation problems because it is well-known that shrinkage estimators may have smaller mean squared errors than unshrunk estimators. Inspired by this idea, we propose a computationally tractable shrinkage rule that selects the shrinkage factor by minimizing an upper bound of the maximum regret. Then, we compare the maximum regret of the proposed shrinkage rule with those of the CES and pooling rules when the space of conditional average treatment effects (CATEs) is correctly specified or misspecified. Our theoretical results demonstrate that the shrinkage rule performs well in many cases and these findings are further supported by numerical experiments. Specifically, we show that the maximum regret of the shrinkage rule can be strictly smaller than those of the CES and pooling rules in certain cases when the space of CATEs is correctly specified. In addition, we find that the shrinkage rule is robust against misspecification of the space of CATEs. Finally, we apply our method to experimental data from the National Job Training Partnership Act Study.

econ.EM

Nonparametric regression for locally stationary random fields under stochastic sampling design

In this study, we develop an asymptotic theory of nonparametric regression for locally stationary random fields (LSRFs) $\{{\bf X}_{{\bf s}, A_{n}}: {\bf s} \in R_{n} \}$ in $\mathbb{R}^{p}$ observed at irregularly spaced locations in $R_{n} =[0,A_{n}]^{d} \subset \mathbb{R}^{d}$. We first derive the uniform convergence rate of general kernel estimators, followed by the asymptotic normality of an estimator for the mean function of the model. Moreover, we consider additive models to avoid the curse of dimensionality arising from the dependence of the convergence rate of estimators on the number of covariates. Subsequently, we derive the uniform convergence rate and joint asymptotic normality of the estimators for additive functions. We also introduce approximately $m_{n}$-dependent RFs to provide examples of LSRFs. We find that these RFs include a wide class of Lévy-driven moving average RFs.

math.ST

Nonparametric regression for locally stationary functional time series

In this study, we develop an asymptotic theory of nonparametric regression for a locally stationary functional time series. First, we introduce the notion of a locally stationary functional time series (LSFTS) that takes values in a semi-metric space. Then, we propose a nonparametric model for LSFTS with a regression function that changes smoothly over time. We establish the uniform convergence rates of a class of kernel estimators, the Nadaraya-Watson (NW) estimator of the regression function, and a central limit theorem of the NW estimator.

math.ST