SearcharxivSearch

arXiv subjects

Grigory Franguridi

Publications and source records attributed to Grigory Franguridi.

12 recordsLinked to original sources

Testing selection on observables in parametric models with refreshment samples

In panels with sample selection (that may occur due to attrition, nonresponse, etc.), the assumption of selection on observables (missing at random, MAR) is commonly imposed despite often being implausible. However, this assumption becomes testable when a refreshment sample is available. We develop a statistical test of MAR based on a distance between two estimated distributions: one obtained using the standard inverse probability weighting (IPW) that is valid under MAR and the other obtained using an alternative weighting that is valid under a weaker assumption of additive nonignorability of Hirano et al. (2001). This test implicitly compares the distribution of the IPW-weighted sample in the attrition period with the distribution of the refreshment sample, which coincide if the MAR assumption holds. We establish that, when the input distributions are parametric, our test statistic converges to the generalized chi-squared distribution under the null of MAR. This limit distribution can be estimated using the recursive formulas derived by Franguridi et al. (2026). We illustrate the performance of our test in Monte Carlo simulations. Finally, we apply our test to an empirical example using a subsample of the Understanding America Study (UAS) dataset.

econ.EM

Inference in partially identified moment models via regularized optimal transport

Many statistical and econometric problems involve parameters defined by moments of a joint distribution when only marginal distributions are observed, leading naturally to partial identification. We develop a methodology for identification, estimation, and inference in the corresponding partially identified GMM model. We characterize the sharp identified set for the parameter of interest via a support-function/optimal-transport (OT) representation. To estimate the identified set, we employ entropic regularization, which yields a smooth approximation to the classical OT problem that can be computed efficiently using the Sinkhorn algorithm. We also propose a test statistic for hypothesis testing and the construction of confidence regions for the identified set. To derive its asymptotic distribution, we establish a novel central limit theorem for the entropic OT value under general smooth cost functions. We then obtain valid critical values using the bootstrap for directionally differentiable functionals of Fang and Santos (2019). The resulting testing procedure controls size locally uniformly, including at parameter values on the boundary of the identified set. We demonstrate good finite-sample performance of our methodology in Monte Carlo simulations. Finally, as an empirical illustration, we estimate a panel logit model of self-reported happiness with attrition and refreshment, using data from the Understanding America Study.

econ.EM

Raking for estimation and inference in panel models with nonignorable attrition and refreshment

In panel data subject to nonignorable attrition, auxiliary (refreshment) sampling may restore full identification under weak assumptions on the attrition process. Despite their generality, these identification strategies have seen limited empirical use, largely because the implied estimation procedure requires solving a functional minimization problem for the target density. We show that this problem can be solved using the iterative proportional fitting (raking) algorithm, which converges rapidly even with continuous and moderately high-dimensional data. This resulting density estimator is then used as input into a parametric moment condition. We establish consistency and convergence rates for both the raking-based density estimator and the resulting moment estimator when the distributions of the observed data are parametric. We also derive a simple recursive procedure for estimating the asymptotic variance. Finally, we demonstrate the satisfactory performance of our estimator in simulations and provide an empirical illustration using data from the Understanding America Study panel.

econ.EM

Generalized method of moments with partially missing data

We consider a generalized method of moments framework in which a part of the data vector is missing for some units in a completely unrestricted, potentially endogenous way. In this setup, the parameters of interest are usually only partially identified. We characterize the identified set for such parameters using the support function of the convex set of moment predictions consistent with the data. This identified set is sharp, valid for both continuous and discrete data, and straightforward to estimate. We also propose a statistic for testing hypotheses and constructing confidence regions for the true parameter, show that standard nonparametric bootstrap may not be valid, and suggest a fix using the bootstrap for directionally differentiable functionals of Fang and Santos (2019). A set of Monte Carlo simulations demonstrates that both our estimator and the confidence region perform well when samples are moderately large and the data have bounded supports.

econ.EM

Debiasing Functions of Private Statistics in Postprocessing

Given a differentially private unbiased estimate $\tilde{q}=q(D) +\nu$ of a statistic $q(D)$, we wish to obtain unbiased estimates of functions of $q(D)$, such as $1/q(D)$, solely through post-processing of $\tilde{q}$, with no further access to the confidential dataset $D$. To this end, we adapt the deconvolution method used for unbiased estimation in the statistical literature, deriving unbiased estimators for a broad family of twice-differentiable functions when the privacy-preserving noise $\nu$ is drawn from the Laplace distribution (Dwork et al., 2006). We further extend this technique to a more general class of functions, deriving approximately optimal estimators that are unbiased for values in a user-specified interval (possibly extending to $\pm \infty$). We use these results to derive an unbiased estimator for private means when the size $n$ of the dataset is not publicly known. In a numerical application, we find that a mechanism that uses our estimator to return an unbiased sample size and mean outperforms a mechanism that instead uses the previously known unbiased privacy mechanism for such means (Kamath et al., 2023). We also apply our estimators to develop unbiased transformation mechanisms for per-record differential privacy, a privacy concept in which the privacy guarantee is a public function of a record's value (Seeman et al., 2024). Our mechanisms provide stronger privacy guarantees than those in prior work (Finley et al., 2024) by using Laplace, rather than Gaussian, noise. Finally, using a different approach, we go beyond Laplace noise by deriving unbiased estimators for polynomials under the weak condition that the noise distribution has sufficiently many moments.

cs.CR

Kotlarski's lemma for dyadic models

We show how to identify the distributions of the latent components in the two-way dyadic model for bipartite networks $y_{i,\ell}= \alpha_i+\eta_{\ell}+\varepsilon_{i,\ell}$. This is achieved by a repeated application of the extension of the classical lemma of Kotlarski (1967) in Evdokimov and White (2012). We provide two separate sets of assumptions under which all the latent distributions are identified. Both rely on some of the latent components being identically distributed.

econ.EM

Nonparametric inference on counterfactuals in first-price auctions

In a classical model of the first-price sealed-bid auction with independent private values, we develop nonparametric estimators for several policy-relevant targets, such as the bidder's surplus and auctioneer's revenue under counterfactual reserve prices. Motivated by the linearity of these targets in the quantile function of bidders' values, we propose an estimator of the latter and derive its Bahadur-Kiefer expansion. This makes it possible to construct uniform confidence bands and test complex hypotheses about the auction design. Using the data on U.S. Forest Service timber auctions, we test whether setting zero reserve prices in these auctions was revenue maximizing.

econ.EM

Closed-form estimation and inference for panels with attrition and refreshment samples

It has long been established that, if a panel dataset suffers from attrition, auxiliary (refreshment) sampling restores full identification under additional assumptions that still allow for nontrivial attrition mechanisms. Such identification results rely on implausible assumptions about the attrition process or lead to theoretically and computationally challenging estimation procedures. We propose an alternative identifying assumption that, despite its nonparametric nature, suggests a simple estimation algorithm based on a transformation of the empirical cumulative distribution function of the data. This estimation procedure requires neither tuning parameters nor optimization in the first step, i.e., it has a closed form. We prove that our estimator is consistent and asymptotically normal and demonstrate its good performance in simulations. We provide an empirical illustration with income data from the Understanding America Study.

econ.EM

Bias correction and uniform inference for the quantile density function

For the kernel estimator of the quantile density function (the derivative of the quantile function), I show how to perform the boundary bias correction, establish the rate of strong uniform consistency of the bias-corrected estimator, and construct the confidence bands that are asymptotically exact uniformly over the entire domain $[0,1]$. The proposed procedures rely on the pivotality of the studentized bias-corrected estimator and known anti-concentration properties of the Gaussian approximation for its supremum.

econ.EM

Efficient counterfactual estimation in semiparametric discrete choice models: a note on Chiong, Hsieh, and Shum (2017)

I suggest an enhancement of the procedure of Chiong, Hsieh, and Shum (2017) for calculating bounds on counterfactual demand in semiparametric discrete choice models. Their algorithm relies on a system of inequalities indexed by cycles of a large number $M$ of observed markets and hence seems to require computationally infeasible enumeration of all such cycles. I show that such enumeration is unnecessary because solving the "fully efficient" inequality system exploiting cycles of all possible lengths $K=1,\dots,M$ can be reduced to finding the length of the shortest path between every pair of vertices in a complete bidirected weighted graph on $M$ vertices. The latter problem can be solved using the Floyd--Warshall algorithm with computational complexity $O\left(M^3\right)$, which takes only seconds to run even for thousands of markets. Monte Carlo simulations illustrate the efficiency gain from using cycles of all lengths, which turns out to be positive, but small.

econ.EM

Bias correction for quantile regression estimators

We study the bias of classical quantile regression and instrumental variable quantile regression estimators. While being asymptotically first-order unbiased, these estimators can have non-negligible second-order biases. We derive a higher-order stochastic expansion of these estimators using empirical process theory. Based on this expansion, we derive an explicit formula for the second-order bias and propose a feasible bias correction procedure that uses finite-difference estimators of the bias components. The proposed bias correction method performs well in simulations. We provide an empirical illustration using Engel's classical data on household food expenditure.

econ.EM

A Uniform Bound on the Operator Norm of Sub-Gaussian Random Matrices and Its Applications

For an $N \times T$ random matrix $X(\beta)$ with weakly dependent uniformly sub-Gaussian entries $x_{it}(\beta)$ that may depend on a possibly infinite-dimensional parameter $\beta\in \mathbf{B}$, we obtain a uniform bound on its operator norm of the form $\mathbb{E} \sup_{\beta \in \mathbf{B}} ||X(\beta)|| \leq CK \left(\sqrt{\max(N,T)} + \gamma_2(\mathbf{B},d_\mathbf{B})\right)$, where $C$ is an absolute constant, $K$ controls the tail behavior of (the increments of) $x_{it}(\cdot)$, and $\gamma_2(\mathbf{B},d_\mathbf{B})$ is Talagrand's functional, a measure of multi-scale complexity of the metric space $(\mathbf{B},d_\mathbf{B})$. We illustrate how this result may be used for estimation that seeks to minimize the operator norm of moment conditions as well as for estimation of the maximal number of factors with functional data.

econ.EM