SearcharxivSearch

arXiv subjects

Amilcar Velez

Publications and source records attributed to Amilcar Velez.

9 recordsLinked to original sources

Debiased Machine Learning with Many Cross-Fitting Folds

This paper studies debiased machine learning (DML) when the number of cross-fitting folds, $K_n$, may grow with the sample size $n$. Existing fixed-$K$ asymptotic theory implies that DML1 and DML2, the two main DML variants, are asymptotically equivalent, providing no guidance on which variant to use or how to choose $K_n$. We show that this equivalence can break down when $K_n$ grows proportionally to $\sqrt{n}$: DML1 can exhibit asymptotic bias, in which case standard inference based on DML1 fails---as can occur, for instance, for the local average treatment effect (LATE)---whereas inference based on DML2 remains valid. Moreover, we show that, under an algorithmic-stability condition, estimation and inference based on DML2 are valid for any $2\le K_n \le n$, including the leave-one-out case, $K_n=n$. Finally, for scalar DML2 estimators whose first-step estimators admit a stochastic linear expansion, we derive a second-order approximation showing that larger values of $K_n$ reduce the second-order asymptotic bias and mean-squared error, although the marginal improvements diminish.

econ.EM

A Simple Approximation to the Distribution of the Ridge Regression Estimator

We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.

econ.EM

On the power properties of inference for parameters with interval identified sets

This paper studies the power properties of confidence intervals (CIs) for a partially-identified parameter of interest with an interval identified set. We assume the researcher has bounds estimators needed to construct the CIs proposed by Imbens and Manski (2004), Stoye (2009), and Stoye (2020), denoted by CI_alpha^1, CI_alpha^2, CI_alpha^3, and CI_alpha^4. We also assume these bounds estimators are ``ordered'': the lower bound estimator is less than or equal to the upper bound estimator. This setup arises in economic applications involving missing data and treatment effects. Under these conditions, we establish two results. First, we show that CI_alpha^1 and CI_alpha^2 are equally powerful, and both dominate CI_alpha^3 and CI_alpha^4. Second, we consider a favorable situation in which there are two possible bounds estimators to construct these CIs, and one is more efficient than the other. One would expect that the more efficient bounds estimator yields more powerful inference. We prove that this desirable result holds for CI_alpha^1 and CI_alpha^2, but not necessarily for CI_alpha^3 or CI_alpha^4. In summary, within the class of models considered, CI_alpha^1 and CI_alpha^2 have identical power properties, and both compare favorably to CI_alpha^3 or CI_alpha^4.

econ.EM

Identification and Inference for Algorithmic Frontiers with Selective Labels

This paper provides identification results to characterize a fairness-accuracy (FA) frontier, and statistical inference tools to test hypotheses and build a confidence set for the FA-frontier, when outcomes are observed only for selected individuals. When the selection process is unrestricted but loss is measured in specific ways, we provide a characterization of the sharp identification region of the FA-frontier. Under an assumption of unconfoundedness conditional on observables (and unrestricted loss functions), we obtain point identification and propose a debiased machine learning estimator, derive its asymptotic distribution, and show how this can be used to carry out inference for the FA-frontier. In work in progress, we extend the partial identification results to a broader class of loss functions.

econ.EM

Decision Theory for the Archetype Discovery Problem

In the archetype discovery problem a researcher wants to summarize N heterogeneous policy effects of interest that vary over a discrete set of covariates. The goal is to partition the set of covariates into K<N groups -- the archetype sets -- and to provide a summary of the policy effects for each group. We use decision theory to show that, under a weighted mean-squared-error criterion, a procedure analogous to the Sorted Group Average Treatment Effects (GATES) solves the archetype discovery problem. The key difference is that, in the optimal procedure, archetype sets are obtained by weighted K-means clustering of the N heterogeneous policy effects, instead of relying on K equally-spaced quantiles. We show that the procedure that minimizes average risk for a given prior can be obtained by clustering the different values of the posterior mean estimate of the policy effects of interest. Similarly, an approximately minimax procedure in large samples can be obtained by clustering a consistent estimator of the policy effects. In both of these cases, an exact solution to the weighted K-means clustering problem can be found using a simple and well-known dynamic programming algorithm.

econ.EM

The Local Projection Residual Bootstrap for AR(1) Models

This paper proposes a local projection residual bootstrap method to construct confidence intervals for impulse response coefficients of AR(1) models. Our bootstrap method is based on the local projection (LP) approach and involves a residual bootstrap procedure applied to AR(1) models. We present theoretical results for our bootstrap method and proposed confidence intervals. First, we prove the uniform consistency of the LP-residual bootstrap over a large class of AR(1) models that allow for a unit root, conditional heteroskedasticity of unknown form, and martingale difference shocks. Then, we prove the asymptotic validity of our confidence intervals over the same class of AR(1) models. Finally, we show that the LP-residual bootstrap provides asymptotic refinements for confidence intervals on a restricted class of AR(1) models relative to those required for the uniform consistency of our bootstrap.

econ.EM

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the average treatment effect (ATE) and the average treatment effect on the treated (ATT) in such RCTs with a binary treatment. We first develop characterizations of the identified sets for both estimands. Since data are generally not i.i.d. under CAR, these characterizations do not follow from existing results. We then provide consistent estimators of the identified sets and asymptotically valid confidence intervals for the parameters. Our asymptotic analysis leads to concrete practical recommendations regarding how to estimate the treatment assignment probabilities that enter the estimated bounds. For the ATE bounds, using sample analog assignment frequencies is more efficient than relying on the true assignment probabilities. For the ATT bounds, the most efficient approach is to use the true assignment probability for the probabilities in the numerator and the sample analog for those in the denominator.

econ.EM

The out-of-sample prediction error of the square-root-LASSO and related estimators

We study the classical problem of predicting an outcome variable, $Y$, using a linear combination of a $d$-dimensional covariate vector, $\mathbf{X}$. We are interested in linear predictors whose coefficients solve: % \begin{align*} \inf_{\boldsymbolβ \in \mathbb{R}^d} \left( \mathbb{E}_{\mathbb{P}_n} \left[ \left(Y-\mathbf{X}^{\top}β\right)^r \right] \right)^{1/r} +δ\, ρ\left(\boldsymbolβ\right), \end{align*} where $δ>0$ is a regularization parameter, $ρ:\mathbb{R}^d\to \mathbb{R}_+$ is a convex penalty function, $\mathbb{P}_n$ is the empirical distribution of the data, and $r\geq 1$. We present three sets of new results. First, we provide conditions under which linear predictors based on these estimators % solve a \emph{distributionally robust optimization} problem: they minimize the worst-case prediction error over distributions that are close to each other in a type of \emph{max-sliced Wasserstein metric}. Second, we provide a detailed finite-sample and asymptotic analysis of the statistical properties of the balls of distributions over which the worst-case prediction error is analyzed. Third, we use the distributionally robust optimality and our statistical analysis to present i) an oracle recommendation for the choice of regularization parameter, $δ$, that guarantees good out-of-sample prediction error; and ii) a test-statistic to rank the out-of-sample performance of two different linear estimators. None of our results rely on sparsity assumptions about the true data generating process; thus, they broaden the scope of use of the square-root lasso and related estimators in prediction problems.

math.ST

On the Robustness to Misspecification of $α$-Posteriors and Their Variational Approximations

$α$-posteriors and their variational approximations distort standard posterior inference by downweighting the likelihood and introducing variational approximation errors. We show that such distortions, if tuned appropriately, reduce the Kullback-Leibler (KL) divergence from the true, but perhaps infeasible, posterior distribution when there is potential parametric model misspecification. To make this point, we derive a Bernstein-von Mises theorem showing convergence in total variation distance of $α$-posteriors and their variational approximations to limiting Gaussian distributions. We use these distributions to evaluate the KL divergence between true and reported posteriors. We show this divergence is minimized by choosing $α$ strictly smaller than one, assuming there is a vanishingly small probability of model misspecification. The optimized value becomes smaller as the the misspecification becomes more severe. The optimized KL divergence increases logarithmically in the degree of misspecification and not linearly as with the usual posterior.

stat.ML