SearcharxivSearch

arXiv subjects

Matias D. Cattaneo

Publications and source records attributed to Matias D. Cattaneo.

At least 19 recordsLinked to original sources

Beta-Sorted Portfolios

Beta-sorted portfolios---portfolios comprised of assets with similar covariation with selected risk factors---are a popular tool in empirical finance to analyze models of (conditional) expected returns. Despite their widespread use, little is known of their econometric properties in contrast to comparable procedures such as two-pass regressions. We formally investigate the properties of beta-sorted portfolio returns by casting the procedure as a two-step nonparametric estimator with a nonparametric first step and a beta-adaptive portfolio construction. Our framework rationalizes the well-known estimation algorithm with precise economic and statistical assumptions on the general data-generating process. We provide conditions which ensure valid estimation and inference allowing for a range of hypotheses of interest in financial applications. We show that the rate of convergence of the estimator changes depending on the value of beta. We demonstrate that valid inference depends critically on the object of interest and discuss drawbacks of the widely used Fama-MacBeth variance estimator. To address these limitations, we propose a new variance estimator. We demonstrate the usefulness of our theoretical results in two empirical applications, including one in which we introduce a novel risk factor that captures the business credit cycle and show that it predicts both the cross-sectional and time-series behavior of U.S. stock returns.

econ.EM

Attention Overload

We introduce an Attention Overload Model (AOM) in which alternatives compete for attention, so each alternative's consideration probability weakly decreases as the choice problem expands. This nonparametric restriction has a search-capacity foundation. We identify exactly which preference rankings are compatible with observed choices and establish sharp bounds on latent attention. We then consider heterogeneous preferences in settings where alternatives are presented in a list, establishing nonparametric identification results for attention and the distribution of preferences. We also develop high-dimensional inference and finite-sample methods for these models. Using the travel-mode experiment of Wang and Zhu (2025), we illustrate the methods and find that recovered pairwise preference shares closely match subjects' self-reported rankings.

econ.TH

Nonlinear Binscatter Methods

Binscatters are a powerful tool for empirical work in the social, behavioral, and biomedical sciences. Available tools rely on least squares estimation of the conditional mean. We introduce novel binscatter methods based on nonlinear, possibly nonsmooth M-estimation, covering generalized linear, robust, and quantile regression models. We provide theoretical results and practical tools, including optimal bin selection, confidence bands, and statistical tests regarding functional form or shape restrictions. We demonstrate our methods by studying the relationship of income and (lack of) health insurance. We provide software for Python, R, and Stata. Our technical results may be of independent interest.

stat.ME

Uniform Estimation and Inference for Nonparametric Partitioning-Based M-Estimators

This paper presents uniform estimation and inference theory for a large class of nonparametric partitioning-based M-estimators. The main theoretical results include: (i) uniform consistency for convex and non-convex objective functions; (ii) rate-optimal uniform Bahadur representations; (iii) rate-optimal uniform (and mean square) convergence rates; (iv) valid strong approximations and feasible uniform inference methods; and (v) extensions to functional transformations of underlying estimators. Uniformity is established over both the evaluation point of the nonparametric functional parameter and a Euclidean parameter indexing the class of loss functions. The results also account explicitly for the smoothness degree of the loss function (if any), and allow for a possibly non-identity (inverse) link function. We illustrate the theoretical and methodological results in four examples: quantile regression, distribution regression, $L_p$ regression, and logistic regression. Many other possibly non-smooth, nonlinear, generalized, robust M-estimation settings are covered by our results. We provide detailed comparisons with the existing literature and demonstrate substantive improvements: we achieve the best (in some cases optimal) known results under improved (in some cases minimal) requirements in terms of regularity conditions and side rate restrictions. The supplemental appendix reports complementary technical results that may be of independent interest, including a novel uniform strong approximation result based on Yurinskii's coupling.

math.ST

Accuracy Limits of Causal Trees for Individualized Treatment Effects

Recursive decision trees are widely used to estimate heterogeneous causal treatment effects in experimental and observational studies. These methods are typically implemented using CART-type recursive partitioning, with splitting criteria designed to identify variation in treatment effects across covariate-defined subgroups. We study causal tree estimators based on adaptive recursive partitioning and establish lower bounds on their estimation accuracy. The class we analyze includes versions with and without sample splitting, based on common treatment effect and squared-error splitting criteria. Even in a constant-effect benchmark with randomized treatment assignment, causal trees constructed via standard CART-type splitting rules can have uniform-norm errors that decrease more slowly than any power of the sample size. The underlying mechanism is that greedy recursive partitioning selects highly imbalanced splits with nonvanishing probability, producing terminal nodes containing very few observations and leading to large estimation variance. We further show that sample splitting, often called ``honesty,'' does not remove this limitation. As a consequence, causal tree estimators may converge arbitrarily slowly uniformly over the covariate space. At the same time, these estimators can have small integrated mean squared error, showing that average accuracy can mask local inaccuracy. Our results also clarify the role of balanced partition assumptions in existing theoretical guarantees for causal forests and related ensemble methods.

math.ST

rd2d: Causal Inference in Boundary Discontinuity Designs

Boundary Discontinuity (BD) designs are used in empirical research to learn about causal treatment effects along a continuous assignment boundary defined by a bivariate score. These designs are also known as multi-score regression discontinuity (RD) designs, and include geographic RD designs as a prominent example. This article introduces \pkg{rd2d}, a statistical software package for \proglang{R}, \proglang{Python}, and \proglang{Stata} that implements local polynomial estimation and inference for BD designs using either the bivariate score or a univariate signed distance-to-boundary score. The software covers sharp and fuzzy BD designs, providing automatic bandwidth selection, robust bias-corrected pointwise inference, uniform confidence bands, cluster-robust inference with joint or separate fitting conventions, covariate-adjusted efficiency improvements, mass-point checks, and covariance regularization, among other features. We illustrate the package with an empirical application to Opportunity Zones, where eligibility has a strong first-stage effect on designation but no significant effects on early workplace-job growth.

stat.ME

Robust Inference for Convex Pairwise Difference Estimators

This paper develops distribution theory and bootstrap-based inference methods for a broad class of convex pairwise difference estimators. These estimators minimize a kernel-weighted convex-in-parameter function over observation pairs with similar covariates, where the similarity is governed by a localization (bandwidth) parameter. While classical results establish asymptotic normality under restrictive bandwidth conditions, we show that valid Gaussian and bootstrap-based inference remains possible under substantially weaker assumptions. First, we extend the theory of small bandwidth asymptotics to convex pairwise difference estimation settings, deriving robust Gaussian approximations even when a smaller than standard bandwidth is used. Second, we employ a debiasing procedure based on generalized jackknifing to enable inference with larger bandwidths, while preserving convexity of the objective function. Third, we construct a novel bootstrap method that adjusts for bandwidth-induced variance distortions, yielding valid inference across a wide range of bandwidth choices. Our proposed inference method enjoys demonstrably greater robustness, while retaining the practical appeal of convex pairwise difference estimators.

econ.EM

Estimation and Inference in Boundary Discontinuity Designs: Distance-Based Methods

We study nonparametric distance-based (isotropic) local polynomial methods for estimating the boundary average treatment effect curve, a causal functional that captures treatment effect heterogeneity in boundary discontinuity designs. We establish identification, estimation, and inference results both pointwise and uniformly along the treatment assignment boundary. We show that the geometric regularity of the boundary, a one-dimensional manifold, plays a central role in determining feasible convergence rates and valid inference procedures. Our theoretical contributions are threefold. First, we derive uniform lower and upper bounds on the convergence rate of the misspecification bias of isotropic local polynomial estimators. Second, we obtain uniform distributional approximations that justify boundary-robust inference. Third, we establish minimax lower bounds for a broad class of nonparametric isotropic regression estimators. These results yield practical guidance for empirical implementation, including new bandwidth selection rules that adapt to local irregularities of the treatment-assignment boundary. We illustrate the proposed methods using simulation evidence and an empirical application, and provide companion general-purpose software.

econ.EM

Estimation and Inference in Boundary Discontinuity Designs: Location-Based Methods

Boundary discontinuity designs are used to learn about causal treatment effects along a continuous assignment boundary that splits units into control and treatment groups according to a bivariate location score. We analyze location-based local polynomial treatment effect estimators that directly employ the bivariate score of each unit. We develop pointwise and uniform estimation and inference methods for the \textit{Boundary Average Treatment Effect Curve} (BATEC), as well as for two aggregated causal parameters: the \textit{Weighted Boundary Average Treatment Effect} (WBATE) and the \textit{Largest Boundary Average Treatment Effect} (LBATE). Our results cover both sharp and fuzzy (imperfect compliance) designs. We illustrate the methods with an empirical application, and provide companion general-purpose software. The supplemental appendix includes additional substantive theoretical results, methodological details, and simulation evidence.

econ.EM

The Effect of Mini-Batch Noise on the Implicit Bias of Adam

With limited high-quality data and growing compute, multi-epoch training is gaining back its importance across sub-areas of deep learning. Adam(W), versions of which are go-to optimizers for many tasks such as next token prediction, has two momentum hyperparameters $(β_1, β_2)$ controlling memory and one very important hyperparameter, batch size, controlling (in particular) the amount mini-batch noise. We introduce a theoretical framework to understand how mini-batch noise influences the implicit bias of memory in Adam (depending on $β_1$, $β_2$) towards sharper or flatter regions of the loss landscape, which is commonly observed to correlate with the generalization gap in multi-epoch training. We find that in the case of large batch sizes, higher $β_2$ increases the magnitude of anti-regularization by memory (hurting generalization), but as the batch size becomes smaller, the dependence of (anti-)regulariation on $β_2$ is reversed. A similar monotonicity shift (in the opposite direction) happens in $β_1$. In particular, the commonly "default" pair $(β_1, β_2) = (0.9, 0.999)$ is a good choice if batches are small; for larger batches, in many settings moving $β_1$ closer to $β_2$ is much better in terms of validation accuracy in multi-epoch training. Moreover, our theoretical derivations connect the scale of the batch size at which the shift happens to the scale of the critical batch size. We illustrate this effect in experiments with small-scale data in the about-to-overfit regime.

cs.LG

Boundary Discontinuity Designs: Theory and Practice

The boundary discontinuity (BD) design is a non-experimental method for identifying causal effects that exploits a thresholding rule based on a bivariate score and a boundary curve. This widely used method generalizes the univariate regression discontinuity design but introduces unique challenges arising from its multidimensional nature. We synthesize over 80 empirical papers that use the BD design, tracing the method's application from its formative stages to its implementation in modern research. We also overview ongoing theoretical and methodological research on identification, estimation, and inference for BD designs employing local polynomial regression, and offer recommendations for practice.

econ.EM

How Memory in Optimization Algorithms Implicitly Modifies the Loss

In modern optimization methods used in deep learning, each update depends on the history of previous iterations, often referred to as memory, and this dependence decays fast as the iterates go further into the past. For example, gradient descent with momentum has exponentially decaying memory through exponentially averaged past gradients. We introduce a general technique for identifying a memoryless algorithm that approximates an optimization algorithm with memory. It is obtained by replacing all past iterates in the update by the current one, and then adding a correction term arising from memory (also a function of the current iterate). This correction term can be interpreted as a perturbation of the loss, and the nature of this perturbation can inform how memory implicitly (anti-)regularizes the optimization dynamics. As an application of our theory, we find that Lion does not have the kind of implicit anti-regularization induced by memory that AdamW does, providing a theory-based explanation for Lion's better generalization performance recently documented.

cs.LG

Inference with Mondrian Random Forests

Random forests are popular methods for regression and classification analysis, and many different variants have been proposed in recent years. One interesting example is the Mondrian random forest, in which the underlying constituent trees are constructed via a Mondrian process. We give precise bias and variance characterizations, along with a Berry-Esseen-type central limit theorem, for the Mondrian random forest regression estimator. By combining these results with a carefully crafted debiasing approach and an accurate variance estimator, we present valid statistical inference methods for the unknown regression function. These methods come with explicitly characterized error bounds in terms of the sample size, tree complexity parameter, and number of trees in the forest, and include coverage error rates for feasible confidence interval estimators. Our debiasing procedure for the Mondrian random forest also allows it to achieve the minimax-optimal point estimation convergence rate in mean squared error for multivariate $β$-Hölder regression functions, for all $β> 0$, provided that the underlying tuning parameters are chosen appropriately. Efficient and implementable algorithms are devised for both batch and online learning settings, and we study the computational complexity of different Mondrian random forest implementations. Finally, simulations with synthetic data validate our theory and methodology, demonstrating their excellent finite-sample properties.

math.ST

Continuity of the Distribution Function of the argmax of a Gaussian Process

Certain extremum estimators have asymptotic distributions that are non-Gaussian, yet characterizable as the distribution of the $\argmax$ of a Gaussian process. This paper presents high-level sufficient conditions under which such asymptotic distributions admit a continuous distribution function. The plausibility of the sufficient conditions is demonstrated by verifying them in three examples, namely maximum score estimation, empirical risk minimization, and threshold regression estimation. In turn, the continuity result buttresses several recently proposed inference procedures whose validity seems to require a result of the kind established herein. A notable feature of the high-level assumptions is that one of them is designed to enable us to employ the Cameron-Martin theorem. In a leading special case, the assumption in question is demonstrably weak and appears to be close to minimal.

econ.EM

Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis

We analyze gradient descent with Polyak heavy-ball momentum (HB) whose fixed momentum parameter $β\in (0, 1)$ provides exponential decay of memory. Building on Kovachki and Stuart (2021), we prove that on an exponentially attractive invariant manifold the algorithm is exactly plain gradient descent with a modified loss, provided that the step size $h$ is small enough. Although the modified loss does not admit a closed-form expression, we describe it with arbitrary precision and prove global (finite "time" horizon) approximation bounds $O(h^{R})$ for any finite order $R \geq 2$. We then conduct a fine-grained analysis of the combinatorics underlying the memoryless approximations of HB, in particular, finding a rich family of polynomials in $β$ hidden inside which contains Eulerian and Narayana polynomials. We derive continuous modified equations of arbitrary approximation order (with rigorous bounds) and the principal flow that approximates the HB dynamics, generalizing Rosca et al. (2023). Approximation theorems cover both full-batch and mini-batch HB. Our theoretical results shed new light on the main features of gradient descent with heavy-ball momentum, and outline a road-map for similar analysis of other optimization algorithms.

cs.LG

Randomization Inference for Before-and-After Studies with Multiple Units: An Application to a Criminal Procedure Reform in Uruguay

Learning about the immediate causal effects of large-scale policy interventions poses a significant challenge for quasi-experimental methods that rely on long-term trends or parametric modeling assumptions. As an alternative, we develop a randomization inference framework for before-and-after studies with multiple units, designed specifically for short-term causal inference and allowing for general assignment mechanisms. The method provides finite-sample-valid statistical inferences without relying on parametric time series models or extrapolation. We demonstrate its utility by analyzing a major criminal justice reform in Uruguay that switched from an inquisitorial to an adversarial system in November 2017. Our method relies on the key assumption of no local time trends near the policy adoption time, which is supported by several falsification tests in our empirical study. We find a statistically significant short-term causal effect: an increase of approximately 25 daily police reports (an 8% rise) in the first week of the new justice system. Our randomization inference framework provides a robust and flexible methodology for evaluating policy adoptions in before-and-after studies with multiple units.

stat.ME

The Regression Discontinuity Design in Medical Science

This article provides an introduction to the Regression Discontinuity (RD) design, and its application to empirical research in the medical sciences. While the main focus of this article is on causal interpretation, key concepts of estimation and inference are also briefly mentioned. A running medical empirical example is provided.

stat.ME

Yurinskii's Coupling for Martingales

Yurinskii's coupling is a popular theoretical tool for non-asymptotic distributional analysis in mathematical statistics and applied probability, offering a Gaussian strong approximation with an explicit error bound under easily verifiable conditions. Originally stated in $\ell_2$-norm for sums of independent random vectors, it has recently been extended both to the $\ell_p$-norm, for $1 \leq p \leq \infty$, and to vector-valued martingales in $\ell_2$-norm, under some strong conditions. We present as our main result a Yurinskii coupling for approximate martingales in $\ell_p$-norm, under substantially weaker conditions than those previously imposed. Our formulation further allows for the coupling variable to follow a more general Gaussian mixture distribution, and we provide a novel third-order coupling method which gives tighter approximations in certain settings. We specialize our main result to mixingales, martingales, and independent data, and derive uniform Gaussian mixture strong approximations for martingale empirical processes. Applications to nonparametric partitioning-based and local polynomial regression procedures are provided, alongside central limit theorems for high-dimensional martingale vectors.

math.ST