Searcharxiv⌕ Search

arXiv subjects

Florian Gunsilius

Publications and source records attributed to Florian Gunsilius.

14 recordsLinked to original sources

Return to Office and the Tenure Distribution

Debates over return-to-office mandates have intensified since the COVID-19 pandemic ended, though their economic implications remain poorly understood. Using large-scale resume data, we analyze how these policies affect employee tenure and seniority at three large U.S. tech companies: Microsoft, SpaceX, and Apple. Employing a distributional synthetic controls framework, we estimate a reduction in tenure and seniority at firms returning to the office. We extend the framework with a functional bootstrap procedure and prove its validity for constructing uniform confidence bands. Return-to-office mandates appear to drive senior employees away, posing a risk to firm productivity and innovation.

econ.GN↗

disco: Distributional Synthetic Controls

The method of synthetic controls is widely used for evaluating causal effects of policy changes in settings with observational data. Often, researchers aim to estimate the causal impact of policy interventions on a treated unit at an aggregate level while also possessing data at a finer granularity. In this article, we introduce the new disco command, which implements the Distributional Synthetic Controls method introduced in Gunsilius (2023, Econometrica 91: 1105-1117). This command allows researchers to construct entire synthetic distributions for the treated unit based on an optimally weighted average of the distributions of the control units. Several aggregation schemes are provided to facilitate clear reporting of the distributional effects of the treatment. The package offers both quantile-based and cumulative distribution function-based approaches, comprehensive inference procedures via bootstrap and permutation methods, and visualization capabilities. We empirically illustrate the use of the package by replicating the results in Van Dijcke, Gunsilius, and Wright (2026, Review of Economics and Statistics, forthcoming).

econ.EM↗

A condition for the identification of multivariate models with binary instruments -- with Corrigendum and Addendum

This article introduces an empirical condition for the nonparametric point-identification of multivariate instrumental variable models with continuous endogenous variables using binary instruments. Verifying this condition can confirm point-identification in settings in which traditional approaches are not applicable. In particular, it shows that nonlinear instrumental variable models with general heterogeneity can be point-identified with only a binary instrument. This generalizes existing identification results which either restrict the unobserved heterogeneity substantially or require the instrument to have a large support. The main assumption on the instrumental variable model is cyclic monotonicity of its first stage, a multivariate generalization of the classical rank-invariance assumption for univariate models. Asymptotic convergence results for the empirical observable distributions are derived that allow to check the condition in practice. The identification rests on a fixed-set convergence result of cyclically monotone maps between quasi-concave functions. The corrigendum corrects the proof of Lemma 1. The proof given there incorrectly identifies preservation of distributional level sets with preservation of the underlying probability measure via Brenier maps. We replace that argument by one based on inverse Brenier maps, which play the role of multivariate ranks. The corrected argument applies to a different but significantly more flexible class of distributions than the quasi-concave class considered in the original paper. In particular, it allows for smooth non-quasi-concave and multimodal densities on compact supports, provided the associated rank fixed set satisfies a nondegeneracy condition. Moreover, it is generically satisfied for smooth parmetric classes of distributions.

econ.EM↗

On the Differential-Geometric Equivalence of Hellinger-Kantorovich and Cone-Wasserstein Spaces

The Hellinger-Kantorovich (HK) space provides a natural geometry for nonnegative measures with varying total mass, but its differential-geometric structure is less well understood than that of the closely related Wasserstein space of probability measures. In this paper, we take a step toward resolving this issue. We show that the cone representation of the HK geometry via the Wasserstein metric preserves the local Riemannian geometry along a class of lifted geodesics. Specifically, we give a constructive procedure that produces a Wasserstein geodesic on the cone along which the HK Riemannian geometry is preserved pointwise, yielding an explicit isometry of tangent spaces between HK geodesics and their Wasserstein lifts. This connection makes many Wasserstein-geometric tools available for HK computations. Concretely, we use it to approximate parallel transport on HK space by lifting to the cone and applying recently developed Wasserstein parallel transport tools, circumventing the high-dimensional PDE arising from the HK covariant derivative. We also derive closed-form expressions for the covariant derivative and parallel transport on Euclidean metric cones, using the theory of warped-product manifolds. Finally, we present simulations illustrating the behavior of parallel geodesics in HK space, which reveal that the HK geometry couples spatial and mass variation through the geometry of the cone -- a feature with nontrivial implications for applied use of the framework.

math.MG↗

Wasserstein Parallel Transport for Predicting the Dynamics of Statistical Systems

Many scientific systems, such as cellular populations or economic cohorts, are naturally described by probability distributions that evolve over time. Predicting how such a system would have evolved under different forces or initial conditions is fundamental to causal inference, domain adaptation, and counterfactual prediction. However, the space of distributions often lacks the vector space structure on which classical methods rely. To address this, we introduce a general notion of parallel dynamics at a distributional level. We base this principle on parallel transport of tangent dynamics along optimal transport geodesics and call it ``Wasserstein Parallel Trends''. By replacing the vector subtraction of classic methods with geodesic parallel transport, we can provide counterfactual comparisons of distributional dynamics in applications such as causal inference, domain adaptation, and batch-effect correction in experimental settings. The main mathematical contribution is a novel notion of fanning scheme on the Wasserstein manifold that allows us to efficiently approximate parallel transport along geodesics while also providing the first theoretical guarantees for parallel transport in the Wasserstein space. We also show that Wasserstein Parallel Trends recovers the classic parallel trends assumption for averages as a special case and derive closed-form parallel transport for Gaussian measures. We deploy the method on synthetic data and two single-cell RNA sequencing datasets to impute gene-expression dynamics across biological systems.

stat.ML↗

Nonparametric Testability of Slutsky Symmetry

Economic theory implies strong limitations on what types of consumption behavior are considered rational. Rationality implies that the Slutsky matrix, which captures the substitution effects of compensated price changes on demand for different goods, is symmetric and negative semi-definite. While empirically informed versions of negative semi-definiteness have been shown to be nonparametrically testable, the analogous question for Slutsky symmetry has remained open. Recently, it has even been shown that the symmetry condition is not testable via the average Slutsky matrix, prompting conjectures about its non-testability. We settle this question by deriving nonparametric conditional quantile restrictions on observable data that constitute a testable implication of Slutsky symmetry in an empirical setting with individual heterogeneity and endogeneity. The theoretical contribution is a multivariate generalization of identification results for partial effects in nonseparable models without monotonicity, which is of independent interest. This result has implications for different areas in econometric theory, including nonparametric welfare analysis with individual heterogeneity for which, in the case of more than two goods, the symmetry condition introduces nonlinear correction factors.

econ.EM↗

Free Discontinuity Regression: With an Application to the Economic Effects of Internet Shutdowns

Sharp, multidimensional changepoints-abrupt shifts in a regression surface whose locations and magnitudes are unknown-arise in settings as varied as gene-expression profiling, financial covariance breaks, climate-regime detection, and urban socioeconomic mapping. Despite their prevalence, there are no current approaches that jointly estimate the location and size of the discontinuity set in a one-shot approach with statistical guarantees. We therefore introduce Free Discontinuity Regression (FDR), a fully nonparametric estimator that simultaneously (i) smooths a regression surface, (ii) segments it into contiguous regions, and (iii) provably recovers the precise locations and sizes of its jumps. By extending a convex relaxation of the Mumford-Shah functional to random spatial sampling and correlated noise, FDR overcomes the fixed-grid and i.i.d. noise assumptions of classical image-segmentation approaches, thus enabling its application to real-world data of any dimension. This yields the first identification and uniform consistency results for multivariate jump surfaces: under mild SBV regularity, the estimated function, its discontinuity set, and all jump sizes converge to their true population counterparts. Hyperparameters are selected automatically from the data using Stein's Unbiased Risk Estimate, and large-scale simulations up to three dimensions validate the theoretical results and demonstrate good finite-sample performance. Applying FDR to an internet shutdown in India reveals a 25-35% reduction in economic activity around the estimated shutdown boundaries-much larger than previous estimates. By unifying smoothing, segmentation, and effect-size recovery in a general statistical setting, FDR turns free-discontinuity ideas into a practical tool with formal guarantees for modern multivariate data.

econ.EM↗

An Optimal Transport Approach to Estimating Causal Effects via Nonlinear Difference-in-Differences

We propose a nonlinear difference-in-differences method to estimate multivariate counterfactual distributions in classical treatment and control study designs with observational data. Our approach sheds a new light on existing approaches like the changes-in-changes and the classical semiparametric difference-in-differences estimator and generalizes them to settings with multivariate heterogeneity in the outcomes. The main benefit of this extension is that it allows for arbitrary dependence and heterogeneity in the joint outcomes. We demonstrate its utility both on synthetic and real data. In particular, we revisit the classical Card \& Krueger dataset, examining the effect of a minimum wage increase on employment in fast food restaurants; a reanalysis with our method reveals that restaurants tend to substitute full-time with part-time labor after a minimum wage increase at a faster pace. A previous version of this work was entitled "An optimal transport approach to causal inference.

stat.ME↗

Tangential Wasserstein Projections

We develop a notion of projections between sets of probability measures using the geometric properties of the 2-Wasserstein space. It is designed for general multivariate probability measures, is computationally efficient to implement, and provides a unique solution in regular settings. The idea is to work on regular tangent cones of the Wasserstein space using generalized geodesics. Its structure and computational properties make the method applicable in a variety of settings, from causal inference to the analysis of object data. An application to estimating causal effects yields a generalization of the notion of synthetic controls to multivariate data with individual-level heterogeneity, as well as a way to estimate optimal weights jointly over all time periods.

stat.ML↗

Matching for causal effects via multimarginal unbalanced optimal transport

Matching on covariates is a well-established framework for estimating causal effects in observational studies. The principal challenge stems from the often high-dimensional structure of the problem. Many methods have been introduced to address this, with different advantages and drawbacks in computational and statistical performance as well as interpretability. This article introduces a natural optimal matching method based on multimarginal unbalanced optimal transport that possesses many useful properties in this regard. It provides interpretable weights based on the distance of matched individuals, can be efficiently implemented via the iterative proportional fitting procedure, and can match several treatment arms simultaneously. Importantly, the proposed method only selects good matches from either group, hence is competitive with the classical k-nearest neighbors approach in terms of bias and variance in finite samples. Moreover, we prove a central limit theorem for the empirical process of the potential functions of the optimal coupling in the unbalanced optimal transport problem with a fixed penalty term. This implies a parametric rate of convergence of the empirically obtained weights to the optimal weights in the population for a fixed penalty term.

stat.ME↗

Distributional synthetic controls

This article extends the widely-used synthetic controls estimator for evaluating causal effects of policy changes to quantile functions. The proposed method provides a geometrically faithful estimate of the entire counterfactual quantile function of the treated unit. Its appeal stems from an efficient implementation via a constrained quantile-on-quantile regression. This constitutes a novel concept of independent interest. The method provides a unique counterfactual quantile function in any scenario: for continuous, discrete or mixed distributions. It operates in both repeated cross-sections and panel data with as little as a single pre-treatment period. The article also provides abstract identification results by showing that any synthetic controls method, classical or our generalization, provides the correct counterfactual for causal models that preserve distances between the outcome distributions. Working with whole quantile functions instead of aggregate values allows for tests of equality and stochastic dominance of the counterfactual- and the observed distribution. It can provide causal inference on standard outcomes like average- or quantile treatment effects, but also more general concepts such as counterfactual Lorenz curves or interquartile ranges.

econ.EM↗

Non-testability of instrument validity under continuous endogenous variables

This note presents a proof of the conjecture in \citet*{pearl1995testability} about testing the validity of an instrumental variable in hidden variable models. It implies that instrument validity cannot be tested in the case where the endogenous treatment is continuously distributed. This stands in contrast to the classical testability results for instrument validity when the treatment is discrete. However, imposing weak structural assumptions on the model, such as continuity between the observable variables, can re-establish theoretical testability in the continuous setting.

econ.EM↗

A path-sampling method to partially identify causal effects in instrumental variable models

Partial identification approaches are a flexible and robust alternative to standard point-identification approaches in general instrumental variable models. However, this flexibility comes at the cost of a ``curse of cardinality'': the number of restrictions on the identified set grows exponentially with the number of points in the support of the endogenous treatment. This article proposes a novel path-sampling approach to this challenge. It is designed for partially identifying causal effects of interest in the most complex models with continuous endogenous treatments. A stochastic process representation allows to seamlessly incorporate assumptions on individual behavior into the model. Some potential applications include dose-response estimation in randomized trials with imperfect compliance, the evaluation of social programs, welfare estimation in demand models, and continuous choice models. As a demonstration, the method provides informative nonparametric bounds on household expenditures under the assumption that expenditure is continuous. The mathematical contribution is an approach to approximately solving infinite dimensional linear programs on path spaces via sampling.

econ.EM↗

Point-identification in multivariate nonseparable triangular models

In this article we introduce a general nonparametric point-identification result for nonseparable triangular models with a multivariate first- and second stage. Based on this we prove point-identification of Hedonic models with multivariate heterogeneity and endogenous observable characteristics, extending and complementing identification results from the literature which all require exogeneity. As an additional application of our theoretical result, we show that the BLP model (Berry et al. 1995) can also be identified without index restrictions.

econ.EM↗