Searcharxiv⌕ Search

arXiv · 2609.40156

Partial identification with entropy regularized optimal transport

Abstract

In many statistical settings, the available data and maintained assumptions do not suffice to uniquely identify the model parameters of interest. In such cases, one can only identify sets which are guaranteed to contain the true parameters. These are often characterized through linear programs that optimize over models compatible with the observed data. These programs can be infinite-dimensional in the optimizer and the number of constraints. We provide a unified way to characterize and solve such optimization problems by phrasing them as optimal transport problems on path spaces. This allows us to regularize the problem with an entropy penalty, recasting it as a multi-marginal entropic optimal transport problem, which can be solved efficiently via Sinkhorn iterations. In addition, it allows us to establish convergence of the regularized value to the sharpest bound, derive consistency rates for a plug-in estimator, and obtain asymptotic distribution for approximate bounds. The method is general and accommodates settings ranging from instrumental variable models with continuous variables to welfare estimation in heterogeneous demand models. We verify the statistical and computational properties in simulations and provide an application to demand estimation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bruno N. Costa, Florian F. Gunsilius. 2026-09-30. Partial identification with entropy regularized optimal transport. https://arxiv.org/abs/2609.40156

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Causal Inference in Possibly Nonlinear Factor Models

This paper develops a causal inference method for treatment effects models with noisily measured confounders. The key feature is that a large number of noisy proxies are available and linked with the underlying latent confounders through an unknown, possibly nonlinear factor structure. The main building block is a local principal subspace approximation procedure that combines K-nearest-neighbor matching and principal component analysis. Estimators of many causal parameters, including average treatment effects and counterfactual distributions, are constructed based on doubly-robust score functions, and their large-sample properties are established. These results require the collection of proxies to be jointly informative about the latent confounders relevant to the outcome and treatment, while allowing some proxies to be uninformative, their identities to be unknown, and measurement errors to be correlated with the outcome or treatment. We also obtain uniformly consistent estimators of the conditional average treatment effect at each unit's confounder values. The results are illustrated with an empirical application studying the effect of political connections on stock returns of financial firms and a Monte Carlo experiment.

econ.EM↗

Constrained Classification and Policy Learning

Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empirical classification risk. These techniques are also useful for causal policy learning problems, since estimation of individualized treatment rules can be cast as a weighted (cost-sensitive) classification problem. Consistency of the surrogate loss approaches studied in Zhang (2004) and Bartlett et al. (2006) relies on the assumption of correct specification, which means that the specified set of classifiers is rich enough to contain a first-best classifier. This assumption is, however, less credible when interpretability or fairness constraints restrict the set of classifiers. Consequently, the applicability of surrogate-loss-based algorithms in such second-best scenarios remains unknown. This paper studies the consistency of surrogate loss procedures under a constrained set of classifiers without assuming correct specification. We show that in settings where the constraint restricts the classifier's prediction set only, hinge losses (i.e., $\ell_1$-support vector machines) are the only surrogate losses that preserve consistency in second-best scenarios. If the constraint additionally restricts the functional form of the classifier, consistency of a surrogate loss approach is not guaranteed, even with hinge loss. We therefore characterize conditions on the constrained set of classifiers that can guarantee consistency of hinge-risk-minimizing classifiers. Exploiting our theoretical results, we develop robust and computationally attractive hinge-loss-based procedures for a monotone classification problem.

econ.EM↗

Bounds on Average Effects in Discrete Choice Panel Data Models

In discrete choice panel data, estimation of average effects is crucial for quantifying the effect of covariates, and for policy evaluation and counterfactual analysis. However, in short panels with individual-specific effects, challenges arise due to partial identification and the incidental parameter problem. In particular, estimating the sharp identified set of average effects becomes impractical when covariates have large support sets, such as when they are continuous. This paper proposes a method for estimating outer bounds on the identified set of average effects, which are easy to construct, converge at the parametric rate, and remain computationally feasible even for moderately large samples. Asymptotically valid confidence intervals are also provided.

econ.EM↗