Searcharxiv⌕ Search

arXiv subjects

Wayne Yuan Gao

Publications and source records attributed to Wayne Yuan Gao.

At least 19 recordsLinked to original sources

Single-Network Finite-Sample Inference in Strategic Network Formation Models

We develop a finite-sample valid inference procedure for strategic network formation models with endogenous network statistics, using only a single observed network. We impose no restrictions on network density, equilibrium selection, or the dependence structure induced by strategic interaction. Using a "bounding-by-c" technique, we obtain realization-wise sandwich inequalities whose middle term depends only on exogenous covariates and i.i.d. pairwise shocks. These inequalities deliver pathwise identifying restrictions and simulable finite-sample critical values in both parametric and semiparametric settings. Our proposed procedure is also computationally tractable: it does not involve solving, simulating, or enumerating equilibrium (sub)networks, and easily scales to networks with 10000 agents in simulations. In two empirical applications, we find statistical evidence for positive link interdependence at the 95% confidence level.

econ.EM↗

Dynamic Heterogeneous Distribution Regression Panel Models, with an Application to Labor Income Processes

We introduce a dynamic distribution regression panel data model with heterogeneous coefficients across units. The objects of primary interest are functionals of these coefficients, including predicted one-step-ahead and stationary cross-sectional distributions of the outcome variable. Coefficients and their functionals are estimated via fixed effect methods. We investigate how these functionals vary in response to counterfactual changes in initial conditions or covariate values. We also identify a uniformity problem related to the robustness of inference to the unknown degree of coefficient heterogeneity, and propose a cross-sectional bootstrap method for uniformly valid inference on function-valued objects. We showcase the utility of our approach through an empirical application to individual income dynamics. Employing the annual Panel Study of Income Dynamics data, we establish the presence of substantial coefficient heterogeneity. We then highlight some important empirical questions that our methodology can address. First, we quantify the impact of a negative labor income shock on the distribution of future labor income. Second, we demonstrate the existence of heterogeneity in income mobility, and its implications for an individuals' incidence to be trapped in poverty. Simulation evidence confirms that our procedures work well in small samples.

econ.EM↗

Variance or Standard Deviation? Shell Geometry and Global-Scale Priors in High-Dimensional Shrinkage

We study how the choice of default prior for a common Gaussian scale affects high-dimensional shrinkage risk, highlighting the role played by high-dimensional geometry. Formally, we consider a high-dimensional setting in which the near-zero behavior of the common scale prior has first-order consequences for shrinkage risk, and show that priors that are flat on the variance and those flat on the standard deviation allocate markedly different mass near the zero-scale boundary, leading to distinct shrinkage behavior and informing principled default prior selection. Specifically, under a radial-power benchmark, we establish that the SD-flat benchmark has a one-unit asymptotic risk advantage near the origin, crosses over in the critical regime, and is second-order equivalent to the variance-flat benchmark for strong signals. Proper single global-scale hyperpriors and bounded coordinate-multiplier mixtures inherit these limits through the near-zero exponent of their SD-scale density. For heavier-tailed or sparse priors, that exponent still classifies the common global-scale component, while local-scale tails, model-size priors, or allocation priors can also affect risk.

stat.ME↗

Identification in Dynamic Dyadic Network Formation Models with Fixed Effects

This paper establishes (set) identification results in a dynamic dyadic network formation model with time-varying observed covariates, lagged local network statistics, and unobserved heterogeneity in the form of fixed effects. Our framework accommodates observed-covariate homophily, transitivity through common friends, second-order or indirect-friend effects, and more general local subgraph statistics within a single dynamic index model. The analysis combines two complementary ways of handling fixed effects: inequalities that integrate out time-invariant dyad heterogeneity by treating each dyad as a short panel, and signed-subgraph comparisons that difference out fixed effects algebraically through intertemporal variation within each dyad. We show that the semiparametric identifying restrictions can be sharpened using either or both of the following assumptions: (i) error distribution is serially independent with a known distribution, (ii) pairwise fixed effect takes the form of additive individual fixed effects. Combining (i) and (ii) under i.i.d. logit shocks, we obtain an exact conditional logit representation and provide sufficient conditions for point identification.

econ.EM↗

Tractable Identification of Strategic Network Formation Models with Unobserved Heterogeneity

We develop a tractable identification approach for strategic network formation models with both strategic link interdependence and individual unobserved heterogeneity (fixed effects). The key challenge is that endogenous network statistics (e.g. number of common friends) enter the link formation equation, while the mapping from model primitives to equilibrium network structure is generally intractable. Our approach sidesteps this difficulty using a ``bounding-by-$c$'' technique that treats endogenous covariates as random variables and exploits monotonicity restrictions to obtain identifying information. A central contribution is to develop a spectrum of fixed-effects handling strategies based on subnetwork configurations: tetrad-based restrictions that difference out all individual fixed effects, triad-based and weighted restrictions that combine ``difference-out'' and ``integrate-out'' steps by differencing out some fixed effects and profiling over the remainder conditional on observed characteristics, and general weighted cycle-based restrictions that unify these cases. We also provide point identification results. Preliminary simulations show that the approach can deliver informative bounds on the structural parameters.

econ.EM↗

Model Restrictiveness in Functional and Structural Settings

We extend the restrictiveness measure of Fudenberg, Gao & Liang (2026) to functional and structural econometric settings using Gaussian process priors. We find that models evaluated over continuum domains appear more restrictive than when evaluated over finite sets of observations. We also extend the restrictiveness framework to structural models with endogeneity, instrumental variables, multiple equilibria, and nonparametric nuisance components. We explain why the choice of discrepancy function is a substantive modeling decision, and why the Rademacher complexity and GMM criterion functions are unsuitable as discrepancies. We further show that restrictiveness equals the normalized limit of the noise-free average-case learning curve. In applications to preferences under risk, and multinomial choice under exogenous and endogenous settings, we find that the same models exhibit uniformly higher restrictiveness when evaluated over continuum domains than based on their predictions on finite sets, and that moment restrictions from endogeneity substantially increase restrictiveness and alter model rankings.

econ.GN↗

Thin Sets Are Not Equally Thin: Minimax Learning of Submanifold Integrals

Many economic parameters are identified by ``thin sets'' (submanifolds with Lebesgue measure zero) and hence difficult to recover from data in an ambient space. This paper provides a unified theory for estimation and inference of such ``thin-set'' identified functionals. We show that thin sets are \emph{not} equally thin: their intrinsic dimensionality $m$ matters in a precise manner. For a nonparametric regression $h_0$ with Hölder smoothness $s$ and $d$-dimensional covariates in the ambient space, we show that $n^{-\frac{s}{2s+d-m}}$ is the minimax optimal rate of estimating linear and nonlinear (e.g., quadratic, upper contour) integrals of $h_0$ on an $m$-dimensional submanifold ($0\leq m < d$), which is the fastest possible attainable rate among all estimators. The minimax lower bound rate result is generalized to estimating submanifold integrals when $h_0$ is a nonparametric density and a nonparametric instrumental variable function. The asymptotic normality of t statistics is established via sieve Riesz representation, and the corresponding inference is computed using Sobol points.

econ.EM↗

Identification in Nonlinear Dynamic Panel Models under Partial Stationarity

This paper provides a general identification approach for a wide range of nonlinear panel data models, including binary choice, ordered response, and other types of limited dependent variable models. Our approach accommodates dynamic models with any number of lagged dependent variables as well as other types of endogenous covariates. Our identification strategy relies on a partial stationarity condition, which allows for not only an unknown distribution of errors, but also temporal dependencies in errors. We derive partial identification results under flexible model specifications and establish sharpness of our identified set in the binary choice setting. We demonstrate the robust finite-sample performance of our approach using Monte Carlo simulations, and apply the approach to the empirical analysis of income categories using various ordered choice models.

econ.EM↗

Identification of Semiparametric Panel Multinomial Choice Models with Infinite-Dimensional Fixed Effects

This paper proposes a robust method for semiparametric identification and estimation in panel multinomial choice models, where we allow for infinite-dimensional fixed effects that enter into consumer utilities in an additively nonseparable way, thus incorporating rich forms of unobserved heterogeneity. Our identification strategy exploits multivariate monotonicity in parametric indexes, and uses the logical contraposition of an intertemporal inequality on choice probabilities to obtain identifying restrictions. We provide a consistent estimation procedure, and demonstrate the practical advantages of our method with Monte Carlo simulations and an empirical illustration on popcorn sales with the Nielsen data.

econ.EM↗

ReLU-Based and DNN-Based Generalized Maximum Score Estimators

We propose a new formulation of the maximum score estimator that uses compositions of rectified linear unit (ReLU) functions, instead of indicator functions as in Manski (1975,1985), to encode the sign alignment restrictions. Since the ReLU function is Lipschitz, our new ReLU-based maximum score criterion function is substantially easier to optimize using standard gradient-based optimization pacakges. We also show that our ReLU-based maximum score (RMS) estimator can be generalized to an umbrella framework defined by multi-index single-crossing (MISC) conditions, while the original maximum score estimator cannot be applied. We establish the $n^{-s/(2s+1)}$ convergence rate and asymptotic normality for the RMS estimator under order-$s$ Holder smoothness. In addition, we propose an alternative estimator using a further reformulation of RMS as a special layer in a deep neural network (DNN) architecture, which allows the estimation procedure to be implemented via state-of-the-art software and hardware for DNN.

econ.EM↗

Optimization via Strategic Law of Large Numbers

This paper proposes a unified framework for the global optimization of a continuous function in a bounded rectangular domain. Specifically, we show that: (1) under the optimal strategy for a two-armed decision model, the sample mean converges to a global optimizer under the Strategic Law of Large Numbers, and (2) a sign-based strategy built upon the solution of a parabolic PDE is asymptotically optimal. Motivated by this result, we propose a class of {\bf S}trategic {\bf M}onte {\bf C}arlo {\bf O}ptimization (SMCO) algorithms, which uses a simple strategy that makes coordinate-wise two-armed decisions based on the signs of the partial gradient of the original function being optimized over (without the need of solving PDEs). While this simple strategy is not generally optimal, we show that it is sufficient for our SMCO algorithm to converge to local optimizer(s) from a single starting point, and to global optimizers under a growing set of starting points. Numerical studies demonstrate the suitability of our SMCO algorithms for global optimization, and illustrate the promise of our theoretical framework and practical approach. For a wide range of test functions with challenging optimization landscapes (including ReLU neural networks with square and hinge loss), our SMCO algorithms converge to the global maximum accurately and robustly, using only a small set of starting points (at most 100 for dimensions up to 1000) and a small maximum number of iterations (200). In fact, our algorithms outperform many state-of-the-art global optimizers, as well as local algorithms augmented with the same set of starting points as ours.

math.OC↗

Inference on Welfare and Value Functionals under Optimal Treatment Assignment

We provide theoretical results for the estimation and inference of a class of welfare and value functionals of the nonparametric conditional average treatment effect (CATE) function under optimal treatment assignment, i.e., treatment is assigned to an observed type if and only if its CATE is nonnegative. For the optimal welfare functional defined as the average value of CATE on the subpopulation with nonnegative CATE, we establish the $\sqrt{n}$ asymptotic normality of the semiparametric plug-in estimators and provide an analytical asymptotic variance formula. For more general value functionals, we show that the plug-in estimators are typically asymptotically normal at the 1-dimensional nonparametric estimation rate, and we provide a consistent variance estimator based on the sieve Riesz representer, as well as a proposed computational procedure for numerical integration on submanifolds. The key reason underlying the different convergence rates for the welfare functional versus the general value functional lies in that, on the boundary subpopulation for whom CATE is zero, the integrand vanishes for the welfare functional but does not for general value functionals. We demonstrate in Monte Carlo simulations the good finite-sample performance of our estimation and inference procedures, and conduct an empirical application of our methods on the effectiveness of job training programs on earnings using the JTPA data set.

econ.EM↗

IV Regressions without Exclusion Restrictions

We study identification and estimation of endogenous linear and nonlinear regression models without excluded instrumental variables, based on the standard mean independence condition and a nonlinear relevance condition. Based on the identification results, we propose two semiparametric estimators as well as a discretization-based estimator that does not require any nonparametric regressions. We establish their asymptotic normality and demonstrate via simulations their robust finite-sample performances with respect to exclusion restrictions violations and endogeneity. Our approach is applied to study the returns to education, and to test the direct effects of college proximity indicators as well as family background variables on the outcome.

econ.EM↗

A Partial Order on Preference Profiles

We propose a theoretical framework under which preference profiles can be meaningfully compared. Specifically, given a finite set of feasible allocations and a preference profile, we first define a ranking vector of an allocation as the vector of all individuals' rankings of this allocation. We then define a partial order on preference profiles and write "$P \geq P^{'}$", if there exists an onto mapping $ψ$ from the Pareto frontier of $P^{'}$ onto the Pareto frontier of $P$, such that the ranking vector of any Pareto efficient allocation $x$ under $P^{'}$ is weakly dominated by the ranking vector of the image allocation $ψ(x)$ under $P$. We provide a characterization of the maximal and minimal elements under the partial order. In particular, we illustrate how an individualistic form of social preferences can be maximal in a specific setting. We also discuss how the framework can be further generalized to incorporate additional economic ingredients.

econ.TH↗

Two-Stage Maximum Score Estimator

This paper considers the asymptotic theory of a semiparametric M-estimator that is generally applicable to models that satisfy a monotonicity condition in one or several parametric indexes. We call the estimator two-stage maximum score (TSMS) estimator since our estimator involves a first-stage nonparametric regression when applied to the binary choice model of Manski (1975, 1985). We characterize the asymptotic distribution of the TSMS estimator, which features phase transitions depending on the dimension and thus the convergence rate of the first-stage estimation. Effectively, the first-stage nonparametric estimator serves as an imperfect smoothing function on a non-smooth criterion function, leading to the pivotality of the first-stage estimation error with respect to the second-stage convergence rate and asymptotic distribution

econ.EM↗

Using Monotonicity Restrictions to Identify Models with Partially Latent Covariates

This paper develops a new method for identifying econometric models with partially latent covariates. Such data structures arise in industrial organization and labor economics settings where data are collected using an input-based sampling strategy, e.g., if the sampling unit is one of multiple labor input factors. We show that the latent covariates can be nonparametrically identified, if they are functions of a common shock satisfying some plausible monotonicity assumptions. With the latent covariates identified, semiparametric estimation of the outcome equation proceeds within a standard IV framework that accounts for the endogeneity of the covariates. We illustrate the usefulness of our method using a new application that focuses on the production functions of pharmacies. We find that differences in technology between chains and independent pharmacies may partially explain the observed transformation of the industry structure.

econ.EM↗

Logical Differencing in Dyadic Network Formation Models with Nontransferable Utilities

This paper considers a semiparametric model of dyadic network formation under nontransferable utilities (NTU). NTU arises frequently in real-world social interactions that require bilateral consent, but by its nature induces additive non-separability. We show how unobserved individual heterogeneity in our model can be canceled out without additive separability, using a novel method we call logical differencing. The key idea is to construct events involving the intersection of two mutually exclusive restrictions on the unobserved heterogeneity, based on multivariate monotonicity. We provide a consistent estimator and analyze its performance via simulation, and apply our method to the Nyakatoke risk-sharing networks.

econ.EM↗

Nonparametric Identification in Index Models of Link Formation

We consider an index model of dyadic link formation with a homophily effect index and a degree heterogeneity index. We provide nonparametric identification results in a single large network setting for the potentially nonparametric homophily effect function, the realizations of unobserved individual fixed effects and the unknown distribution of idiosyncratic pairwise shocks, up to normalization, for each possible true value of the unknown parameters. We propose a novel form of scale normalization on an arbitrary interquantile range, which is not only theoretically robust but also proves particularly convenient for the identification analysis, as quantiles provide direct linkages between the observable conditional probabilities and the unknown index values. We then use an inductive "in-fill and out-expansion" algorithm to establish our main results, and consider extensions to more general settings that allow nonseparable dependence between homophily and degree heterogeneity, as well as certain extents of network sparsity and weaker assumptions on the support of unobserved heterogeneity. As a byproduct, we also propose a concept called "modeling equivalence" as a refinement of "observational equivalence", and use it to provide a formal discussion about normalization, identification and their interplay with counterfactuals.

econ.EM↗