SearcharxivSearch

arXiv subjects

Youngki Shin

Publications and source records attributed to Youngki Shin.

At least 19 recordsLinked to original sources

SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models

Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose Stochastic Approximation with Unbiased Simulated Scores (SAUSS), an averaged stochastic approximation based on conditionally unbiased mini-batch score estimates. Each iteration uses a fixed mini-batch regardless of sample size. For multinomial probit, accept-reject sampling provides exact conditional draws and unbiased score estimates for any fixed number of accepted draws. Under local conditions, asymptotic theory for the averaged estimator and the partial-sum process of the SAUSS iterates incorporates mini-batch and simulation variability and supports random-scaling and plug-in inference. In simulations and an application, SAUSS gives comparable results in less than 1% of the computation time of simulated maximum likelihood. SAUSS extends to limited dependent variable models with conditional-expectation score representations and exact conditional sampling.

stat.ME

Ranking Treatment Saturations under Clustered Network Interference

In this paper, we study how to rank a finite set of treatment saturations for a target population with clustered network interference. We propose an empirical success (ES) ranking rule that, for each pair of saturations, selects the saturation level with the higher estimated welfare using data from a two-stage randomized saturation design. We adopt the statistical decision theory framework with additively separable regret loss to assess the performance of the ES ranking rule. We derive non-asymptotic upper bounds on the maximum regret of the ES ranking rule that depend on the within-cluster network only through a single combinatorial summary of its dependency structure. We exploit these bounds to characterize a quasi-optimal first-stage saturation distribution within the two-stage randomized saturation design. We further show that the ES ranking rule is asymptotically optimal among threshold ranking rules in the sense of minimizing an upper bound on the worst-case regret.

econ.EM

The Evolution of Unobserved Skill Returns in the U.S.: A New Approach Using Panel Data

Economists disagree about the factors driving the substantial increase in residual wage inequality in the US over the past few decades. To identify changes in the returns to unobserved skills, we make a novel assumption about the dynamics of skills rather than about the stability of skill distributions across cohorts, as is standard. We show that our assumption is supported by data on test score dynamics for older workers in the HRS. Using survey data from the PSID and administrative data from the IRS and SSA, we estimate that the returns to unobserved skills $declined$ substantially in the late-1980s and 1990s despite an increase in residual inequality. Accounting for firm-specific pay differences yields similar results. Extending our framework to consider occupational differences in returns to skill and multiple unobserved skills, we further show that skill returns display similar patterns for workers employed in each of cognitive, routine, and social occupations. Finally, our results suggest that increasing skill dispersion, driven by rising skill volatility, explains most of the growth in residual wage inequality since the 1980s.

econ.GN

Optimal Wage Band for Job Matching with Signaling

We study an optimal wage band problem in a competitive matching labor market where education signals worker ability. We prove uniqueness of the competitive signaling equilibrium under a general class of utility and profit functions and show that the optimal wage band problem is isomorphic to a simpler optimal ability threshold problem. Using a parametric model, we analyze how wage bands improve welfare relative to no intervention. Our results highlight novel mechanisms driven by asymmetric information, contrasting with existing literature. The framework is broadly applicable to settings where agents invest in costly signals under asymmetric information in competitive matching environments.

econ.TH

Fast Inference for Quantile Regression with Tens of Millions of Observations

Big data analytics has opened new avenues in economic research, but the challenge of analyzing datasets with tens of millions of observations is substantial. Conventional econometric methods based on extreme estimators require large amounts of computing resources and memory, which are often not readily available. In this paper, we focus on linear quantile regression applied to "ultra-large" datasets, such as U.S. decennial censuses. A fast inference framework is presented, utilizing stochastic subgradient descent (S-subGD) updates. The inference procedure handles cross-sectional data sequentially: (i) updating the parameter estimate with each incoming "new observation", (ii) aggregating it as a $\textit{Polyak-Ruppert}$ average, and (iii) computing a pivotal statistic for inference using only a solution path. The methodology draws from time-series regression to create an asymptotically pivotal statistic through random scaling. Our proposed test statistic is calculated in a fully online fashion and critical values are calculated without resampling. We conduct extensive numerical studies to showcase the computational merits of our proposed inference. For inference problems as large as $(n, d) \sim (10^7, 10^3)$, where $n$ is the sample size and $d$ is the number of regressors, our method generates new insights, surpassing current inference methods in computation. Our method specifically reveals trends in the gender gap in the U.S. college wage premium using millions of observations, while controlling over $10^3$ covariates to mitigate confounding effects.

econ.EM

SGMM: Stochastic Approximation to Generalized Method of Moments

We introduce a new class of algorithms, Stochastic Generalized Method of Moments (SGMM), for estimation and inference on (overidentified) moment restriction models. Our SGMM is a novel stochastic approximation alternative to the popular Hansen (1982) (offline) GMM, and offers fast and scalable implementation with the ability to handle streaming datasets in real time. We establish the almost sure convergence, and the (functional) central limit theorem for the inefficient online 2SLS and the efficient SGMM. Moreover, we propose online versions of the Durbin-Wu-Hausman and Sargan-Hansen tests that can be seamlessly integrated within the SGMM framework. Extensive Monte Carlo simulations show that as the sample size increases, the SGMM matches the standard (offline) GMM in terms of estimation accuracy and gains over computational efficiency, indicating its practical value for both large-scale and online datasets. We demonstrate the efficacy of our approach by a proof of concept using two well known empirical examples with large sample sizes.

econ.EM

csa2sls: A complete subset approach for many instruments using Stata

We develop a Stata command $\texttt{csa2sls}$ that implements the complete subset averaging two-stage least squares (CSA2SLS) estimator in Lee and Shin (2021). The CSA2SLS estimator is an alternative to the two-stage least squares estimator that remedies the bias issue caused by many correlated instruments. We conduct Monte Carlo simulations and confirm that the CSA2SLS estimator reduces both the mean squared error and the estimation bias substantially when instruments are correlated. We illustrate the usage of $\texttt{csa2sls}$ in Stata by an empirical application.

econ.EM

Optimal Delegation in Markets for Matching with Signaling

This paper studies a delegation problem faced by the planner who wants to regulate receivers' reaction choices in markets for matching between receivers and senders with signaling. We provide a noble insight into the planner's willingness to delegate and the design of optimal (reaction) interval delegation as a solution to the planner's general mechanism design problem. The relative heterogeneity of receiver types and the productivity of the sender' signal are crucial in deriving optimal interval delegation in the presence of the trade-off between matching efficiency and signaling costs.

econ.TH

Statistical Treatment Rules under Social Interaction

In this paper we study treatment assignment rules in the presence of social interaction. We construct an analytical framework under the anonymous interaction assumption, where the decision problem becomes choosing a treatment fraction. We propose a multinomial empirical success (MES) rule that includes the empirical success rule of Manski (2004) as a special case. We investigate the non-asymptotic bounds of the expected utility based on the MES rule. Finally, we prove that the MES rule achieves the asymptotic optimality with the minimax regret criterion.

econ.EM

Fast and Robust Online Inference with Stochastic Gradient Descent via Random Scaling

We develop a new method of online inference for a vector of parameters estimated by the Polyak-Ruppert averaging procedure of stochastic gradient descent (SGD) algorithms. We leverage insights from time series regression in econometrics and construct asymptotically pivotal statistics via random scaling. Our approach is fully operational with online data and is rigorously underpinned by a functional central limit theorem. Our proposed inference method has a couple of key advantages over the existing methods. First, the test statistic is computed in an online fashion with only SGD iterates and the critical values can be obtained without any resampling methods, thereby allowing for efficient implementation suitable for massive online data. Second, there is no need to estimate the asymptotic variance and our inference method is shown to be robust to changes in the tuning parameters for SGD algorithms in simulation experiments with synthetic data.

stat.ML

Monotone Equilibrium in Matching Markets with Signaling

We introduce a notion of competitive signaling equilibrium (CSE) in one-to-one matching markets with a continuum of heterogeneous senders and receivers. We then study monotone CSE where equilibrium outcomes - sender actions, receiver reactions, beliefs, and matching - are all monotone in the stronger set order. We show that if the sender utility is monotone-supermodular and the receiver's utility is weakly monotone-supermodular, a CSE is stronger monotone if and only if it passes Criterion D1 (Cho and Kreps (1987), Banks and Sobel (1987)). Given any interval of feasible reactions that receivers can take, we fully characterize a unique stronger monotone CSE and establishes its existence with quasilinear utility functions.

econ.TH

Complete Subset Averaging for Quantile Regressions

We propose a novel conditional quantile prediction method based on complete subset averaging (CSA) for quantile regressions. All models under consideration are potentially misspecified and the dimension of regressors goes to infinity as the sample size increases. Since we average over the complete subsets, the number of models is much larger than the usual model averaging method which adopts sophisticated weighting schemes. We propose to use an equal weight but select the proper size of the complete subset based on the leave-one-out cross-validation method. Building upon the theory of Lu and Su (2015), we investigate the large sample properties of CSA and show the asymptotic optimality in the sense of Li (1987). We check the finite sample performance via Monte Carlo simulations and empirical applications.

econ.EM

Predictive Quantile Regression with Mixed Roots and Increasing Dimensions: The ALQR Approach

In this paper we propose the adaptive lasso for predictive quantile regression (ALQR). Reflecting empirical findings, we allow predictors to have various degrees of persistence and exhibit different signal strengths. The number of predictors is allowed to grow with the sample size. We study regularity conditions under which stationary, local unit root, and cointegrated predictors are present simultaneously. We next show the convergence rates, model selection consistency, and asymptotic distributions of ALQR. We apply the proposed method to the out-of-sample quantile prediction problem of stock returns and find that it outperforms the existing alternatives. We also provide numerical evidence from additional Monte Carlo experiments, supporting the theoretical results.

econ.EM

Exact Computation of Maximum Rank Correlation Estimator

In this paper we provide a computation algorithm to get a global solution for the maximum rank correlation estimator using the mixed integer programming (MIP) approach. We construct a new constrained optimization problem by transforming all indicator functions into binary parameters to be estimated and show that it is equivalent to the original problem. We also consider an application of the best subset rank prediction and show that the original optimization problem can be reformulated as MIP. We derive the non-asymptotic bound for the tail probability of the predictive performance measure. We investigate the performance of the MIP algorithm by an empirical example and Monte Carlo simulations.

econ.EM

Factor-Driven Two-Regime Regression

We propose a novel two-regime regression model where regime switching is driven by a vector of possibly unobservable factors. When the factors are latent, we estimate them by the principal component analysis of a panel data set. We show that the optimization problem can be reformulated as mixed integer optimization, and we present two alternative computational algorithms. We derive the asymptotic distribution of the resulting estimator under the scheme that the threshold effect shrinks to zero. In particular, we establish a phase transition that describes the effect of first-stage factor estimation as the cross-sectional dimension of panel data increases relative to the time-series dimension. Moreover, we develop bootstrap inference and illustrate our methods via numerical studies.

econ.EM

Sparse HP Filter: Finding Kinks in the COVID-19 Contact Rate

In this paper, we estimate the time-varying COVID-19 contact rate of a Susceptible-Infected-Recovered (SIR) model. Our measurement of the contact rate is constructed using data on actively infected, recovered and deceased cases. We propose a new trend filtering method that is a variant of the Hodrick-Prescott (HP) filter, constrained by the number of possible kinks. We term it the $\textit{sparse HP filter}$ and apply it to daily data from five countries: Canada, China, South Korea, the UK and the US. Our new method yields the kinks that are well aligned with actual events in each country. We find that the sparse HP filter provides a fewer kinks than the $\ell_1$ trend filter, while both methods fitting data equally well. Theoretically, we establish risk consistency of both the sparse HP and $\ell_1$ trend filters. Ultimately, we propose to use time-varying $\textit{contact growth rates}$ to document and monitor outbreaks of COVID-19.

econ.EM

Desperate times call for desperate measures: government spending multipliers in hard times

We investigate state-dependent effects of fiscal multipliers and allow for endogenous sample splitting to determine whether the US economy is in a slack state. When the endogenized slack state is estimated as the period of the unemployment rate higher than about 12 percent, the estimated cumulative multipliers are significantly larger during slack periods than non-slack periods and are above unity. We also examine the possibility of time-varying regimes of slackness and find that our empirical results are robust under a more flexible framework. Our estimation results point out the importance of the heterogenous effects of fiscal policy and shed light on the prospect of fiscal policy in response to economic shocks from the current COVID-19 pandemic.

econ.GN

Complete Subset Averaging with Many Instruments

We propose a two-stage least squares (2SLS) estimator whose first stage is the equal-weighted average over a complete subset with $k$ instruments among $K$ available, which we call the complete subset averaging (CSA) 2SLS. The approximate mean squared error (MSE) is derived as a function of the subset size $k$ by the Nagar (1959) expansion. The subset size is chosen by minimizing the sample counterpart of the approximate MSE. We show that this method achieves the asymptotic optimality among the class of estimators with different subset sizes. To deal with averaging over a growing set of irrelevant instruments, we generalize the approximate MSE to find that the optimal $k$ is larger than otherwise. An extensive simulation experiment shows that the CSA-2SLS estimator outperforms the alternative estimators when instruments are correlated. As an empirical illustration, we estimate the logistic demand function in Berry, Levinsohn, and Pakes (1995) and find the CSA-2SLS estimate is better supported by economic theory than the alternative estimates.

econ.EM