SearcharxivSearch

arXiv subjects

Serena Ng

Publications and source records attributed to Serena Ng.

At least 19 recordsLinked to original sources

Cross-Section Estimation of Long-Run Relations Using Time-Compressed Data

Many empirical investigations of long-run relations are based on cross-section regressions in averaged or long differenced data that effectively have the time dimension of a $T\times N$ panel compressed. We analyze a class of time-compressed I(1) data and show that they have magnified variability stemming from the fact that the cross-section variance of a non-stationary panel `fans out' with time. Cross-section regressions in time compressed data can potentially yield estimates that are super-consistent and asymptotically normal, whether the regressors are stationary, non-stationary, or highly persistent. The fastest convergence rate of $\sqrt{N}T$ requires a compression scheme that not only magnifies the non-stationary signal, but also dilutes the regression noise. Omitted fixed effects preclude noise dilution but the estimates remain super-consistent. However, the fanning out effect can be weakened when the data have a strong force for mean-reversion or convergence, a problem that seems relevant for temperature data. We consider three applications and find that the long-run relation between consumption and income, and between growth/inflation and demographic variables are reasonably well determined, but the estimated relation between growth and warming temperature is fragile.

econ.EM

The Economic Impact of Low- and High-Frequency Temperature Changes

Variations in the low- and high-frequency components of temperature may have distinct impacts on economic outcomes. Parametric and non-parametric estimates from three panels of data all find significant heterogeneity in the relative importance of the two components, but there is clear evidence in each panel of a common, slowly evolving low-frequency factor that is highly correlated with the low-frequency factor of economic activity. In regressions that quantify the output effects of the components, we find that one-way clustered standard errors often lead to size distortions, and that an additive fixed effect specification does not adequately control for common time effects. Using bootstrap inference to assess estimates from our preferred interactive fixed effect specification, we only find a marginally significant effect of the high-frequency component on growth in the U.S. panel. However, the effect of the low-frequency component is significant in the European and International panels, suggesting that the increase in the low-frequency temperature component over the post-1980 period is associated with a reduction in economic growth of approximately 1.3 percentage points. The findings are corroborated by time series estimation using data at the unit and national levels.

econ.GN

Empirical Bayes Estimation in Heterogeneous Coefficient Panel Models

We develop an empirical Bayes (EB) G-modeling framework for short-panel linear models with nonparametric prior for the random intercepts, slopes, dynamics, and non-spherical error variances. We establish identification and consistency of the nonparametric maximum likelihood estimator (NPMLE) under general conditions, and provide low-level sufficient conditions for several models of empirical interest. Conditions for regret consistency of the EB estimators are also established. The NPMLE is computed using a Wasserstein-Fisher-Rao gradient flow algorithm adapted to panel regressions. Using data from the Panel Study of Income Dynamics, we find that the slope coefficient for potential experience is substantially heterogeneous and negatively correlated with the random intercept, and that error variances and autoregressive coefficients vary significantly across individuals. The EB estimates reduce mean squared prediction errors relative to individual maximum likelihood estimates.

econ.EM

Imputation of Counterfactual Outcomes when the Errors are Predictable

A crucial input into causal inference is the imputed counterfactual outcome. Imputation error can arise because of sampling uncertainty from estimating the prediction model using the untreated observations, or from out-of-sample information not captured by the model. While the literature has focused on sampling uncertainty, it vanishes with the sample size. Often overlooked is the possibility that the out-of-sample error can be informative about the missing counterfactual outcome if it is mutually or serially correlated. Motivated by the best linear unbiased predictor (\blup) of \citet{goldberger:62} in a time series setting, we propose an improved predictor of potential outcome when the errors are correlated. The proposed \pup\; is practical as it is not restricted to linear models, can be used with consistent estimators already developed, and improves mean-squared error for a large class of strong mixing error processes. Ignoring predictability in the errors can distort conditional inference. However, the precise impact will depend on the choice of estimator as well as the realized values of the residuals.

econ.EM

Constructing High Frequency Economic Indicators by Imputation

Monthly and weekly economic indicators are often taken to be the largest common factor estimated from high and low frequency data, either separately or jointly. To incorporate mixed frequency information without directly modeling them, we target a low frequency diffusion index that is already available, and treat high frequency values as missing. We impute these values using multiple factors estimated from the high frequency data. In the empirical examples considered, static matrix completion that does not account for serial correlation in the idiosyncratic errors yields imprecise estimates of the missing values irrespective of how the factors are estimated. Single equation and systems-based dynamic procedures that account for serial correlation yield imputed values that are closer to the observed low frequency ones. This is the case in the counterfactual exercise that imputes the monthly values of consumer sentiment series before 1978 when the data was released only on a quarterly basis. This is also the case for a weekly version of the CFNAI index of economic activity that is imputed using seasonally unadjusted data. The imputed series reveals episodes of increased variability of weekly economic information that are masked by the monthly data, notably around the 2014-15 collapse in oil prices.

econ.EM

Approximate Factor Models with Weaker Loadings

Pervasive cross-section dependence is increasingly recognized as a characteristic of economic data and the approximate factor model provides a useful framework for analysis. Assuming a strong factor structure where $\Lop\Lo/N^α$ is positive definite in the limit when $α=1$, early work established convergence of the principal component estimates of the factors and loadings up to a rotation matrix. This paper shows that the estimates are still consistent and asymptotically normal when $α\in(0,1]$ albeit at slower rates and under additional assumptions on the sample size. The results hold whether $α$ is constant or varies across factor loadings. The framework developed for heterogeneous loadings and the simplified proofs that can be also used in strong factor analysis are of independent interest.

econ.EM

Least Squares Estimation Using Sketched Data with Heteroskedastic Errors

Researchers may perform regressions using a sketch of data of size $m$ instead of the full sample of size $n$ for a variety of reasons. This paper considers the case when the regression errors do not have constant variance and heteroskedasticity robust standard errors would normally be needed for test statistics to provide accurate inference. We show that estimates using data sketched by random projections will behave `as if' the errors were homoskedastic. Estimation by random sampling would not have this property. The result arises because the sketched estimates in the case of random projections can be expressed as degenerate $U$-statistics, and under certain conditions, these statistics are asymptotically normal with homoskedastic variance. We verify that the conditions hold not only in the case of least squares regression when the covariates are exogenous, but also in instrumental variables estimation when the covariates are endogenous. The result implies that inference, including first-stage F tests for instrument relevance, can be simpler than the full sample case if the sketching scheme is appropriately chosen.

stat.ML

Time Series Estimation of the Dynamic Effects of Disaster-Type Shock

This paper provides three results for SVARs under the assumption that the primitive shocks are mutually independent. First, a framework is proposed to accommodate a disaster-type variable with infinite variance into a SVAR. We show that the least squares estimates of the SVAR are consistent but have non-standard asymptotics. Second, the disaster shock is identified as the component with the largest kurtosis and whose impact effect is negative. An estimator that is robust to infinite variance is used to recover the mutually independent components. Third, an independence test on the residuals pre-whitened by the Choleski decomposition is proposed to test the restrictions imposed on a SVAR. The test can be applied whether the data have fat or thin tails, and to over as well as exactly identified models. Three applications are considered. In the first, the independence test is used to shed light on the conflicting evidence regarding the role of uncertainty in economic fluctuations. In the second, disaster shocks are shown to have short term economic impact arising mostly from feedback dynamics. The third uses the framework to study the dynamic effects of economic shocks post-covid.

stat.ME

Factor-Based Imputation of Missing Values and Covariances in Panel Data of Large Dimensions

Economists are blessed with a wealth of data for analysis, but more often than not, values in some entries of the data matrix are missing. Various methods have been proposed to handle missing observations in a few variables. We exploit the factor structure in panel data of large dimensions. Our \textsc{tall-project} algorithm first estimates the factors from a \textsc{tall} block in which data for all rows are observed, and projections of variable specific length are then used to estimate the factor loadings. A missing value is imputed as the estimated common component which we show is consistent and asymptotically normal without further iteration. Implications for using imputed data in factor augmented regressions are then discussed. To compensate for the downward bias in covariance matrices created by an omitted noise when the data point is not observed, we overlay the imputed data with re-sampled idiosyncratic residuals many times and use the average of the covariances to estimate the parameters of interest. Simulations show that the procedures have desirable finite sample properties.

econ.EM

Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor structure holds for the full panel of data and its sub-blocks, it is shown that the common component can be consistently estimated at four different rates of convergence without requiring regularization or iteration. An asymptotic analysis of the estimation error is obtained. An application of our analysis is estimation of counterfactuals when potential outcomes have a factor structure. We study the estimation of average and individual treatment effects on the treated and establish a normal distribution theory that can be useful for hypothesis testing.

econ.EM

Modeling Macroeconomic Variations After COVID-19

The coronavirus is a global event of historical proportions and just a few months changed the time series properties of the data in ways that make many pre-covid forecasting models inadequate. It also creates a new problem for estimation of economic factors and dynamic causal effects because the variations around the outbreak can be interpreted as outliers, as shifts to the distribution of existing shocks, or as addition of new shocks. I take the latter view and use covid indicators as controls to 'de-covid' the data prior to estimation. I find that economic uncertainty remains high at the end of 2020 even though real economic activity has recovered and covid uncertainty has receded. Dynamic responses of variables to shocks in a VAR similar in magnitude and shape to the ones identified before 2020 can be recovered by directly or indirectly modeling covid and treating it as exogenous. These responses to economic shocks are distinctly different from those to a covid shock which are much larger but shorter lived. Disentangling the two types of shocks can be important in macroeconomic modeling post-covid.

econ.EM

Estimation and Inference by Stochastic Optimization: Three Examples

This paper illustrates two algorithms designed in Forneron & Ng (2020): the resampled Newton-Raphson (rNR) and resampled quasi-Newton (rqN) algorithms which speed-up estimation and bootstrap inference for structural models. An empirical application to BLP shows that computation time decreases from nearly 5 hours with the standard bootstrap to just over 1 hour with rNR, and only 15 minutes using rqN. A first Monte-Carlo exercise illustrates the accuracy of the method for estimation and inference in a probit IV regression. A second exercise additionally illustrates statistical efficiency gains relative to standard estimation for simulation-based estimation using a dynamic panel regression example.

econ.EM

Inference by Stochastic Optimization: A Free-Lunch Bootstrap

Assessing sampling uncertainty in extremum estimation can be challenging when the asymptotic variance is not analytically tractable. Bootstrap inference offers a feasible solution but can be computationally costly especially when the model is complex. This paper uses iterates of a specially designed stochastic optimization algorithm as draws from which both point estimates and bootstrap standard errors can be computed in a single run. The draws are generated by the gradient and Hessian computed from batches of data that are resampled at each iteration. We show that these draws yield consistent estimates and asymptotically valid frequentist inference for a large class of regular problems. The algorithm provides accurate standard errors in simulation examples and empirical applications at low computational costs. The draws from the algorithm also provide a convenient way to detect data irregularities.

econ.EM

Simpler Proofs for Approximate Factor Models of Large Dimensions

Estimates of the approximate factor model are increasingly used in empirical work. Their theoretical properties, studied some twenty years ago, also laid the ground work for analysis on large dimensional panel data models with cross-section dependence. This paper presents simplified proofs for the estimates by using alternative rotation matrices, exploiting properties of low rank matrices, as well as the singular value decomposition of the data in addition to its covariance structure. These simplifications facilitate interpretation of results and provide a more friendly introduction to researchers new to the field. New results are provided to allow linear restrictions to be imposed on factor models.

econ.EM

Latent Dirichlet Analysis of Categorical Survey Responses

Beliefs are important determinants of an individual's choices and economic outcomes, so understanding how they comove and differ across individuals is of considerable interest. Researchers often rely on surveys that report individual beliefs as qualitative data. We propose using a Bayesian hierarchical latent class model to analyze the comovements and observed heterogeneity in categorical survey responses. We show that the statistical model corresponds to an economic structural model of information acquisition, which guides interpretation and estimation of the model parameters. An algorithm based on stochastic optimization is proposed to estimate a model for repeated surveys when responses follow a dynamic structure and conjugate priors are not appropriate. Guidance on selecting the number of belief types is also provided. Two examples are considered. The first shows that there is information in the Michigan survey responses beyond the consumer sentiment index that is officially published. The second shows that belief types constructed from survey responses can be used in a subsequent analysis to estimate heterogeneous returns to education.

econ.EM

An Econometric Perspective on Algorithmic Subsampling

Datasets that are terabytes in size are increasingly common, but computer bottlenecks often frustrate a complete analysis of the data. While more data are better than less, diminishing returns suggest that we may not need terabytes of data to estimate a parameter or test a hypothesis. But which rows of data should we analyze, and might an arbitrary subset of rows preserve the features of the original data? This paper reviews a line of work that is grounded in theoretical computer science and numerical linear algebra, and which finds that an algorithmically desirable sketch, which is a randomly chosen subset of the data, must preserve the eigenstructure of the data, a property known as a subspace embedding. Building on this work, we study how prediction and inference can be affected by data sketching within a linear regression setup. We show that the sketching error is small compared to the sample size effect which a researcher can control. As a sketch size that is algorithmically optimal may not be suitable for prediction and inference, we use statistical arguments to provide 'inference conscious' guides to the sketch size. When appropriately implemented, an estimator that pools over different sketches can be nearly as efficient as the infeasible one using the full sample.

econ.EM

Boosting High Dimensional Predictive Regressions with Time Varying Parameters

High dimensional predictive regressions are useful in wide range of applications. However, the theory is mainly developed assuming that the model is stationary with time invariant parameters. This is at odds with the prevalent evidence for parameter instability in economic time series, but theories for parameter instability are mainly developed for models with a small number of covariates. In this paper, we present two $L_2$ boosting algorithms for estimating high dimensional models in which the coefficients are modeled as functions evolving smoothly over time and the predictors are locally stationary. The first method uses componentwise local constant estimators as base learner, while the second relies on componentwise local linear estimators. We establish consistency of both methods, and address the practical issues of choosing the bandwidth for the base learners and the number of boosting iterations. In an extensive application to macroeconomic forecasting with many potential predictors, we find that the benefits to modeling time variation are substantial and they increase with the forecast horizon. Furthermore, the timing of the benefits suggests that the Great Moderation is associated with substantial instability in the conditional mean of various economic series.

econ.EM

Principal Components and Regularized Estimation of Factor Models

It is known that the common factors in a large panel of data can be consistently estimated by the method of principal components, and principal components can be constructed by iterative least squares regressions. Replacing least squares with ridge regressions turns out to have the effect of shrinking the singular values of the common component and possibly reducing its rank. The method is used in the machine learning literature to recover low-rank matrices. We study the procedure from the perspective of estimating a minimum-rank approximate factor model. We show that the constrained factor estimates are biased but can be more efficient in terms of mean-squared errors. Rank consideration suggests a data-dependent penalty for selecting the number of factors. The new criterion is more conservative in cases when the nominal number of factors is inflated by the presence of weak factors or large measurement noise. The framework is extended to incorporate a priori linear constraints on the loadings. We provide asymptotic results that can be used to test economic hypotheses.

stat.ME