SearcharxivSearch

arXiv subjects

Qingliang Fan

Publications and source records attributed to Qingliang Fan.

13 recordsLinked to original sources

Identification and Robust Inference for Multiple Treatments with Possibly Invalid Instruments

The instrumental variable (IV) method is widely used to infer causal effects in observational studies with unmeasured confounding, but invalid instruments can compromise both population identification and finite-sample inference. This paper studies linear IV models with multiple endogenous treatments and possibly invalid instruments. Identification of multiple effects is more delicate than in the single-treatment setting because a single instrument no longer identifies a single candidate effect; instead, each relevant instrument defines a hyperplane in the multidimensional effect space. For identification of multiple treatment effects, we introduce generalized plurality and majority rules which require a sufficiently large number of IVs to be valid. For inference, data-dependent instrument selection may fail to separate certain invalid IVs from valid ones, leading to undercoverage of confidence intervals when these invalid instruments are mistakenly selected as valid. We propose a sampling confidence interval for each treatment effect, which is robust to IV selection errors. We establish asymptotic coverage and parametric-rate length of our sampling confidence interval under regularity conditions and illustrate this method in Monte Carlo simulations and a Mendelian randomization application.

stat.ME

Adaptive Multi-task Learning for Multi-sector Portfolio Optimization

Accurate transfer of information across multiple sectors to enhance model estimation is both significant and challenging in multi-sector portfolio optimization involving a large number of assets in different classes. Within the framework of factor modeling, we propose a novel data-adaptive multi-task learning methodology that quantifies and learns the relatedness among the principal temporal subspaces (spanned by factors) across multiple sectors under study. This approach not only improves the simultaneous estimation of multiple factor models but also enhances multi-sector portfolio optimization, which heavily depends on the accurate recovery of these factor models. Additionally, a novel and easy-to-implement algorithm, termed projection-penalized principal component analysis, is developed to accomplish the multi-task learning procedure. Diverse simulation designs and practical application on daily return data from Russell 3000 index demonstrate the advantages of multi-task learning methodology.

stat.ME

Single-Index Quantile Factor Model with Observed Characteristics

We propose a characteristics-augmented quantile factor (QCF) model, where unknown factor loading functions are linked to a large set of observed individual-level (e.g., bond- or stock-specific) covariates via a single-index projection. The single-index specification offers a parsimonious, interpretable, and statistically efficient way to nonparametrically characterize the time-varying loadings, while avoiding the curse of dimensionality in flexible nonparametric models. Using a three-step sieve estimation procedure, the QCF model demonstrates high in-sample and out-of-sample accuracy in simulations. We establish asymptotic properties for estimators of the latent factor, loading functions, and index parameters. In an empirical study, we analyze the dynamic distributional structure of U.S. corporate bond returns from 2003 to 2020. Our method outperforms the benchmark quantile Fama-French five-factor model and quantile latent factor model, particularly in the tails ($\tau=0.05, 0.95$). The model reveals state-dependent risk exposures driven by characteristics such as bond and equity volatility, coupon, and spread. Finally, we provide economic interpretations of the latent factors.

econ.EM

Cost-aware Portfolios in a Large Universe of Assets

This paper considers the finite horizon portfolio rebalancing problem in terms of mean-variance optimization, where decisions are made based on current information on asset returns and transaction costs. The study's novelty is that the transaction costs are integrated within the optimization problem in a high-dimensional portfolio setting where the number of assets is larger than the sample size. We propose portfolio construction and rebalancing models with nonconvex penalty considering two types of transaction cost, the proportional transaction cost and the quadratic transaction cost. We establish the desired theoretical properties under mild regularity conditions. Monte Carlo simulations and empirical studies using S&P 500 and Russell 2000 stocks show the satisfactory performance of the proposed portfolio and highlight the importance of involving the transaction costs when rebalancing a portfolio.

stat.ME

Robust Bond Risk Premia Predictability Test in the Quantiles

Different from existing literature on testing the macro-spanning hypothesis of bond risk premia, which only considers mean regressions, this paper investigates whether the yield curve represented by CP factor (Cochrane and Piazzesi, 2005) contains all available information about future bond returns in a predictive quantile regression with many other macroeconomic variables. In this study, we introduce the Trend in Debt Holding (TDH) as a novel predictor, testing it alongside established macro indicators such as Trend Inflation (TI) (Cieslak and Povala, 2015), and macro factors from Ludvigson and Ng (2009). A significant challenge in this study is the invalidity of traditional quantile model inference approaches, given the high persistence of many macro variables involved. Furthermore, the existing methods addressing this issue do not perform well in the marginal test with many highly persistent predictors. Thus, we suggest a robust inference approach, whose size and power performance are shown to be better than existing tests. Using data from 1980-2022, the macro-spanning hypothesis is strongly supported at center quantiles by the empirical finding that the CP factor has predictive power while all other macro variables have negligible predictive power in this case. On the other hand, the evidence against the macro-spanning hypothesis is found at tail quantiles, in which TDH has predictive power at right tail quantiles while TI has predictive power at both tails quantiles. Finally, we show the performance of in-sample and out-of-sample predictions implemented by the proposed method are better than existing methods.

econ.EM

Shocks-adaptive Robust Minimum Variance Portfolio for a Large Universe of Assets

This paper proposes a robust, shocks-adaptive portfolio in a large-dimensional assets universe where the number of assets could be comparable to or even larger than the sample size. It is well documented that portfolios based on optimizations are sensitive to outliers in return data. We deal with outliers by proposing a robust factor model, contributing methodologically through the development of a robust principal component analysis (PCA) for factor model estimation and a shrinkage estimation for the random error covariance matrix. This approach extends the well-regarded Principal Orthogonal Complement Thresholding (POET) method (Fan et al., 2013), enabling it to effectively handle heavy tails and sudden shocks in data. The novelty of the proposed robust method is its adaptiveness to both global and idiosyncratic shocks, without the need to distinguish them, which is useful in forming portfolio weights when facing outliers. We develop the theoretical results of the robust factor model and the robust minimum variance portfolio. Numerical and empirical results show the superior performance of the new portfolio.

q-fin.PM

Portfolio Analysis in High Dimensions with TE and Weight Constraints

This paper explores the statistical properties of forming constrained optimal portfolios within a high-dimensional set of assets. We examine portfolios with tracking error constraints, those with simultaneous tracking error and weight restrictions, and portfolios constrained solely by weight. Tracking error measures portfolio performance against a benchmark (typically an index), while weight constraints determine asset allocation based on regulatory requirements or fund prospectuses. Our approach employs a novel statistical learning technique that integrates factor models with nodewise regression, named the Constrained Residual Nodewise Optimal Weight Regression (CROWN) method. We demonstrate its estimation consistency in large dimensions, even when assets outnumber the portfolio's time span. Convergence rate results for constrained portfolio weights, risk, and Sharpe Ratio are provided, and simulation and empirical evidence highlight the method's outstanding performance.

q-fin.PM

Robust Inference for Multiple Predictive Regressions with an Application on Bond Risk Premia

We propose a robust hypothesis testing procedure for the predictability of multiple predictors that could be highly persistent. Our method improves the popular extended instrumental variable (IVX) testing (Phillips and Lee, 2013; Kostakis et al., 2015) in that, besides addressing the two bias effects found in Hosseinkouchack and Demetrescu (2021), we find and deal with the variance-enlargement effect. We show that two types of higher-order terms induce these distortion effects in the test statistic, leading to significant over-rejection for one-sided tests and tests in multiple predictive regressions. Our improved IVX-based test includes three steps to tackle all the issues above regarding finite sample bias and variance terms. Thus, the test statistics perform well in size control, while its power performance is comparable with the original IVX. Monte Carlo simulations and an empirical study on the predictability of bond risk premia are provided to demonstrate the effectiveness of the newly proposed approach.

stat.ME

Inference for Nonlinear Endogenous Treatment Effects Accounting for High-Dimensional Covariate Complexity

Nonlinearity and endogeneity are prevalent challenges in causal analysis using observational data. This paper proposes an inference procedure for a nonlinear and endogenous marginal effect function, defined as the derivative of the nonparametric treatment function, with a primary focus on an additive model that includes high-dimensional covariates. Using the control function approach for identification, we implement a regularized nonparametric estimation to obtain an initial estimator of the model. Such an initial estimator suffers from two biases: the bias in estimating the control function and the regularization bias for the high-dimensional outcome model. Our key innovation is to devise the double bias correction procedure that corrects these two biases simultaneously. Building on this debiased estimator, we further provide a confidence band of the marginal effect function. Simulations and an empirical study of air pollution and migration demonstrate the validity of our procedures.

econ.EM

On the instrumental variable estimation with many weak and invalid instruments

We discuss the fundamental issue of identification in linear instrumental variable (IV) models with unknown IV validity. With the assumption of the "sparsest rule", which is equivalent to the plurality rule but becomes operational in computation algorithms, we investigate and prove the advantages of non-convex penalized approaches over other IV estimators based on two-step selections, in terms of selection consistency and accommodation for individually weak IVs. Furthermore, we propose a surrogate sparsest penalty that aligns with the identification condition and provides oracle sparse structure simultaneously. Desirable theoretical properties are derived for the proposed estimator with weaker IV strength conditions compared to the previous literature. Finite sample properties are demonstrated using simulations and the selection and estimation method is applied to an empirical study concerning the effect of BMI on diastolic blood pressure.

stat.ME

A Heteroskedasticity-Robust Overidentifying Restriction Test with High-Dimensional Covariates

This paper proposes an overidentifying restriction test for high-dimensional linear instrumental variable models. The novelty of the proposed test is that it allows the number of covariates and instruments to be larger than the sample size. The test is scale-invariant and is robust to heteroskedastic errors. To construct the final test statistic, we first introduce a test based on the maximum norm of multiple parameters that could be high-dimensional. The theoretical power based on the maximum norm is higher than that in the modified Cragg-Donald test (Koles\'{a}r, 2018), the only existing test allowing for large-dimensional covariates. Second, following the principle of power enhancement (Fan et al., 2015), we introduce the power-enhanced test, with an asymptotically zero component used to enhance the power to detect some extreme alternatives with many locally invalid instruments. Finally, an empirical example of the trade and economic growth nexus demonstrates the usefulness of the proposed test.

econ.EM

Estimation of Conditional Average Treatment Effects with High-Dimensional Data

Given the unconfoundedness assumption, we propose new nonparametric estimators for the reduced dimensional conditional average treatment effect (CATE) function. In the first stage, the nuisance functions necessary for identifying CATE are estimated by machine learning methods, allowing the number of covariates to be comparable to or larger than the sample size. The second stage consists of a low-dimensional local linear regression, reducing CATE to a function of the covariate(s) of interest. We consider two variants of the estimator depending on whether the nuisance functions are estimated over the full sample or over a hold-out sample. Building on Belloni at al. (2017) and Chernozhukov et al. (2018), we derive functional limit theory for the estimators and provide an easy-to-implement procedure for uniform inference based on the multiplier bootstrap. The empirical application revisits the effect of maternal smoking on a baby's birth weight as a function of the mother's age.

econ.EM

Endogenous Treatment Effect Estimation with some Invalid and Irrelevant Instruments

Instrumental variables (IV) regression is a popular method for the estimation of the endogenous treatment effects. Conventional IV methods require all the instruments are relevant and valid. However, this is impractical especially in high-dimensional models when we consider a large set of candidate IVs. In this paper, we propose an IV estimator robust to the existence of both the invalid and irrelevant instruments (called R2IVE) for the estimation of endogenous treatment effects. This paper extends the scope of Kang et al. (2016) by considering a true high-dimensional IV model and a nonparametric reduced form equation. It is shown that our procedure can select the relevant and valid instruments consistently and the proposed R2IVE is root-n consistent and asymptotically normal. Monte Carlo simulations demonstrate that the R2IVE performs favorably compared to the existing high-dimensional IV estimators (such as, NAIVE (Fan and Zhong, 2018) and sisVIVE (Kang et al., 2016)) when invalid instruments exist. In the empirical study, we revisit the classic question of trade and growth (Frankel and Romer, 1999).

econ.EM