SearcharxivSearch

arXiv subjects

Ming-Yen Cheng

Publications and source records attributed to Ming-Yen Cheng.

9 recordsLinked to original sources

Distributionally Robust Recovery of Omitted Factors from Forecast Residuals with Application to Interest Rate Risk Management

A forecasting model compresses its predictors into an estimate of a conditional mean, and the systematic structure that estimate omits survives in the second moment of its forecast errors. Accuracy comparisons do not measure this structure, and variance-based extraction does not recover the part of it that a given decision bears. In this paper, we propose a distributionally robust framework that recovers the omitted structure from the residuals of a fixed forecaster: a decision is made robust over a two-layer moment ambiguity set on the standardized residual cross-section, and the discovery statistic is the covariance forcing, the component of the decision's residual risk transverse to its exposure. We demonstrate that the forcing is invariant to shrinkage and to isotropic inflation of the covariance, so the recovered direction is a property of the residuals rather than of the regularization, the sense in which the recovery is ground truth. This is confirmed on monthly U.S. Treasury zero-coupon yields from 2006 to 2025, where the recovered factor is named by factor-adjusted robust selection against a panel of 111 macroeconomic, Treasury supply-and-demand, and financial indicators. From the residuals of the linear factor-augmented dynamic Nelson-Siegel benchmark the factor names as a leading business-cycle factor, anchored on the Conference Board leading index and certified by a block-permutation test; from those of the more accurate nonlinear random forest benchmark the same procedure selects the same real-activity family without certification. A neutralization test completes the evidence: removing the recovered factor from the deployed duration position leaves volatility essentially unchanged and worsens the tail, so the factor is a material systematic risk the position bears.

q-fin.MF

Greedy Forward Regression for Variable Screening

Two popular variable screening methods under the ultra-high dimensional setting with the desirable sure screening property are the sure independence screening (SIS) and the forward regression (FR). Both are classical variable screening methods and recently have attracted greater attention under the new light of high-dimensional data analysis. We consider a new and simple screening method that incorporates multiple predictors in each step of forward regression, with decision on which variables to incorporate based on the same criterion. If only one step is carried out, it actually reduces to the SIS. Thus it can be regarded as a generalization and unification of the FR and the SIS. More importantly, it preserves the sure screening property and has similar computational complexity as FR in each step, yet it can discover the relevant covariates in fewer steps. Thus, it reduces the computational burden of FR drastically while retaining advantages of the latter over SIS. Furthermore, we show that it can find all the true variables if the number of steps taken is the same as the correct model size, even when using the original FR. An extensive simulation study and application to two real data examples demonstrate excellent performance of the proposed method.

stat.ME

Efficient estimation in semivarying coefficient models for longitudinal/clustered data

In semivarying coefficient models for longitudinal/clustered data, usually of primary interest is usually the parametric component which involves unknown constant coefficients. First, we study semiparametric efficiency bound for estimation of the constant coefficients in a general setup. It can be achieved by spline regression provided that the within-cluster covariance matrices are all known, which is an unrealistic assumption in reality. Thus, we propose an adaptive estimator of the constant coefficients when the covariance matrices are unknown and depend only on the index random variable, such as time, and when the link function is the identity function. After preliminary estimation, based on working independence and both spline and local linear regression, we estimate the covariance matrices by applying local linear regression to the resulting residuals. Then we employ the covariance matrix estimates and spline regression to obtain our final estimators of the constant coefficients. The proposed estimator achieves the semiparametric efficiency bound under normality assumption, and it has the smallest covariance matrix among a class of estimators even when normality is violated. We also present results of numerical studies. The simulation results demonstrate that our estimator is superior to the one based on working independence. When applied to the CD4 count data, our method identifies an interesting structure that was not found by previous analyses.

stat.ME

Forward variable selection for sparse ultra-high dimensional varying coefficient models

Varying coefficient models have numerous applications in a wide scope of scientific areas. While enjoying nice interpretability, they also allow flexibility in modeling dynamic impacts of the covariates. But, in the new era of big data, it is challenging to select the relevant variables when there are a large number of candidates. Recently several work are focused on this important problem based on sparsity assumptions; they are subject to some limitations, however. We introduce an appealing forward variable selection procedure. It selects important variables sequentially according to a sum of squares criterion, and it employs an EBIC- or BIC-based stopping rule. Clearly it is simple to implement and fast to compute, and it possesses many other desirable properties from both theoretical and numerical viewpoints. We establish rigorous selection consistency results when either EBIC or BIC is used as the stopping criterion, under some mild regularity conditions. Notably, unlike existing methods, an extra screening step is not required to ensure selection consistency. Even if the regularity conditions fail to hold, our procedure is still useful as an effective screening procedure in a less restrictive setup. We carried out simulation and empirical studies to show the efficacy and usefulness of our procedure.

stat.ME

Nonparametric independence screening and structure identification for ultra-high dimensional longitudinal data

Ultra-high dimensional longitudinal data are increasingly common and the analysis is challenging both theoretically and methodologically. We offer a new automatic procedure for finding a sparse semivarying coefficient model, which is widely accepted for longitudinal data analysis. Our proposed method first reduces the number of covariates to a moderate order by employing a screening procedure, and then identifies both the varying and constant coefficients using a group SCAD estimator, which is subsequently refined by accounting for the within-subject correlation. The screening procedure is based on working independence and B-spline marginal models. Under weaker conditions than those in the literature, we show that with high probability only irrelevant variables will be screened out, and the number of selected variables can be bounded by a moderate order. This allows the desirable sparsity and oracle properties of the subsequent structure identification step. Note that existing methods require some kind of iterative screening in order to achieve this, thus they demand heavy computational effort and consistency is not guaranteed. The refined semivarying coefficient model employs profile least squares, local linear smoothing and nonparametric covariance estimation, and is semiparametric efficient. We also suggest ways to implement the proposed methods, and to select the tuning parameters. An extensive simulation study is summarized to demonstrate its finite sample performance and the yeast cell cycle data is analyzed.

stat.ME

A New Test for One-Way ANOVA with Functional Data and Application to Ischemic Heart Screening

We propose and study a new global test, namely the $F_{\max}$-test, for the one-way ANOVA problem in functional data analysis. The test statistic is taken as the maximum value of the usual pointwise $F$-test statistics over the interval the functional responses are observed. A nonparametric bootstrap method is employed to approximate the null distribution of the test statistic and to obtain an estimated critical value for the test. The asymptotic random expression of the test statistic is derived and the asymptotic power is studied. In particular, under mild conditions, the $F_{\max}$-test asymptotically has the correct level and is root-$n$ consistent in detecting local alternatives. Via some simulation studies, it is found that in terms of both level accuracy and power, the $F_{\max}$-test outperforms the Globalized Pointwise F (GPF) test of \cite{Zhang_Liang:2013} when the functional data are highly or moderately correlated, and its performance is comparable with the latter otherwise. An application to an ischemic heart real dataset suggests that, after proper manipulation, resting electrocardiogram (ECG) signals can be used as an effective tool in clinical ischemic heart screening, without the need of further stress tests as in the current standard procedure.

math.ST

Nonparametric and adaptive modeling of dynamic seasonality and trend with heteroscedastic and dependent errors

Seasonality (or periodicity) and trend are features describing an observed sequence, and extracting these features is an important issue in many scientific fields. However, it is not an easy task for existing methods to analyze simultaneously the trend and {\it dynamics} of the seasonality such as time-varying frequency and amplitude, and the {\it adaptivity} of the analysis to such dynamics and robustness to heteroscedastic, dependent errors is not guaranteed. These tasks become even more challenging when there exist multiple seasonal components. We propose a nonparametric model to describe the dynamics of multi-component seasonality, and investigate the recently developed Synchrosqueezing transform (SST) in extracting these features in the presence of a trend and heteroscedastic, dependent errors. The identifiability problem of the nonparametric seasonality model is studied, and the adaptivity and robustness properties of the SST are theoretically justified in both discrete- and continuous-time settings. Consequently we have a new technique for de-coupling the trend, seasonality and heteroscedastic, dependent error process in a general nonparametric setup. Results of a series of simulations are provided, and the incidence time series of varicella and herpes zoster in Taiwan and respiratory signals observed from a sleep study are analyzed.

math.ST

Local Linear Regression on Manifolds and its Geometric Interpretation

High-dimensional data analysis has been an active area, and the main focuses have been variable selection and dimension reduction. In practice, it occurs often that the variables are located on an unknown, lower-dimensional nonlinear manifold. Under this manifold assumption, one purpose of this paper is regression and gradient estimation on the manifold, and another is developing a new tool for manifold learning. To the first aim, we suggest directly reducing the dimensionality to the intrinsic dimension $d$ of the manifold, and performing the popular local linear regression (LLR) on a tangent plane estimate. An immediate consequence is a dramatic reduction in the computation time when the ambient space dimension $p\gg d$. We provide rigorous theoretical justification of the convergence of the proposed regression and gradient estimators by carefully analyzing the curvature, boundary, and non-uniform sampling effects. A bandwidth selector that can handle heteroscedastic errors is proposed. To the second aim, we analyze carefully the behavior of our regression estimator both in the interior and near the boundary of the manifold, and make explicit its relationship with manifold learning, in particular estimating the Laplace-Beltrami operator of the manifold. In this context, we also make clear that it is important to use a smaller bandwidth in the tangent plane estimation than in the LLR. Simulation studies and the Isomap face data example are used to illustrate the computational speed and estimation accuracy of our methods.

math.ST

Reducing variance in univariate smoothing

A variance reduction technique in nonparametric smoothing is proposed: at each point of estimation, form a linear combination of a preliminary estimator evaluated at nearby points with the coefficients specified so that the asymptotic bias remains unchanged. The nearby points are chosen to maximize the variance reduction. We study in detail the case of univariate local linear regression. While the new estimator retains many advantages of the local linear estimator, it has appealing asymptotic relative efficiencies. Bandwidth selection rules are available by a simple constant factor adjustment of those for local linear estimation. A simulation study indicates that the finite sample relative efficiency often matches the asymptotic relative efficiency for moderate sample sizes. This technique is very general and has a wide range of applications.

math.ST