SearcharxivSearch

arXiv subjects

Wenyang Zhang

Publications and source records attributed to Wenyang Zhang.

18 recordsLinked to original sources

Feature Screening for High-Dimensional Structural Break Predictive Regression

Predictive regression is a crucial tool for exploring return predictability. In this study, we introduce an efficient procedure for selecting and estimating active predictors and change points in structural break predictive regression. Our approach allows the number of change points to increase with the sample size and accommodates sparse active predictors that may be stationary or cointegrated. We begin by identifying the active predictors using a Sure Independence Canonical Screening (SICS) procedure. Next, we estimate the change points through a Ratio-Controlled Regression Screening (RCRS) method. Finally, we reduce redundancy by eliminating unnecessary breakpoints and predictors using information criteria (IC). This approach allows for consistent estimation and selection of true breakpoints and active predictors. Our simulations and empirical studies demonstrate that the proposed procedure performs effectively.

stat.ME

Two-Stage Robust Sparse Gradient Methods for Regression Under Heavy-Tailed Designs

We study high-dimensional sparse regression under simultaneous heavy-tailed covariates and noise. Heavy-tailed data affect sparse optimization in two different ways: extreme covariates can destabilize the gradient field during global localization, while heavy-tailed noise limits the final statistical accuracy during local refinement. Motivated by this two-phase structure, we propose two-stage RIGHT, a robust sparse first-order method based on coordinate-wise median-of-means (MoM) gradient estimation and delayed sample splitting. The MoM gradient estimator is computationally simple, compatible with hard-thresholded updates, and admits phase-adaptive concentration bounds whose rates depend on the current localization radius. Delayed splitting reuses data during global localization and reserves fresh batches for the shorter refinement stage, reducing the sample-splitting cost. The theoretical results reveal a decoupled rate structure: the design-tail index controls gradient stability and sample complexity, whereas the noise-tail index controls the final statistical rate. We also provide phase-wise lower-bound benchmarks showing that the design-driven localization barrier is intrinsic. Extensive simulation experiments and real data analysis showcase the efficacy of the proposed method over existing competitors.

stat.ME

Scale-Invariant Robust Estimation of High-Dimensional Kronecker-Structured Matrices

High-dimensional Kronecker-structured estimation faces a conflict between non-convex scaling ambiguities and statistical robustness. The arbitrary factor scaling distorts gradient magnitudes, rendering standard fixed-threshold robust methods ineffective. We resolve this via Scaled Robust Gradient Descent (SRGD), which stabilizes optimization by de-scaling gradients before truncation. To further enforce interpretability, we introduce Scaled Hard Thresholding (SHT) for invariant variable selection. A two-step estimation procedure, built upon robust initialization and SRGD--SHT iterative updates, is proposed for canonical matrix problems, such as trace regression, matrix GLMs, and bilinear models. The convergence rates are established for heavy-tailed predictors and noise, identifying a phase transition where optimal convergence rates recover under finite noise variance and degrade optimally for heavier tails. Experiments on simulated data and two real-world applications confirm superior robustness and efficiency of the proposed procedure.

stat.ME

High-dimensional low-rank matrix regression with unknown latent structures

We study low-rank matrix regression in settings where matrix-valued predictors and scalar responses are observed across multiple individuals. Rather than assuming a fully homogeneous coefficient matrices across individuals, we accommodate shared low-dimensional structure alongside individual-specific deviations. To this end, we introduce a tensor-structured homogeneity pursuit framework, wherein each coefficient matrix is represented as a product of shared low-rank subspaces and individualized low-rank loadings. We propose a scalable estimation procedure based on scaled gradient descent, and establish non-asymptotic bounds demonstrating that the proposed estimator attains improved convergence rates by leveraging shared information while preserving individual-specific signals. The framework is further extended to incorporate scaled hard thresholding for recovering sparse latent structures, with theoretical guarantees in both linear and generalized linear model settings. Our approach provides a principled middle ground between fully pooled and fully separate analyses, achieving strong theoretical performance, computational tractability, and interpretability in high-dimensional multi-individual matrix regression problems.

stat.ME

Workflow for High-Fidelity Dynamic Analysis of Structures with Pile Foundation

The demand for high-fidelity numerical simulations in soil-structure interaction analysis is on the rise, yet a standardized workflow to guide the creation of such simulations remains elusive. This paper aims to bridge this gap by presenting a step-by-step guideline proposing a workflow for dynamic analysis of structures with pile foundations. The proposed workflow encompasses instructions on how to use Domain Reduction Method for loading, Perfectly Matched Layer elements for wave absorption, soil-structure interaction modeling using Embedded interface elements, and domain decomposition for efficient use of processing units. Through a series of numerical simulations, we showcase the practical application of this workflow. Our results reveal the efficacy of the Domain Reduction Method in reducing simulation size without compromising model fidelity, show the precision of Perfectly Matched Layer elements in modeling infinite domains, highlight the efficiency of Embedded Interface elements in establishing connections between structures and the soil domain, and demonstrate the overall effectiveness of the proposed workflow in conducting high-fidelity simulations. While our study focuses on simplified geometries and loading scenarios, it serves as a foundational framework for future research endeavors aimed at exploring more intricate structural configurations and dynamic loading conditions

math.NA

Robust estimation for high-dimensional time series with heavy tails

We study in this paper the problem of least absolute deviation (LAD) regression for high-dimensional heavy-tailed time series which have finite $α$-th moment with $α\in (1,2]$. To handle the heavy-tailed dependent data, we propose a Catoni type truncated minimization problem framework and obtain an $\mathcal{O}\big( \big( (d_1+d_2) (d_1\land d_2) \log^2 n / n \big)^{(α- 1)/α} \big)$ order excess risk, where $d_1$ and $d_2$ are the dimensionality and $n$ is the number of samples. We apply our result to study the LAD regression on high-dimensional heavy-tailed vector autoregressive (VAR) process. Simulations for the VAR($p$) model show that our new estimator with truncation are essential because the risk of the classical LAD has a tendency to blow up. We further apply our estimation to the real data and find that ours fits the data better than the classical LAD.

math.ST

Complex trend inference for high-dimensional piecewise locally stationary time series

This paper studies high-dimensional trend inference for piecewise smooth signals under nonstationary noise and asynchronous structural breaks by first detecting asynchronous changes without assuming stationarity and then further exploiting latent group structures to estimate trend functions. In the first step, we propose AJDN (Asynchronous Jump Detection under Nonstationary Noise), a multiscale framework for the identification and localization of jumps in high-dimensional time series. We show that AJDN consistently recovers the number of jumps with a prescribed asymptotic probability and achieves nearly optimal localization rates in the presence of asynchronicity and nonstationarity, both of which often violate the assumptions of existing high-dimensional change point methods and thereby deteriorate their performance. In the second step, we augment AJDN with a homogeneity pursuit step and obtain AJDN-H, which identifies latent groups of dimensions that share common jump structures and trend parameters given the detected jumps. This allows for efficient information pooling and improves the accuracy of trend estimation under both asynchronicity and nonstationarity. The robustness and finite-sample performance of the proposed methodology are examined by extensive simulation studies. An application to financial data demonstrates the practical utility of the AJDN-H framework in complex, high-dimensional settings.

stat.ME

A Flexible and Parsimonious Modelling Strategy for Clustered Data Analysis

Statistical modelling strategy is the key for success in data analysis. The trade-off between flexibility and parsimony plays a vital role in statistical modelling. In clustered data analysis, in order to account for the heterogeneity between the clusters, certain flexibility is necessary in the modelling, yet parsimony is also needed to guard against the complexity and account for the homogeneity among the clusters. In this paper, we propose a flexible and parsimonious modelling strategy for clustered data analysis. The strategy strikes a nice balance between flexibility and parsimony, and accounts for both heterogeneity and homogeneity well among the clusters, which often come with strong practical meanings. In fact, its usefulness has gone beyond clustered data analysis, it also sheds promising lights on transfer learning. An estimation procedure is developed for the unknowns in the resulting model, and asymptotic properties of the estimators are established. Intensive simulation studies are conducted to demonstrate how well the proposed methods work, and a real data analysis is also presented to illustrate how to apply the modelling strategy and associated estimation procedure to answer some real problems arising from real life.

stat.ME

Model Averaging based Semiparametric Modelling for Conditional Quantile Prediction

In real data analysis, the underlying model is usually unknown, modelling strategy plays a key role in the success of data analysis. Stimulated by the idea of model averaging, we propose a novel semiparametric modelling strategy for conditional quantile prediction, without assuming the underlying model is any specific parametric or semiparametric model. Thanks the optimality of the selected weights by cross-validation, the proposed modelling strategy results in a more accurate prediction than that based on some commonly used semiparametric models, such as the varying coefficient models and additive models. Asymptotic properties are established of the proposed modelling strategy together with its estimation procedure. Intensive simulation studies are conducted to demonstrate how well the proposed method works, compared with its alternatives under various circumstances. The results show the proposed method indeed leads to more accurate predictions than its alternatives. Finally, the proposed modelling strategy together with its prediction procedure are applied to the Boston housing data, which result in more accurate predictions of the quantiles of the house prices than that based on some commonly used alternative methods, therefore, present us a more accurate picture of the housing market in Boston.

stat.ME

Multi-Kink Quantile Regression for Longitudinal Data with Application to the Progesterone Data Analysis

Motivated by investigating the relationship between progesterone and the days in a menstrual cycle in a longitudinal study, we propose a multi-kink quantile regression model for longitudinal data analysis. It relaxes the linearity condition and assumes different regression forms in different regions of the domain of the threshold covariate. In this paper, we first propose a multi-kink quantile regression for longitudinal data. Two estimation procedures are proposed to estimate the regression coefficients and the kink points locations: one is a computationally efficient profile estimator under the working independence framework while the other one considers the within-subject correlations by using the unbiased generalized estimation equation approach. The selection consistency of the number of kink points and the asymptotic normality of two proposed estimators are established. Secondly, we construct a rank score test based on partial subgradients for the existence of kink effect in longitudinal studies. Both the null distribution and the local alternative distribution of the test statistic have been derived. Simulation studies show that the proposed methods have excellent finite sample performance. In the application to the longitudinal progesterone data, we identify two kink points in the progesterone curves over different quantiles and observe that the progesterone level remains stable before the day of ovulation, then increases quickly in five to six days after ovulation and then changes to stable again or even drops slightly

stat.ME

Estimation and Inference for Multi-Kink Quantile Regression

The Multi-Kink Quantile Regression (MKQR) model is an important tool for analyzing data with heterogeneous conditional distributions, especially when quantiles of response variable are of interest, due to its robustness to outliers and heavy-tailed errors in the response. It assumes different linear quantile regression forms in different regions of the domain of the threshold covariate but are still continuous at kink points. In this paper, we investigate parameter estimation, kink point detection and statistical inference in MKQR models. We propose an iterative segmented quantile regression algorithm for estimating both the regression coefficients and the locations of kink points. The proposed algorithm is much more computationally efficient than the grid search algorithm and not sensitive to the selection of initial values. Asymptotic properties, such as selection consistency of the number of kink points, asymptotic normality of the estimators of both regression coefficients and kink effects, are established to justify the proposed method theoretically. A score test, based on partial subgradients, is developed to verify whether the kink effects exist or not. Test-inversion confidence intervals for kink location parameters are also constructed. Intensive simulation studies conducted show the proposed methods work very well when sample size is finite. Finally, we apply the MKQR models together with the proposed methods to the dataset about secondary industrial structure of China and the dataset about triceps skinfold thickness of Gambian females, which leads to some very interesting findings. A new R package MultiKink is developed to implement the proposed methods.

stat.ME

Homogeneity Pursuit in Single Index Models based Panel Data Analysis

Panel data analysis is an important topic in statistics and econometrics. Traditionally, in panel data analysis, all individuals are assumed to share the same unknown parameters, e.g. the same coefficients of covariates when the linear models are used, and the differences between the individuals are accounted for by cluster effects. This kind of modelling only makes sense if our main interest is on the global trend, this is because it would not be able to tell us anything about the individual attributes which are sometimes very important. In this paper, we proposed a modelling based on the single index models embedded with homogeneity for panel data analysis, which builds the individual attributes in the model and is parsimonious at the same time. We develop a data driven approach to identify the structure of homogeneity, and estimate the unknown parameters and functions based on the identified structure. Asymptotic properties of the resulting estimators are established. Intensive simulation studies conducted in this paper also show the resulting estimators work very well when sample size is finite. Finally, the proposed modelling is applied to a public financial dataset and a UK climate dataset, the results reveal some interesting findings.

math.ST

Factor Models for Asset Returns Based on Transformed Factors

The Fama-French three factor models are commonly used in the description of asset returns in finance. Statistically speaking, the Fama-French three factor models imply that the return of an asset can be accounted for directly by the Fama-French three factors, i.e. market, size and value factor, through a linear function. A natural question is: would some kind of transformed Fama-French three factors work better than the three factors? If so, what kind of transformation should be imposed on each factor in order to make the transformed three factors better account for asset returns? In this paper, we are going to address these questions through nonparametric modelling. We propose a data driven approach to construct the transformation for each factor concerned. A generalised maximum likelihood ratio based hypothesis test is also proposed to test whether transformations on the Fama-French three factors are needed for a given data set. Asymptotic properties are established to justify the proposed methods. Intensive simulation studies are conducted to show how the proposed methods work when sample size is finite. Finally, we apply the proposed methods to a real data set, which leads to some interesting findings.

stat.ME

Model selection and structure specification in ultra-high dimensional generalised semi-varying coefficient models

In this paper, we study the model selection and structure specification for the generalised semi-varying coefficient models (GSVCMs), where the number of potential covariates is allowed to be larger than the sample size. We first propose a penalised likelihood method with the LASSO penalty function to obtain the preliminary estimates of the functional coefficients. Then, using the quadratic approximation for the local log-likelihood function and the adaptive group LASSO penalty (or the local linear approximation of the group SCAD penalty) with the help of the preliminary estimation of the functional coefficients, we introduce a novel penalised weighted least squares procedure to select the significant covariates and identify the constant coefficients among the coefficients of the selected covariates, which could thus specify the semiparametric modelling structure. The developed model selection and structure specification approach not only inherits many nice statistical properties from the local maximum likelihood estimation and nonconcave penalised likelihood method, but also computationally attractive thanks to the computational algorithm that is proposed to implement our method. Under some mild conditions, we establish the asymptotic properties for the proposed model selection and estimation procedure such as the sparsity and oracle property. We also conduct simulation studies to examine the finite sample performance of the proposed method, and finally apply the method to analyse a real data set, which leads to some interesting findings.

math.ST

A Dynamic Structure for High Dimensional Covariance Matrices and its Application in Portfolio Allocation

Estimation of high dimensional covariance matrices is an interesting and important research topic. In this paper, we propose a dynamic structure and develop an estimation procedure for high dimensional covariance matrices. Asymptotic properties are derived to justify the estimation procedure and simulation studies are conducted to demonstrate its performance when the sample size is finite. By exploring a financial application, an empirical study shows that portfolio allocation based on dynamic high dimensional covariance matrices can significantly outperform the market from 1995 to 2014. Our proposed method also outperforms portfolio allocation based on the sample covariance matrix and the portfolio allocation proposed in Fan, Fan and Lv (2008).

stat.ME

A semiparametric spatial dynamic model

Stimulated by the Boston house price data, in this paper, we propose a semiparametric spatial dynamic model, which extends the ordinary spatial autoregressive models to accommodate the effects of some covariates associated with the house price. A profile likelihood based estimation procedure is proposed. The asymptotic normality of the proposed estimators are derived. We also investigate how to identify the parametric/nonparametric components in the proposed semiparametric model. We show how many unknown parameters an unknown bivariate function amounts to, and propose an AIC/BIC of nonparametric version for model selection. Simulation studies are conducted to examine the performance of the proposed methods. The simulation results show our methods work very well. We finally apply the proposed methods to analyze the Boston house price data, which leads to some interesting findings.

math.ST

A semiparametric model for cluster data

In the analysis of cluster data, the regression coefficients are frequently assumed to be the same across all clusters. This hampers the ability to study the varying impacts of factors on each cluster. In this paper, a semiparametric model is introduced to account for varying impacts of factors over clusters by using cluster-level covariates. It achieves the parsimony of parametrization and allows the explorations of nonlinear interactions. The random effect in the semiparametric model also accounts for within-cluster correlation. Local, linear-based estimation procedure is proposed for estimating functional coefficients, residual variance and within-cluster correlation matrix. The asymptotic properties of the proposed estimators are established, and the method for constructing simultaneous confidence bands are proposed and studied. In addition, relevant hypothesis testing problems are addressed. Simulation studies are carried out to demonstrate the methodological power of the proposed methods in the finite sample. The proposed model and methods are used to analyse the second birth interval in Bangladesh, leading to some interesting findings.

math.ST

Estimation of the covariance matrix of random effects in longitudinal studies

Longitudinal studies are often conducted to explore the cohort and age effects in many scientific areas. The within cluster correlation structure plays a very important role in longitudinal data analysis. This is because not only can an estimator be improved by incorporating the within cluster correlation structure into the estimation procedure, but also the within cluster correlation structure can sometimes provide valuable insights in practical problems. For example, it can reveal the correlation strengths among the impacts of various factors. Motivated by data typified by a set from Bangladesh pertinent to the use of contraceptives, we propose a random effect varying-coefficient model, and an estimation procedure for the within cluster correlation structure of the proposed model. The estimation procedure is optimization-free and the proposed estimators enjoy asymptotic normality under mild conditions. Simulations suggest that the proposed estimation is practicable for finite samples and resistent against mild forms of model misspecification. Finally, we analyze the data mentioned above with the new random effect varying-coefficient model together with the proposed estimation procedure, which reveals some interesting sociological dynamics.

math.ST