SearcharxivSearch

arXiv subjects

Zetai Cen

Publications and source records attributed to Zetai Cen.

9 recordsLinked to original sources

Learning Perturbations to Extrapolate Your LLM

Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose a framework where token prefixes are perturbed by a learnable transformation of a continuous latent vector within an embedding space. To overcome the challenge of an intractable marginal likelihood, we derive unbiased estimating equations for model parameters and optimize them via stochastic gradient descent. We establish the statistical properties of the resulting estimator in over-parameterized regimes. Empirical evaluations on both synthetic and real-world datasets demonstrate that our proposal yields significant gains in out-of-domain settings over a range of state-of-the-art baseline methods.

stat.ML

Perturbation is All You Need for Extrapolating Language Models

This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models. In contrast to the standard autoregressive next-token prediction based on an exact prefix, we introduce a perturbation-based procedure that first transforms the prefix into a semantic neighbour and then conditions on this perturbed variant for next-token prediction. This yields a hierarchical model with a pre-post-additive noise structure. Within this framework, we develop a rigorous theory of extrapolability, namely, the capacity of a model class to make reliable predictions for token sequences that lie outside the empirical support of the training corpus, by establishing five properties of the proposed procedure: adaptivity, contractivity, robustness, extrapolability, and double robustness. We evaluate the finite sample performance of the proposed procedure using both synthetic and real world language data. Results show that the proposed method consistently improves out-of-support prediction while maintaining competitive in-support performance, demonstrating that perturbation offers a practical route to language modelling.

stat.ML

Detection and Mode-Identification of Multiple Change Points in Tensor Factor Models

We study the problems arising from modeling high-dimensional tensor-valued time series under a Tucker decomposition-based factor model with multiple structural change points. First, we propose an algorithm for detecting the multiple change points, which utilizes the low-rank structure of the data for statistical and computational efficiency. Also, the multi-dimensional array setting poses unique challenges, as some changes are associated with a subset of the modes, and the changes in different modes may interact with one another. Recognizing these, we investigate the problem of identifying each change with the tensor modes post-segmentation. To this end, we formalize the mode-identifiability of each change and propose an algorithm for detecting the modes at which the data are undergoing a mode-identifiable shift. We establish the consistency of both change point detection and mode-identification methods under a weak moment condition, and demonstrate their good performance on simulated datasets where, in particular, it is shown that the mode-identification step can improve the post-segmentation estimation of the mode-wise loading space. Additionally we analyze the datasets on New York City taxi usage and Fama--French portfolio returns using the proposed suite of methods.

math.ST

Identification and Estimation of Multi-order Tensor Factor Models

We propose a novel framework in high-dimensional factor models to simultaneously analyse multiple tensor time series, each with potentially different tensor orders and dimensionality. The connection between different tensor time series is through their global factors that are correlated to each other. A salient feature of our model is that when all tensor time series have the same order, it can be regarded as an extension of multilevel factor models from vectors to general tensors. Under very mild conditions, we separate the global and local components in the proposed model. Parameter estimation is thoroughly discussed, including a consistent factor number estimator. With strong correlation between global factors and noise allowed, we derive the rates of convergence of our estimators, which can be more superior than those of existing methods for multilevel factor models. We also develop estimators that are more computationally efficient, with rates of convergence spelt out. Extensive experiments are performed under various settings, corroborating with the pronounced theoretical results. As a real application example, we analyse a set of taxi data to study the traffic flow between Times Squares and its neighbouring areas.

stat.ME

Main Effect Factor Models in High-Dimensional Matrix Time Series: Identification and Sparsity

We propose a general identification framework for main effect factor models for matrix-valued time series. The classical sum-to-zero restriction on the row and column main effects is replaced by a broad class of shift-varying functions, which includes weighted averages, quantiles, and reference-unit anchors as special cases. We prove that the shift-varying property is necessary and sufficient for parameter identification, and we derive closed-form estimators under minimal assumptions. Asymptotic convergence rates of all estimated components are derived under weak, heterogeneous factor strengths. Building on this flexible identification, we address sparsity in the main effects by selecting a minimum-based shift-varying function that sets the smallest main effect to zero, and we introduce a doubly adaptive fused Lasso estimator that consistently recovers the true sparse and dense blocks. A modified Mallows's statistic is developed for tuning parameter selection, and the entire procedure is computationally practical via an equivalent generalized lasso formulation. Simulation experiments are performed under a variety of settings, showing our proposed methods work well. Two real applications on employment growth and economic indices are demonstrated to illustrate the empirical value of our approach.

math.ST

Inference on Dynamic Spatial Autoregressive Models with Change Point Detection

We analyze a varying-coefficient dynamic spatial autoregressive model with spatial fixed effects. One salient feature of the model is the incorporation of multiple spatial weight matrices through their linear combinations with varying coefficients, which help solve the problem of choosing the most ``correct'' one for applied econometricians who often face the availability of multiple expert spatial weight matrices. We estimate and make inferences on the model coefficients and coefficients in basis expansions of the varying coefficients through penalized estimations, establishing the oracle properties of the estimators and the consistency of the overall estimated spatial weight matrix, which can be time-dependent. We further consider two applications of our model in change point detections in dynamic spatial autoregressive models, providing theoretical justifications in consistent change point locations estimation and practical implementations. Simulation experiments demonstrate the performance of our proposed methodology, and real data analyses are also carried out.

stat.ME

On Testing Kronecker Product Structure in Tensor Factor Models

We propose a test for testing the Kronecker product structure of a factor loading matrix implied by a tensor factor model with Tucker decomposition in the common component. Through defining a Kronecker product structure set, we define if a tensor time series response $\{\mathcal{Y}_t\}$ has a Kronecker product structure, equivalent to the ability to decompose $\{\mathcal{Y}_t\}$ according to a tensor factor model. Our test is built on analysing and comparing the residuals from fitting a full tensor factor model, and the residuals from fitting a (tensor) factor model on a reshaped version of the data. In the most extreme case, the reshaping is the vectorisation of the tensor data, and the factor loading matrix in such a case can be general if there is no Kronecker product structure present. Theoretical results are developed through asymptotic normality results on estimated residuals. Numerical experiments suggest that the size of the tests gets closer to the pre-set nominal value as the sample size or the order of the tensor gets larger, while the power increases with mode dimensions and the number of combined modes. We demonstrate out tests through a NYC taxi traffic data and a Fama-French matrix portfolio of returns.

math.ST

Tensor Time Series Imputation through Tensor Factor Modelling

We propose tensor time series imputation when the missing pattern in the tensor data can be general, as long as any two data positions along a tensor fibre are both observed for enough time points. The method is based on a tensor time series factor model with Tucker decomposition of the common component. One distinguished feature of the tensor time series factor model used is that there can be weak factors in the factor loadings matrix for each mode. This reflects reality better when real data can have weak factors which drive only groups of observed variables, for instance, a sector factor in financial market driving only stocks in a particular sector. Using the data with missing entries, asymptotic normality is derived for rows of estimated factor loadings, while consistent covariance matrix estimation enables us to carry out inferences. As a first in the literature, we also propose a ratio-based estimator for the rank of the core tensor under general missing patterns. Rates of convergence are spelt out for the imputations from the estimated tensor factor models. Simulation results show that our imputation procedure works well, with asymptotic normality and corresponding inferences also demonstrated. Re-imputation performances are also gauged when we demonstrate that using slightly larger rank then estimated gives superior re-imputation performances. A Fama-French portfolio example with matrix returns and an OECD data example with matrix of Economic indicators are presented and analyzed, showing the efficacy of our imputation approach compared to direct vector imputation.

math.ST

Matrix-valued Factor Model with Time-varying Main Effects

We introduce the matrix-valued time-varying Main Effects Factor Model (MEFM). MEFM is a generalization to the traditional matrix-valued factor model (FM). We give rigorous definitions of MEFM and its identifications, and propose estimators for the time-varying grand mean, row and column main effects, and the row and column factor loading matrices for the common component. Rates of convergence for different estimators are spelt out, with asymptotic normality shown. The core rank estimator for the common component is also proposed, with consistency of the estimators presented. We propose a test for testing if FM is sufficient against the alternative that MEFM is necessary, and demonstrate the power of such a test in various simulation settings. We also demonstrate numerically the accuracy of our estimators in extended simulation experiments. A set of NYC Taxi traffic data is analysed and our test suggests that MEFM is indeed necessary for analysing the data against a traditional FM.

math.ST