SearcharxivSearch

arXiv subjects

Xiaofeng Shao

Publications and source records attributed to Xiaofeng Shao.

At least 19 recordsLinked to original sources

Online Change-Point Monitoring for Object-valued Time Series

We develop closed- and open-end procedures for monitoring changes in the marginal distribution of object-valued time series. The method combines two distance-based Hilbert-space embeddings, a monitoring-time-dependent projection, and self-normalization. It is computable entirely from pairwise distances, does not require long-run variance estimation, and admits exact recursive updates. Under weak temporal dependence, the closed-end null limit is a pivotal Brownian functional. For monitoring over an unbounded horizon, we introduce a growing monitoring-time weight, establish a maximal inequality and uniform remote-tail control, and derive an open-end pivotal limit together with an equivalent fixed-interval representation for critical-value simulation. Both procedures are consistent against fixed marginal changes and retain the $n^{-1/2}$ and $n^{-1/4}$ local detection boundaries associated with the linear and quadratic projection signals. Simulations with distribution-valued time series illustrate favorable numerical properties, and an application to monthly stock-return distributions illustrates the usefulness of our procedure. Further simulations with graph-valued time series and an application to spatial point processes of seismic activity are reported in the supplementary material.

math.ST

Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling

Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional central limit theorem whose covariance retains the multiplicative component induced by the random TD matrix and the stationary iterate error. We then derive a joint functional limit for parallel Richardson--Romberg (RR) recursions driven by the same trajectory. A Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified state-value contrasts without estimating the long-run covariance or selecting a bandwidth or batch length. For such a contrast, the procedure admits a one-pass implementation whose memory does not grow with the trajectory length. At a fixed stepsize, the inferential center is the RR stationary target. We also study horizon-indexed designs in which the stepsize remains constant within each run and decreases across longer horizons. Under an explicit RR-dependent rate window, the residual RR target shift, multiplicative remainder, and initialization effect are negligible at the root-$n$ scale, yielding inference for the projected Bellman solution. Experiments on FrozenLake and Garnet illustrate stationary-target coverage, RR target correction, and the finite-sample behavior of the horizon-indexed design.

stat.ML

Testing Equality of Conditional Distributions via Generative Models

We study the problem of testing whether two conditional distributions are equal using generative models. The proposed method learns a conditional generator from each sample and uses it to create responses at covariate values observed in the other sample, allowing generated and observed responses to be compared directly. By aligning covariates through cross-generation, the approach avoids conditional density-ratio estimation and local smoothing over high-dimensional covariates. The population version of this construction yields a conditional discrepancy that characterizes equality of the two conditional distributions under suitable overlap conditions, while the sample version leads to a test statistic defined as the supremum of an RKHS-indexed empirical process with multiplier bootstrap calibration. A computationally efficient algorithm for evaluating the statistic and its bootstrap analogue is developed based on alternating maximization and the kernel trick. Theoretically, we derive the limiting distribution of the test statistic under both the null and alternative hypotheses, prove bootstrap validity and consistency of the resulting test, and show that the proposed procedure attains a double-robustness property with respect to conditional generator estimation errors. Simulations and real data applications suggest that the proposed method performs well for multivariate responses and high-dimensional covariates.

stat.ME

Hypothesis Testing for a Functional Parameter via Self-normalization

Testing simple or composite hypothesis on a functional parameter has attracted considerable attention in time series analysis. To accommodate for the unknown temporal dependence, classical nonparametric approaches such as block bootstrapping and subsampling all involve a bandwidth parameter, the choice of which can substantially affect the finite sample performance. The self normalization (SN) method is tuning parameter free when applied to the inference of a finite-dimensional parameter but its applicability to a functional parameter is unknown. In this paper, we propose a sample splitting based approach to generalize the SN method to hypothesis testing of a functional parameter. Our SS-SN (sample splitting plus self-normalization) idea is broadly applicable to many testing problems for functional parameters, including testing for simple/composite hypothesis on marginal cumulative distribution function, testing for time-reversibility and testing for a change point on the spectral distribution of a multivariate time series. Specifically, we derive the pivotal limiting distributions of our SS-SN test statistics under the null for both simple and composite null hypothesis, and derive the limiting power function under the local alternatives. Numerical simulations show that our new tests tend to yield accurate size with competitive power performance as compared to many existing ones.

stat.ME

Change-Point Detection for Object-valued Time Series

This article is concerned with change point detection for object-valued data that reside in a metric space, which has attracted some recent interests in statistics and econometrics literature. The existing methods either focus on independent data or can only detect change in the Fréchet mean or variance. In this paper, we propose a self-normalization (SN, hereafter) based statistic for detecting a shift in the marginal distribution of object-valued time series. Our test is universally applicable to a wide range of object-valued data, such as distributional and network data, and can accommodate weak serial dependence. In addition the proposed test statistic is almost tuning parameter free, has pivotal limiting null distribution and only uses the pairwise distances. When combined with the Wild Binary Segmentation algorithm (WBS, hereafter), our statistic can be used to estimate the number and locations of multiple change points. Asymptotic results for our SN based statistic are derived under both null and local alternatives in the single change point setting. For the first time, the WBS estimation consistency is shown for a broad class of object-valued time series and in a nonparametric setting, which requires new non-standard theoretical arguments. Extensive numerical experiments and real data analysis are conducted to illustrate the effectiveness and broad applicability of our proposed method.

stat.ME

Another Look at Bandwidth-free Inference: a Sample Splitting Approach

The bandwidth-free tests/inferences for a multi-dimensional parameter have attracted considerable attention in econometrics and statistics literature. These tests can be conveniently implemented due to their tuning-parameter free nature and possess more accurate size as compared to the traditional HAC-based approaches, where consistent long run variance estimation was involved. However, when sample size is small/medium, these bandwidth-free tests exhibit large size distortion when both the dimension of the parameter and the magnitude of temporal dependence are moderate, making them unreliable to use in practice. In this paper, we propose a sample splitting based approach to reduce the dimension of the parameter to one for the subsequent bandwidth-free inference. Our SS-SN (sample splitting plus self-normalization) idea is broadly applicable to many testing problems for time series, including mean testing, testing for zero autocorrelation, linear hypotheses testing in a time series regression model and testing for a change point in multivariate mean. Specifically, we propose $L_{\infty}$-type and $L_2$-type SS-SN test statistics and derive their limiting distributions under both the null and alternatives and show their effectiveness in alleviating size distortion via simulations. As an important theoretical contribution, we obtain the limiting distributions for both SS-SN test statistics in the multivariate mean testing problem when the dimension is allowed to diverge as sample size grows to infinity. In addition we show the asymptotic independence of $L_{\infty}$-type and $L_2$-type SS-SN test statistics under the null in the growing dimensional setting.

stat.ME

Generalized Spectral Testing with Sample Splitting

Residual-based goodness-of-fit tests for parametric time-series models are often complicated by parameter-estimation effects, which can alter the limiting behavior of diagnostic statistics. We propose a sample-splitting generalized spectral test (in the spirit of Escanciano(2006)) for assessing conditional mean specification in linear and nonlinear time-series models. The procedure estimates the model parameter on a fitting subsample and constructs a generalized spectral Cramer-von Mises statistic from residuals computed on a checking/testing subsample. The statistic aggregates pairwise conditional mean restrictions over all lags and is therefore bandwidth-free and free of truncation-lag selection. Under mild regularity conditions and a score-alignment condition, the residual-based process has the same limiting null distribution as the infeasible oracle process based on the true errors. Although the resulting limiting law is still non-pivotal, it can be consistently approximated by a simple multiplier bootstrap that does not require generating bootstrap time series or re-estimating parameters. Such an oracle-equivalence property is in sharp contrast to the original full-sample test, for which parameter estimation contributes an additional first-order term to the limiting process, and requires re-estimating parameters in each bootstrapped sample. We further establish consistency of the proposed test against fixed alternatives and nontrivial power against local alternatives. Extensive simulations and real data analyses show that the proposed test controls size well, has comparable power, and delivers substantial computational savings in models where repeated estimation is costly.

econ.EM

Resampling-free Inference for Time Series via RKHS Embedding

In this article, we study nonparametric inference problems in the context of multivariate or functional time series, including testing for goodness-of-fit, the presence of a change point in the marginal distribution, and the independence of two time series, among others. Most methodologies available in the existing literature address these problems by employing a bandwidth-dependent bootstrap or subsampling approach, which can be computationally expensive and/or sensitive to the choice of bandwidth. To address these limitations, we propose a novel class of kernel-based tests by embedding the data into a reproducing kernel Hilbert space, and construct test statistics using sample splitting, projection, and self-normalization (SN) techniques. Through a new conditioning technique, we demonstrate that our test statistics have pivotal limiting null distributions under strong mixing and mild moment assumptions. We also analyze the limiting power of our tests under local alternatives. Finally, we showcase the superior size accuracy and computational efficiency of our methods as compared to some existing ones.

stat.ME

Diagnostic Checking for Wasserstein Autoregression

Wasserstein autoregression provides a robust framework for modeling serial dependence among probability distributions, with wide-ranging applications in economics, finance, and climate science. In this paper, we develop portmanteau-type diagnostic tests for assessing the adequacy of Wasserstein autoregressive models. By defining autocorrelation functions for model errors and residuals in the Wasserstein space, we construct two related tests: one analogous to the classical McLeod type test, and the other based on the sample-splitting approach of Davis and Fernandes(2025). We establish that, under mild regularity conditions, the corresponding test statistics converge in distribution to chi-square limits. Simulation studies and empirical applications demonstrate that the proposed tests effectively detect model mis-specification, offering a principled and reliable diagnostic tool for distributional time series analysis.

stat.ME

Statistical inference for high-dimensional spectral density matrix

The spectral density matrix is a fundamental object of interest in time series analysis, and it encodes both contemporary and dynamic linear relationships between component processes of the multivariate system. In this paper we develop novel inference procedures for the spectral density matrix in the high-dimensional setting. Specifically, we introduce a new global testing procedure to test the nullity of the cross-spectral density for a given set of frequencies and across pairs of component indices. For the first time, both Gaussian approximation and parametric bootstrap methodologies are employed to conduct inference for a high-dimensional parameter formulated in the frequency domain, and new technical tools are developed to provide asymptotic guarantees of the size accuracy and power for global testing. We further propose a multiple testing procedure for simultaneously testing the nullity of the cross-spectral density at a given set of frequencies. The method is shown to control the false discovery rate. Both numerical simulations and a real data illustration demonstrate the usefulness of the proposed testing methods.

math.ST

Online Generalized Method of Moments for Time Series

Online learning has gained popularity in recent years due to the urgent need to analyse large-scale streaming data, which can be collected in perpetuity and serially dependent. This motivates us to develop the online generalized method of moments (OGMM), an explicitly updated estimation and inference framework in the time series setting. The OGMM inherits many properties of offline GMM, such as its broad applicability to many problems in econometrics and statistics, natural accommodation for over-identification, and achievement of semiparametric efficiency under temporal dependence. As an online method, the key gain relative to offline GMM is the vast improvement in time complexity and memory requirement. Building on the OGMM framework, we propose improved versions of online Sargan--Hansen and structural stability tests following recent work in econometrics and statistics. Through Monte Carlo simulations, we observe encouraging finite-sample performance in online instrumental variables regression, online over-identifying restrictions test, online quantile regression, and online anomaly detection. Interesting applications of OGMM to stochastic volatility modelling and inertial sensor calibration are presented to demonstrate the effectiveness of OGMM.

stat.ME

Testing Conditional Mean Independence Using Generative Neural Networks

Conditional mean independence (CMI) testing is crucial for statistical tasks including model determination and variable importance evaluation. In this work, we introduce a novel population CMI measure and a bootstrap-based testing procedure that utilizes deep generative neural networks to estimate the conditional mean functions involved in the population measure. The test statistic is thoughtfully constructed to ensure that even slowly decaying nonparametric estimation errors do not affect the asymptotic accuracy of the test. Our approach demonstrates strong empirical performance in scenarios with high-dimensional covariates and response variable, can handle multivariate responses, and maintains nontrivial power against local alternatives outside an $n^{-1/2}$ neighborhood of the null hypothesis. We also use numerical simulations and real-world imaging data applications to highlight the efficacy and versatility of our testing procedure.

stat.ML

Doubly Robust Conditional Independence Testing with Generative Neural Networks

This article addresses the problem of testing the conditional independence of two generic random vectors $X$ and $Y$ given a third random vector $Z$, which plays an important role in statistical and machine learning applications. We propose a new non-parametric testing procedure that avoids explicitly estimating any conditional distributions but instead requires sampling from the two marginal conditional distributions of $X$ given $Z$ and $Y$ given $Z$. We further propose using a generative neural network (GNN) framework to sample from these approximated marginal conditional distributions, which tends to mitigate the curse of dimensionality due to its adaptivity to any low-dimensional structures and smoothness underlying the data. Theoretically, our test statistic is shown to enjoy a doubly robust property against GNN approximation errors, meaning that the test statistic retains all desirable properties of the oracle test statistic utilizing the true marginal conditional distributions, as long as the product of the two approximation errors decays to zero faster than the parametric rate. Asymptotic properties of our statistic and the consistency of a bootstrap procedure are derived under both null and local alternatives. Extensive numerical experiments and real data analysis illustrate the effectiveness and broad applicability of our proposed test.

stat.ME

SNSeg: An R Package for Time Series Segmentation via Self-Normalization

Time series segmentation aims to identify potential change-points in a sequence of temporally dependent data, so that the original sequence can be partitioned into several homogeneous subsequences. It is useful for modeling and predicting non-stationary time series and is widely applied in natural and social sciences. Existing segmentation methods primarily focus on only one type of parameter changes such as mean and variance, and they typically depend on laborious tuning or smoothing parameters, which can be challenging to choose in practice. The self-normalization based change-point estimation framework SNCP by Zhao et al. (2022), however, offers users more flexibility and convenience as it allows for change-point estimation of different types of parameters (e.g. mean, variance, quantile and autocovariance) in a unified fashion, and requires effortless tuning. In this paper, the R package SNSeg is introduced to implement SNCP for segmentation of univariate and multivariate time series. An extension of SNCP, named SNHD, is also designed and implemented for change-point estimation in the mean vector of high-dimensional time series. The estimated changepoints as well as segmented time series are available with graphical tools. Detailed examples of SNSeg are given in simulations of multivariate autoregressive processes with change-points.

stat.CO

Dimension-agnostic Change Point Detection

Change point testing for high-dimensional data has attracted a lot of attention in statistics and machine learning owing to the emergence of high-dimensional data with structural breaks from many fields. In practice, when the dimension is less than the sample size but is not small, it is often unclear whether a method that is tailored to high-dimensional data or simply a classical method that is developed and justified for low-dimensional data is preferred. In addition, the methods designed for low-dimensional data may not work well in the high-dimensional environment and vice versa. In this paper, we propose a dimension-agnostic testing procedure targeting a single change point in the mean of a multivariate time series. Specifically, we can show that the limiting null distribution for our test statistic is the same regardless of the dimensionality and the magnitude of cross-sectional dependence. The power analysis is also conducted to understand the large sample behavior of the proposed test. Through Monte Carlo simulations and a real data illustration, we demonstrate that the finite sample results strongly corroborate the theory and suggest that the proposed test can be used as a benchmark for change-point detection of time series of low, medium, and high dimensions.

stat.ME

Change-point Inference for High-dimensional Heteroscedastic Data

We propose a bootstrap-based test to detect a mean shift in a sequence of high-dimensional observations with unknown time-varying heteroscedasticity. The proposed test builds on the U-statistic based approach in Wang et al. (2022), targets a dense alternative, and adopts a wild bootstrap procedure to generate critical values. The bootstrap-based test is free of tuning parameters and is capable of accommodating unconditional time varying heteroscedasticity in the high-dimensional observations, as demonstrated in our theory and simulations. Theoretically, we justify the bootstrap consistency by using the recently proposed unconditional approach in Bucher and Kojadinovic (2019). Extensions to testing for multiple change-points and estimation using wild binary segmentation are also presented. Numerical simulations demonstrate the robustness of the proposed testing and estimation procedures with respect to different kinds of time-varying heteroscedasticity.

stat.ME

Two Sample Testing in High Dimension via Maximum Mean Discrepancy

Maximum Mean Discrepancy (MMD) has been widely used in the areas of machine learning and statistics to quantify the distance between two distributions in the $p$-dimensional Euclidean space. The asymptotic property of the sample MMD has been well studied when the dimension $p$ is fixed using the theory of U-statistic. As motivated by the frequent use of MMD test for data of moderate/high dimension, we propose to investigate the behavior of the sample MMD in a high-dimensional environment and develop a new studentized test statistic. Specifically, we obtain the central limit theorems for the studentized sample MMD as both the dimension $p$ and sample sizes $n,m$ diverge to infinity. Our results hold for a wide range of kernels, including popular Gaussian and Laplacian kernels, and also cover energy distance as a special case. We also derive the explicit rate of convergence under mild assumptions and our results suggest that the accuracy of normal approximation can improve with dimensionality. Additionally, we provide a general theory on the power analysis under the alternative hypothesis and show that our proposed test can detect difference between two distributions in the moderately high dimensional regime. Numerical simulations demonstrate the effectiveness of our proposed test statistic and normal approximation.

math.ST

Testing Serial Independence of Object-Valued Time Series

We propose a novel method for testing serial independence of object-valued time series in metric spaces, which is more general than Euclidean or Hilbert spaces. The proposed method is fully nonparametric, free of tuning parameters, and can capture all nonlinear pairwise dependence. The key concept used in this paper is the distance covariance in metric spaces, which is extended to auto distance covariance for object-valued time series. Furthermore, we propose a generalized spectral density function to account for pairwise dependence at all lags and construct a Cramer-von Mises type test statistic. New theoretical arguments are developed to establish the asymptotic behavior of the test statistic. A wild bootstrap is also introduced to obtain the critical values of the non-pivotal limiting null distribution. Extensive numerical simulations and two real data applications are conducted to illustrate the effectiveness and versatility of our proposed method.

stat.ME