SearcharxivSearch

arXiv subjects

Song Xi Chen

Publications and source records attributed to Song Xi Chen.

At least 19 recordsLinked to original sources

A Partially Functional Dynamic Structural Equation Model for Multi-Resolution Environmental Data

Understanding the complex relationships between atmospheric pollutant emissions and their multifaceted determinants presents a dual challenge: driving factors operate at fundamentally different temporal resolutions, from continuously monitored meteorological variables to annually reported socio-economic indicators, and their interconnections evolve dynamically over time. To address these challenges, we propose a Partially Functional Dynamic Structural Equation Model (PFDSEM) that coherently integrates functional covariates (e.g., high-frequency meteorological data) and scalar predictors (e.g., economic and demographic indicators) within a unified dynamic structural framework. The model captures non-stationary temporal dependencies and inter-variable correlations via a Conditional Autoregressive (CAR) structure combined with a Linear Model of Coregionalization (LMC), while functional covariates are incorporated through basis expansion with Bayesian P-spline smoothing. Bayesian inference via Markov Chain Monte Carlo provides full uncertainty quantification. Comprehensive simulation studies confirm accurate parameter recovery and robustness to prior specifications under diverse conditions. Applying the PFDSEM to pollutant emissions data from 30 Chinese provinces (2015--2020), we identify temporally dynamic and province-specific associations between ten categories of socio-environmental factors and ten major air pollutants, including CO$_2$. The results reveal substantial cross-province heterogeneity in the strength and direction of these associations and pronounced seasonal patterns in meteorological effects, offering a quantitative evidence base for designing temporally adaptive and regionally tailored environmental policies.

math.ST

Localization Estimator for High Dimensional Tensor Covariance Matrices

This paper considers covariance matrix estimation of tensor data under high dimensionality. A multi-bandable covariance class is established to accommodate the need for complex covariance structures of multi-layer lattices and general covariance decay patterns. We propose a high dimensional covariance localization estimator for tensor data, which regulates the sample covariance matrix through a localization function. The statistical properties of the proposed estimator are studied by deriving the minimax rates of convergence under the spectral and the Frobenius norms. Numerical experiments and real data analysis on ocean eddy data are carried out to illustrate the utility of the proposed method in practice.

stat.ME

Errors-in-variables regression for dependent data with estimated error covariance matrix: To prewhiten or not?

We consider statistical inference for errors-in-variables regression models with dependent observations under the high dimensionality of the error covariance matrix. It is tempting to prewhiten the model and data that had led to efficient weighted least squares estimation in the presence of the measurement errors, as being practised in the optimal fingerprinting approach in climate change studies. However, it is unclear to what extent the prewhitened estimator can improve the estimation efficiency of the unprewhitened estimator for errors-in-variables regression. We compare the prewhitening and unprewhitening estimators in terms of their estimation efficiency and computational cost. It shows that while the prewhitening operation does not necessarily improve the estimation efficiency of its unprewhitening counterpart, it demands more on the ensemble size needed in the error-covariance matrix estimation to ensure the asymptotic normality, and hence it would requires much more computationally resource.

stat.AP

High Dimensional Ensemble Kalman Filter

The Ensemble Kalman Filter (EnKF), as a fundamental data assimilation approach, has been widely used in many fields of the sciences and engineering. When the state variable is of high dimensional accompanied with high resolution observations of physical models, some key theoretical aspects of the EnKF are open for investigation. This paper proposes several high dimensional EnKF (HD-EnKF) methods equipped with consistent estimators for the important forecast error covariance and Kalman Gain matrices. It then studies the theoretical properties of the EnKF under both fixed and high dimensional state variables, which provides one-step and multiple-step mean square errors of the analysis states to the underlying oracle states offered by the Kalman Filter and gives the much needed insight to the roles played by the forecast error covariance on the accuracy of the EnKF. The accuracy of the data assimilation under the misspecified physical model is also considered. Numerical studies on the Lorenz-96 and the Shallow Water Equation models illustrate that the proposed HD-EnKF algorithms outperform the standard EnKF and widely used inflation methods.

stat.ME

Partially Functional Dynamic Backdoor Diffusion-based Causal Model

Causal inference in spatio-temporal settings is critically hindered by unmeasured confounders with complex spatio-temporal dynamics and the prevalence of multi-resolution data. While diffusion models present a promising avenue for estimating structural causal models, existing approaches are limited by assumptions of causal sufficiency or static confounding, failing to capture the region-specific, temporally dependent nature of real-world latent variables or to directly handle functional variables. We bridge this gap by introducing the Partially Functional Dynamic Backdoor Diffusion-based Causal Model (PFD-BDCM), a unified generative framework designed to simultaneously tackle causal inference with dynamic confounding and functional data. Our approach formalizes a novel structural causal model that captures spatio-temporal dependencies in latent confounders through conditional autoregressive processes, represents functional variables via basis expansion coefficients treated as standard graph nodes, and integrates valid backdoor adjustment into a diffusion-based generative process. We provide theoretical guarantees on the preservation of causal effects under basis expansion and derive error bounds for counterfactual estimates. Experiments on synthetic data and a real-world air pollution case study demonstrate that PFD-BDCM outperforms existing methods across observational, interventional, and counterfactual queries. This work provides a rigorous and practical tool for robust causal inference in complex spatio-temporal systems characterized by non-stationarity and multi-resolution data.

stat.ML

Likelihood Matching for Diffusion Models

We propose a Likelihood Matching approach for training diffusion models by first establishing an equivalence between the likelihood of the target data distribution and a likelihood along the sample path of the reverse diffusion. To efficiently compute the reverse sample likelihood, a quasi-likelihood is considered to approximate each reverse transition density by a Gaussian distribution with matched conditional mean and covariance, respectively. The score and Hessian functions for the diffusion generation are estimated by maximizing the quasi-likelihood, ensuring a consistent matching of both the first two transitional moments between every two time points. A stochastic sampler is introduced to facilitate computation that leverages both the estimated score and Hessian information. We establish consistency of the quasi-maximum likelihood estimation, and provide non-asymptotic convergence guarantees for the proposed sampler, quantifying the rates of the approximation errors due to the score and Hessian estimation, dimensionality, and the number of diffusion steps. Empirical and simulation evaluations demonstrate the effectiveness of the proposed Likelihood Matching and validate the theoretical results.

stat.ML

Identifying Heterogeneity in Distributed Learning

We study methods for identifying heterogeneous parameter components in distributed M-estimation with minimal data transmission. One is based on a re-normalized Wald test, which is shown to be consistent as long as the number of distributed data blocks $K$ is of a smaller order of the minimum block sample size and the level of heterogeneity is dense. The second one is an extreme contrast test (ECT) based on the difference between the largest and smallest component-wise estimated parameters among data blocks. By introducing a sample splitting procedure, the ECT can avoid the bias accumulation arising from the M-estimation procedures, and exhibits consistency for $K$ being much larger than the sample size while the heterogeneity is sparse. The ECT procedure is easy to operate and communication-efficient. A combination of the Wald and the extreme contrast tests is formulated to attain more robust power under varying levels of sparsity of the heterogeneity. We also conduct intensive numerical experiments to compare the family-wise error rate (FWER) and the power of the proposed methods. Additionally, we conduct a case study to present the implementation and validity of the proposed methods.

stat.ML

Glider Path Design and Control for Reconstructing Three-Dimensional Structures of Oceanic Mesoscale Eddies

Underwater gliders offer effective means in oceanic surveys with a major task in reconstructing the three-dimensional hydrographic field of a mesoscale eddy. This paper considers three key issues in the hydrographic reconstruction of mesoscale eddies with the sampled data from the underwater gliders. It first proposes using the Thin Plate Spline (TPS) as the interpolation method for the reconstruction with a blocking scheme to speed up the computation. It then formulates a procedure for selecting glider path design that minimizes the reconstruction errors among a set of pathway formations. Finally we provide a glider path control procedure to guide the glider to follow to designed pathways as much as possible in the presence of ocean current. A set of optimization algorithms are experimented and several with robust glider control performance on a simulated eddy are identified.

stat.AP

Concentration Inequalities for Statistical Inference

This paper gives a review of concentration inequalities which are widely employed in non-asymptotical analyses of mathematical statistics in a wide range of settings, from distribution-free to distribution-dependent, from sub-Gaussian to sub-exponential, sub-Gamma, and sub-Weibull random variables, and from the mean to the maximum concentration. This review provides results in these settings with some fresh new results. Given the increasing popularity of high-dimensional data and inference, results in the context of high-dimensional linear and Poisson regressions are also provided. We aim to illustrate the concentration inequalities with known constants and to improve existing bounds with sharper constants.

math.ST

Statistical Inference for Four-Regime Segmented Regression Models

Segmented regression models offer model flexibility and interpretability as compared to the global parametric and the nonparametric models, and yet are challenging in both estimation and inference. We consider a four-regime segmented model for temporally dependent data with segmenting boundaries depending on multivariate covariates with non-diminishing boundary effects. A mixed integer quadratic programming algorithm is formulated to facilitate the least square estimation of the regression and the boundary parameters. The rates of convergence and the asymptotic distributions of the least square estimators are obtained for the regression and the boundary coefficients, respectively. We propose a smoothed regression bootstrap to facilitate inference on the parameters and a model selection procedure to select the most suitable model within the model class with at most four segments. Numerical simulations and a case study on air pollution in Beijing are conducted to demonstrate the proposed approach, which shows that the segmented models with three or four regimes are suitable for the modeling of the meteorological effects on the PM2.5 concentration.

stat.ME

Transfer Learning with General Estimating Equations

We consider statistical inference for parameters defined by general estimating equations under the covariate shift transfer learning. Different from the commonly used density ratio weighting approach, we undertake a set of formulations to make the statistical inference semiparametric efficient with simple inference. It starts with re-constructing the estimation equations to make them Neyman orthogonal, which facilitates more robustness against errors in the estimation of two key nuisance functions, the density ratio and the conditional mean of the moment function. We present a divergence-based method to estimate the density ratio function, which is amenable to machine learning algorithms including the deep learning. To address the challenge that the conditional mean is parametric-dependent, we adopt a nonparametric multiple-imputation strategy that avoids regression at all possible parameter values. With the estimated nuisance functions and the orthogonal estimation equation, the inference for the target parameter is formulated via the empirical likelihood without sample splittings. We show that the proposed estimator attains the semiparametric efficiency bound, and the inference can be conducted with the Wilks' theorem. The proposed method is further evaluated by simulations and an empirical study on a transfer learning inference for ground-level ozone pollution

stat.ME

A Review on the Optimal Fingerprinting Approach in Climate Change Studies

We provide a review on the "optimal fingerprinting" approach as summarized in Allen and Tett (1999) from a point view of statistical inference in light of the recent criticism of McKitrick (2021). Our review finds that the "optimal fingerprinting" approach would survive much of McKitrick (2021)'s criticism under two conditions: (i) the null simulation of the climate model is independent of the physical observations and (ii) the null simulation provides consistent estimation of the residual covariance matrix of the physical observations, both depend on the conduction and the quality of the climate models. If the latter condition fails, the estimator would be still unbiased and consistent under routine conditions, but losing the "optimal" aspect of the approach. The residual consistency test suggested by Allen and Tett (1999) is valid for checking the agreement between the residual covariances of the null simulation and the physical observations. We further outline the connection between the "optimal fingerprinting" approach and the Feasible Generalized Least Square.

physics.data-an

Matrix Completion under Low-Rank Missing Mechanism

Matrix completion is a modern missing data problem where both the missing structure and the underlying parameter are high dimensional. Although missing structure is a key component to any missing data problems, existing matrix completion methods often assume a simple uniform missing mechanism. In this work, we study matrix completion from corrupted data under a novel low-rank missing mechanism. The probability matrix of observation is estimated via a high dimensional low-rank matrix estimation procedure, and further used to complete the target matrix via inverse probabilities weighting. Due to both high dimensional and extreme (i.e., very small) nature of the true probability matrix, the effect of inverse probability weighting requires careful study. We derive optimal asymptotic convergence rates of the proposed estimators for both the observation probabilities and the target matrix.

stat.ML

High-dimensional empirical likelihood inference

High-dimensional statistical inference with general estimating equations are challenging and remain less explored. In this paper, we study two problems in the area: confidence set estimation for multiple components of the model parameters, and model specifications test. For the first one, we propose to construct a new set of estimating equations such that the impact from estimating the high-dimensional nuisance parameters becomes asymptotically negligible. The new construction enables us to estimate a valid confidence region by empirical likelihood ratio. For the second one, we propose a test statistic as the maximum of the marginal empirical likelihood ratios to quantify data evidence against the model specification. Our theory establishes the validity of the proposed empirical likelihood approaches, accommodating over-identification and exponentially growing data dimensionality. The numerical studies demonstrate promising performance and potential practical benefits of the new methods.

stat.ME

Multi-level Thresholding Test for High Dimensional Covariance Matrices

We consider testing the equality of two high-dimensional covariance matrices by carrying out a multi-level thresholding procedure, which is designed to detect sparse and faint differences between the covariances. A novel U-statistic composition is developed to establish the asymptotic distribution of the thresholding statistics in conjunction with the matrix blocking and the coupling techniques. We propose a multi-thresholding test that is shown to be powerful in detecting sparse and weak differences between two covariance matrices. The test is shown to have attractive detection boundary and to attain the optimal minimax rate in the signal strength under different regimes of high dimensionality and the sparsity of the signal. Simulation studies are conducted to demonstrate the utility of the proposed test.

math.ST

Analyzing China's Consumer Price Index Comparatively with that of United States

This paper provides a thorough analysis on the dynamic structures and predictability of China's Consumer Price Index (CPI-CN), with a comparison to those of the United States. Despite the differences in the two leading economies, both series can be well modeled by a class of Seasonal Autoregressive Integrated Moving Average Model with Covariates (S-ARIMAX). The CPI-CN series possess regular patterns of dynamics with stable annual cycles and strong Spring Festival effects, with fitting and forecasting errors largely comparable to their US counterparts. Finally, for the CPI-CN, the diffusion index (DI) approach offers improved predictions than the S-ARIMAX models.

econ.EM

Distributed Statistical Inference for Massive Data

This paper considers distributed statistical inference for general symmetric statistics %that encompasses the U-statistics and the M-estimators in the context of massive data where the data can be stored at multiple platforms in different locations. In order to facilitate effective computation and to avoid expensive communication among different platforms, we formulate distributed statistics which can be conducted over smaller data blocks. The statistical properties of the distributed statistics are investigated in terms of the mean square error of estimation and asymptotic distributions with respect to the number of data blocks. In addition, we propose two distributed bootstrap algorithms which are computationally effective and are able to capture the underlying distribution of the distributed statistics. Numerical simulation and real data applications of the proposed approaches are provided to demonstrate the empirical performance.

math.ST

High dimensional generalized empirical likelihood for moment restrictions with dependent data

This paper considers the maximum generalized empirical likelihood (GEL) estimation and inference on parameters identified by high dimensional moment restrictions with weakly dependent data when the dimensions of the moment restrictions and the parameters diverge along with the sample size. The consistency with rates and the asymptotic normality of the GEL estimator are obtained by properly restricting the growth rates of the dimensions of the parameters and the moment restrictions, as well as the degree of data dependence. It is shown that even in the high dimensional time series setting, the GEL ratio can still behave like a chi-square random variable asymptotically. A consistent test for the over-identification is proposed. A penalized GEL method is also provided for estimation under sparsity setting.

math.ST