SearcharxivSearch

arXiv subjects

Xiaoping Shi

Publications and source records attributed to Xiaoping Shi.

At least 19 recordsLinked to original sources

DROP: Distributionally Robust Optimization for Multi-task Learning in Graphical Models

Gaussian Graphical Models (GGMs) are widely used to infer conditional dependence structures in high-dimensional data. However, standard precision matrix estimators are highly sensitive to data contamination, such as extreme outliers and heavy-tailed noise. In this paper, we propose DROP (Distributionally Robust Optimization), a robust estimation method formulated within a multi-task nodewise regression framework. The proposed estimator enforces structural sparsity while resisting the influence of corrupted observations. Theoretically, we establish error bounds for the DROP estimator under general contamination. Through extensive high-dimensional simulations, we demonstrate that DROP consistently controls the rate of false positive edges and outperforms conventional non-robust estimators when data deviate from standard Gaussian assumptions. Furthermore, in a functional MRI (fMRI) application, DROP maintains a stable graph structure and preserves network modularity even when subjected to severe data perturbations, whereas competing methods yield excessively dense networks. To facilitate reproducible research, the DROP R package will be made publicly available on GitHub.

stat.AP

Consistent and powerful CUSUM change-point test for panel data with changes in variance

This paper investigates change-point of variance in panel data models with time series of $\alpha$-mixing. Based on the cumulative sum (CUSUM) method and the individual differences, we construct a CUSUM test for panel data models to detect variance changes. Under the null hypothesis, we derive the limit distribution of this test, which can be used to detect the change-point of variance. Under the alternative hypothesis, the limit behavior of the CUSUM test is also derived. To validate the performance of the test, we conducted simulation analyses on with Gaussian and Gamma errors. The results demonstrate that this testing method significantly outperforms existing approaches, particularly in detecting sparse variance changes. Finally, we conducted a practical case study using panel data from the Shanghai Shenzhen CSI 300 Index Components. Not only did we successfully identify the change-points of variance, but we also delved deeper into the underlying economic drivers behind these changes.

stat.ME

REAMP: A Stochastic Resonance Approach for Multi-Change Point Detection in High-Dimensional Data

Detecting multiple structural breaks in high-dimensional data remains a challenge, particularly when changes occur in higher-order moments or within complex manifold structures. In this paper, we propose REAMP (Resonance-Enhanced Analysis of Multi-change Points), a novel framework that integrates optimal transport theory with the physical principles of stochastic resonance. By utilizing a two-stage dimension reduction via the Earth Movers Distance (EMD) and Shortest Hamiltonian Paths (SHP), we map high-dimensional observations onto a graph-based count statistic. To overcome the locality constraints of traditional search algorithms, we implement a stochastic resonance system that utilizes randomized Beta-density priors to vibrate the objective function. This process allows multiple change points to resonate as global minima across iterative simulations, generating a candidate point cloud. A double-sharpening procedure is then applied to these candidates to pinpoint precise change point locations. We establish the asymptotic consistency of the resonance estimator and demonstrate through simulations that REAMP outperforms state-of-the-art methods, especially in scenarios involving simultaneous mean and variance shifts. The practical utility of the method is further validated through an application to time-lapse embryo monitoring, where REAMP provides both accurate detection and intuitive visualization of cell division stages.

stat.ME

Adaptive Kernel Regression for Constrained Route Alignment: Theory and Iterative Data Sharpening

Route alignment design in surveying and transportation engineering frequently involves fixed waypoint constraints, where a path must precisely traverse specific coordinates. While existing literature primarily relies on geometric optimization or control-theoretic spline frameworks, there is a lack of systematic statistical modeling approaches that balance global smoothness with exact point adherence. This paper proposes an Adaptive Nadaraya-Watson (ANW) kernel regression estimator designed to address the fixed waypoint problem. By incorporating waypoint-specific weight tuning parameters, the ANW estimator decouples global smoothing from local constraint satisfaction, avoiding the "jagged" artifacts common in naive local bandwidth-shrinking strategies. To further enhance estimation accuracy, we develop an iterative data sharpening algorithm that systematically reduces bias while maintaining the stability of the kernel framework. We establish the theoretical foundation for the ANW estimator by deriving its asymptotic bias and variance and proving its convergence properties under the internal constraint model. Numerical case studies in 1D and 2D trajectory planning demonstrate that the method effectively balances root mean square error (RMSE) and curvature smoothness. Finally, we validate the practical utility of the framework through empirical applications to railway and highway route planning. In sum, this work provides a stable, theoretically grounded, and computationally efficient solution for complex, constrained alignment design problems.

stat.ME

Propagation-Distance Limit for a Classical Nonlocal Optical System

We derive closed-form analog quantum-speed-limit (QSL) bounds for highly nonlocal optical beams whose paraxial propagation is mapped to a reversed (inverted) harmonic-oscillator generator. Treating the longitudinal coordinate $z$ as an evolution parameter (propagation distance), we construct the propagator, evaluate the Bures distance, and obtain analytic Mandelstam--Tamm and Margolus--Levitin bounds that fix a propagation-distance limit $z_{\mathrm{PDL}}$ to reach a prescribed mode distinguishability. This distance-domain constraint is the classical optical analogue of the minimal orthogonality time in quantum mechanics. We then propose a compact self-defocusing PDL beam shaper that achieves strong transverse-mode conversion within millimeter scales. We further show that small variations in refractive index, beam power, or temperature shift $z_{\mathrm{SL}}$ with high leverage, enabling speed-limit-based metrology with index sensitivities down to $10^{-7}$ RIU and temperature resolutions of order $1$ mK. The results bridge distance-domain QSL geometry and practical photonic applications.

quant-ph

Product Depth for Temporal Point Processes Observed Only Up to the First k Events

Temporal point processes (TPPs) model the timing of discrete events along a timeline and are widely used in fields such as neuroscience and fi- nance. Statistical depth functions are powerful tools for analyzing centrality and ranking in multivariate and functional data, yet existing depth notions for TPPs remain limited. In this paper, we propose a novel product depth specifically designed for TPPs observed only up to the first k events. Our depth function comprises two key components: a normalized marginal depth, which captures the temporal distribution of the final event, and a conditional depth, which characterizes the joint distribution of the preceding events. We establish its key theoretical properties and demonstrate its practical utility through simulation studies and real data applications.

stat.ME

Asymptotic linear dependence and ellipse statistics for multivariate two-sample homogeneity test

Statistical depth, which measures the center-outward rank of a given sample with respect to its underlying distribution, has become a popular and powerful tool in nonparametric inference. In this paper, we investigate the use of statistical depth in multivariate two-sample problems. We propose a new depth-based nonparametric two-sample test, which has the Chi-square(1) asymptotic distribution under the null hypothesis. Simulations and real-data applications highlight the efficacy and practical value of the proposed test.

stat.ME

DEEPEAST technique to enhance power in two-sample tests via the same-attraction function

Data depth has emerged as an invaluable nonparametric measure for the ranking of multivariate samples. The main contribution of depth-based two-sample comparisons is the introduction of the Q statistic (Liu and Singh, 1993), a quality index. Unlike traditional methods, data depth does not require the assumption of normal distributions and adheres to four fundamental properties. Many existing two-sample homogeneity tests, which assess mean and/or scale changes in distributions often suffer from low statistical power or indeterminate asymptotic distributions. To overcome these challenges, we introduced a DEEPEAST (depth-explored same-attraction sample-to-sample central-outward ranking) technique for improving statistical power in two-sample tests via the same-attraction function. We proposed two novel and powerful depth-based test statistics: the sum test statistic and the product test statistic, which are rooted in Q statistics, share a "common attractor" and are applicable across all depth functions. We further proved the asymptotic distribution of these statistics for various depth functions. To assess the performance of power gain, we apply three depth functions: Mahalanobis depth (Liu and Singh, 1993), Spatial depth (Brown, 1958; Gower, 1974), and Projection depth (Liu, 1992). Through two-sample simulations, we have demonstrated that our sum and product statistics exhibit superior power performance, utilizing a strategic block permutation algorithm and compare favourably with popular methods in literature. Our tests are further validated through analysis on Raman spectral data, acquired from cellular and tissue samples, highlighting the effectiveness of the proposed tests highlighting the effective discrimination between health and cancerous samples.

stat.ME

Q statistics in data depth: fundamental theory revisited and variants

Recently, data depth has been widely used to rank multivariate data. The study of the depth-based $Q$ statistic, originally proposed by Liu and Singh (1993), has become increasingly popular when it can be used as a quality index to differentiate between two samples. Based on the existing theoretical foundations, more and more variants have been developed for increasing power in the two sample test. However, the asymptotic expansion of the $Q$ statistic in the important foundation work of Zuo and He (2006) currently has an optimal rate $m^{-3/4}$ slower than the target $m^{-1}$, leading to limitations in higher-order expansions for developing more powerful tests. We revisit the existing assumptions and add two new plausible assumptions to obtain the target rate by applying a new proof method based on the Hoeffding decomposition and the Cox-Reid expansion. The aim of this paper is to rekindle interest in asymptotic data depth theory, to place Q-statistical inference on a firmer theoretical basis, to show its variants in current research, to open the door to the development of new theories for further variants requiring higher-order expansions, and to explore more of its potential applications.

math.ST

Adaptive Penalized Likelihood method for Markov Chains

Maximum Likelihood Estimation (MLE) and Likelihood Ratio Test (LRT) are widely used methods for estimating the transition probability matrix in Markov chains and identifying significant relationships between transitions, such as equality. However, the estimated transition probability matrix derived from MLE lacks accuracy compared to the real one, and LRT is inefficient in high-dimensional Markov chains. In this study, we extended the adaptive Lasso technique from linear models to Markov chains and proposed a novel model by applying penalized maximum likelihood estimation to optimize the estimation of the transition probability matrix. Meanwhile, we demonstrated that the new model enjoys oracle properties, which means the estimated transition probability matrix has the same performance as the real one when given. Simulations show that our new method behave very well overall in comparison with various competitors. Real data analysis further convince the value of our proposed method.

stat.ME

Modeling Long Sequences in Bladder Cancer Recurrence: A Comparative Evaluation of LSTM,Transformer,and Mamba

Traditional survival analysis methods often struggle with complex time-dependent data,failing to capture and interpret dynamic characteristics adequately.This study aims to evaluate the performance of three long-sequence models,LSTM,Transformer,and Mamba,in analyzing recurrence event data and integrating them with the Cox proportional hazards model.This study integrates the advantages of deep learning models for handling long-sequence data with the Cox proportional hazards model to enhance the performance in analyzing recurrent events with dynamic time information.Additionally,this study compares the ability of different models to extract and utilize features from time-dependent clinical recurrence data.The LSTM-Cox model outperformed both the Transformer-Cox and Mamba-Cox models in prediction accuracy and model fit,achieving a Concordance index of up to 0.90 on the test set.Significant predictors of bladder cancer recurrence,such as treatment stop time,maximum tumor size at recurrence and recurrence frequency,were identified.The LSTM-Cox model aligned well with clinical outcomes,effectively distinguishing between high-risk and low-risk patient groups.This study demonstrates that the LSTM-Cox model is a robust and efficient method for recurrent data analysis and feature extraction,surpassing newer models like Transformer and Mamba.It offers a practical approach for integrating deep learning technologies into clinical risk prediction systems,thereby improving patient management and treatment outcomes.

cs.LG

Information Theoretical Approach to Detecting Quantum Gravitational Corrections

One way to test quantum gravitational corrections is through black hole physics. In this paper, We investigate the scales at which quantum gravitational corrections can be detected in a black hole using information theory. This is done by calculating the Kullback-Leibler divergence for the probability distributions obtained from the Parikh-Wilczek formalism. We observe that the quantum gravitational corrections increase the Kullback-Leibler divergence as the mass of the black hole decreases, which is expected as quantum gravitational corrections can be neglected for larger black holes. However, we further observe that after a certain critical value, quantum gravitational corrections tend to decrease again as the mass of the black hole decreases. To understand the reason behind this behavior, we explicitly obtain Fisher information about such quantum gravitational corrections and find that it also increases as the mass decreases, but again, after a critical value, it decreases. This is because at such a scale, quantum fluctuations dominate the system and we lose information about the system. We obtain these results for higher-dimensional black holes and observe this behavior for Kullback-Leibler divergence and Fisher information depending on the dimensions of the black hole. These results can quantify the scale dependence and dimension dependence of the difficulty in detecting quantum gravitational corrections.

hep-th

Classifying deviation from standard quantum behavior using Kullback Leibler divergence

In this letter, we propose a novel statistical method to measure which system is better suited to probe small deviations from the usual quantum behavior. Such deviations are motivated by a number of theoretical and phenomenological motivations, and various systems have been proposed to test them. We propose that measuring deviations from quantum mechanics for a system would be easier if it has a higher Kullback Leibler divergence. We show this explicitly for a nonlocal Schrodinger equation and argue that it will hold for any modification to standard quantum behaviour. Thus, the results of this letter can be used to classify a wide range of theoretical and phenomenological models.

quant-ph

Multivariate two-sample test statistics based on data depth

Data depth has been applied as a nonparametric measurement for ranking multivariate samples. In this paper, we focus on homogeneity tests to assess whether two multivariate samples are from the same distribution. There are many data depth-based tests for this problem, but they may not be very powerful, or have unknown asymptotic distributions, or have slow convergence rates to asymptotic distributions. Given the recent development of data depth as an important measure in quality assurance, we propose three new test statistics for multivariate two-sample homogeneity tests. The proposed minimum test statistics have simple asymptotic half-normal distribution. We also discuss the generalization of the proposed tests to multiple samples. The simulation study demonstrates the superior performance of the proposed tests. The test procedure is illustrated by two real data examples.

math.ST

Quantum Thermodynamics of an $α^{\prime}$-Corrected Reissner-Nordström Black Hole

In this paper, we will analyze the effects of $α^{\prime} $ corrections on the behavior of a Reissner-Nordström black hole. We will calculate the effects of such corrections on the thermodynamics and thermodynamic stability of such a black hole. We will also derived a novel $α^{\prime}$-corrected first law. We will investigate the effect of such corrections on the Parikh-Wilczek formalism. This will be done using cross entropy and Kullback-Leibler divergence between the original probability distribution and the $α^{\prime}$-corrected probability distribution. We will then analyze the non-equilibrium quantum thermodynamics of this black hole. It will be observed that its quantum thermodynamics is corrected due to quantum gravitational corrections. We will use Ramsey scheme for emitted particles to calculate the quantum work distribution for this system. The average quantum work will be related to the difference of $α^{\prime}$-corrected free energies using the Jarzynski equality.

hep-th

Two edge-count tests and relevance analysis in k high-dimensional samples

For the task of relevance analysis, the conventional Tukey's test may be applied to the set of all pairwise comparisons. However, there were few studies that discuss both nonparametric k-sample comparisons and relevance analysis in high dimensions. Our aim is to capture the degree of relevance between combined samples and provide additional insights and advantages in high-dimensional k-sample comparisons. Our solution is to extend a graph-based two-sample comparison and investigate its availability for large and unequal sample sizes. We propose two distribution-free test statistics based on between-sample edge counts and measure the degree of relevance by standardized counts. The asymptotic permutation null distributions of the proposed statistics are derived, and the power gain is proved when the sample sizes are smaller than the square root of the dimension. We also discuss different edge costs in the graph to compare the parameters of the distributions. Simulation comparisons and real data analysis of tumors and images further convince the value of our proposed method. Software implementing the relevance analysis is available in the R package Relevance.

stat.ME

Ternary primitive LCD BCH codes

Absolute coset leaders were first proposed by the authors which have advantages in constructing binary LCD BCH codes. As a continue work, in this paper we focus on ternary linear codes. Firstly, we find the largest, second largest, and third largest absolute coset leaders of ternary primitive BCH codes. Secondly, we present three classes of ternary primitive BCH codes and determine their weight distributions. Finally, we obtain some LCD BCH codes and calculate some weight distributions. However, the calculation of weight distributions of two of these codes is equivalent to that of Kloosterman sums.

cs.IT

Exploring the space-time pattern of log-transformed infectious count of COVID-19: a clustering-segmented autoregressive sigmoid model

At the end of April 20, 2020, there were only a few new COVID-19 cases remaining in China, whereas the rest of the world had shown increases in the number of new cases. It is of extreme importance to develop an efficient statistical model of COVID-19 spread, which could help in the global fight against the virus. We propose a clustering-segmented autoregressive sigmoid (CSAS) model to explore the space-time pattern of the log-transformed infectious count. Four key characteristics are included in this CSAS model, including unknown clusters, change points, stretched S-curves, and autoregressive terms, in order to understand how this outbreak is spreading in time and in space, to understand how the spread is affected by epidemic control strategies, and to apply the model to updated data from an extended period of time. We propose a nonparametric graph-based clustering method for discovering dissimilarity of the curve time series in space, which is justified with theoretical support to demonstrate how the model works under mild and easily verified conditions. We propose a very strict purity score that penalizes overestimation of clusters. Simulations show that our nonparametric graph-based clustering method is faster and more accurate than the parametric clustering method regardless of the size of data sets. We provide a Bayesian information criterion (BIC) to identify multiple change points and calculate a confidence interval for a mean response. By applying the CSAS model to the collected data, we can explain the differences between prevention and control policies in China and selected countries.

stat.ME