SearcharxivSearch

arXiv subjects

Hangjin Jiang

Publications and source records attributed to Hangjin Jiang.

8 recordsLinked to original sources

Structure-Adaptive E-Value Filter for Detecting Regional Signals in Brain Imaging

Structural MRI provides a noninvasive view of the neuroanatomical differences associated with cognitive impairment and dementia. Using data from the Alzheimer's Disease Neuroimaging Initiative, we investigate which anatomical regions exhibit widespread gray-matter differences and how these regional patterns vary across the clinical spectrum. Addressing this goal requires translating spatially dependent voxel-level evidence into regional conclusions while controlling multiplicity across anatomical regions. We propose the Structure-adaptive E-value FilTer (SEFT), which uses flexible working models to construct spatially adaptive voxel-level scores and aggregates them into regional partial-conjunction e-values. When combined with the e-value Benjamini--Hochberg (e-BH) procedure, these e-values provide finite-sample control of the set-wise false discovery rate under arbitrary interregional dependence. The ADNI analysis reveals a coherent neuroanatomical pattern: the exploratory analysis shows that differences between normal cognition and mild cognitive impairment are concentrated in medial-temporal regions, whereas the differences between mild cognitive impairment and dementia extend more broadly into temporal--limbic and posterior association regions. Both patterns largely overlap the normal-cognition--dementia benchmark, identifying a shared anatomical core across the clinical comparisons.

stat.AP

Split-and-Conquer: Distributed Factor Modeling for High-Dimensional Matrix-Variate Time Series

In this paper, we propose a distributed framework for reducing the dimensionality of high-dimensional, large-scale, heterogeneous matrix-variate time series data using a factor model. The data are first partitioned column-wise (or row-wise) and allocated to node servers, where each node estimates the row (or column) loading matrix via two-dimensional tensor PCA. These local estimates are then transmitted to a central server and aggregated, followed by a final PCA step to obtain the global row (or column) loading matrix estimator. Given the estimated loading matrices, the corresponding factor matrices are subsequently computed. Unlike existing distributed approaches, our framework preserves the latent matrix structure, thereby improving computational efficiency and enhancing information utilization. We also discuss row- and column-wise clustering procedures for settings in which the group memberships are unknown. Furthermore, we extend the analysis to unit-root nonstationary matrix-variate time series. Asymptotic properties of the proposed method are derived for the diverging dimension of the data in each computing unit and the sample size $T$. Simulation results assess the computational efficiency and estimation accuracy of the proposed framework, and real data applications further validate its predictive performance.

stat.ML

Regularized Estimation of High-Dimensional Matrix-Variate Autoregressive Models

Matrix-variate time series data are increasingly popular in economics, statistics, and environmental studies, among other fields. This paper develops regularized estimation methods for analyzing high-dimensional matrix-variate time series using bilinear matrix-variate autoregressive models. The bilinear autoregressive structure is widely used for matrix-variate time series, as it reduces model complexity while capturing interactions between rows and columns. However, when dealing with large dimensions, the commonly used iterated least-squares method results in numerous estimated parameters, making interpretation difficult. To address this, we propose two regularized estimation methods to further reduce model dimensionality. The first assumes banded autoregressive coefficient matrices, where each data point interacts only with nearby points. A two-step estimation method is used: first, traditional iterated least-squares is applied for initial estimates, followed by a banded iterated least-squares approach. A Bayesian Information Criterion (BIC) is introduced to estimate the bandwidth of the coefficient matrices. The second method assumes sparse autoregressive matrices, applying the LASSO technique for regularization. We derive asymptotic properties for both methods as the dimensions diverge and the sample size $T\rightarrow\infty$. Simulations and real data examples demonstrate the effectiveness of our methods, comparing their forecasting performance against common autoregressive models in the literature.

stat.ME

Bayesian Variable Selection for Single Index Logistic Model

In the era of big data, variable selection is a key technology for handling high-dimensional problems with a small sample size but a large number of covariables. Different variable selection methods were proposed for different models, such as linear model, logistic model and generalized linear model. However, fewer works focused on variable selection for single index models, especially, for single index logistic model, due to the difficulty arose from the unknown link function and the slow mixing rate of MCMC algorithm for traditional logistic model. In this paper, we proposed a Bayesian variable selection procedure for single index logistic model by taking the advantage of Gaussian process and data augmentation. Numerical results from simulations and real data analysis show the advantage of our method over the state of arts.

stat.ME

A Goodness-of-Fit Test for Statistical Models

Statistical modeling plays a fundamental role in understanding the underlying mechanism of massive data (statistical inference) and predicting the future (statistical prediction). Although all models are wrong, researchers try their best to make some of them be useful. The question here is how can we measure the usefulness of a statistical model for the data in hand? This is key to statistical prediction. The important statistical problem of testing whether the observations follow the proposed statistical model has only attracted relatively few attentions. In this paper, we proposed a new framework for this problem through building its connection with two-sample distribution comparison. The proposed method can be applied to evaluate a wide range of models. Examples are given to show the performance of the proposed method.

stat.ME

Equitability of Dependence Measure

Measuring dependence between two random variables is very important, and critical in many applied areas such as variable selection, brain network analysis. However, we do not know what kind of functional relationship is between two covariates, which requires the dependence measure to be equitable. That is, it gives similar scores to equally noisy relationship of different types. In fact, the dependence score is a continuous random variable taking values in $[0,1]$, thus it is theoretically impossible to give similar scores. In this paper, we introduce a new definition of equitability of a dependence measure, i.e, power-equitable (weak-equitable) and show by simulation that HHG and Copula Dependence Coefficient (CDC) are weak-equitable.

stat.ML

Dependence Measure for non-additive model

We proposed a new statistical dependency measure called Copula Dependency Coefficient(CDC) for two sets of variables based on copula. It is robust to outliers, easy to implement, powerful and appropriate to high-dimensional variables. These properties are important in many applications. Experimental results show that CDC can detect the dependence between variables in both additive and non-additive models.

stat.ML

The Link between Magnetic-field Orientations and Star Formation Rates

Understanding star formation rates (SFR) is a central goal of modern star-formation models, which mainly involve gravity, turbulence and, in some cases, magnetic fields (B-fields). However, a connection between B-fields and SFR has never been observed. Here, a comparison between the surveys of SFR and a study of cloud-field alignment - which revealed a bimodal (parallel or perpendicular) alignment - shows consistently lower SFR per solar mass for clouds almost perpendicular to the B-fields. This is evidence of B-fields being a primary regulator of SFR. The perpendicular alignment possesses a significantly higher magnetic flux than the parallel alignment and thus a stronger support of the gas against self-gravity. This results in overall lower masses of the fragmented components, which are in agreement with the lower SFR.

astro-ph.GA