Searcharxiv⌕ Search

arXiv subjects

Jian Qing Shi

Publications and source records attributed to Jian Qing Shi.

At least 19 recordsLinked to original sources

Heterogeneity-calibrated Byzantine-robust distributed composite quantile regression

We study sparse composite quantile regression (CQR) for distributed data with heterogeneous honest sites and Byzantine workers. Honest sites share a common slope but may differ in their covariate distributions, error laws, and quantile intercepts. The proposed heterogeneity-calibrated robust CQR (HC-RCQR) profiles local intercepts and calibrates scores using an approximate inverse profile Hessian. Honest workers transmit the resulting vectors, whereas Byzantine workers may send arbitrary vectors. The server updates the estimate by coordinatewise trimming and soft thresholding. A scalar example shows how unequal honest-site curvatures allow intermediate Byzantine reports to survive trimming and how ideal calibration reduces their possible effect. We also establish nonidentification of the mean of unrestricted honest-site slopes when fault identities are unknown. Under suitable conditions, we establish conditional contraction and support-recovery guarantees. The bound separates score offset, sampling fluctuation, contamination, calibration error, and the Newton remainder. Simulations and a bike-demand study examine performance under heterogeneous data and adversarial messages.

stat.ME↗

Causal Inference for Heterogeneous Extreme Quantiles with Heavy-Tailed Outcomes

We propose a framework for estimating conditional extreme quantile treatment effects (CEQTEs) in observational studies with heavy-tailed outcomes. Our procedure first estimates intermediate conditional quantiles using inverse-probability-weighted (IPW) quantile regression and then extrapolates them to extreme levels using extreme value theory. Under a linear conditional quantile model, we show that the conditional and marginal distributions of each potential outcome share a common extreme value index (EVI), motivating two complementary Hill-type EVI estimators based on conditional and marginal information, respectively. On the theoretical front, we introduce an IPW tail quantile score process that bridges regression quantile score processes and uniform tail empirical processes while accounting for treatment assignment. We establish its functional weak convergence under mild regularity conditions, without requiring a max-domain-of-attraction condition. This result provides the probabilistic foundation for the asymptotic analysis of the proposed CEQTE estimators. Simulation studies demonstrate favorable finite-sample performance, and an application to NLSY79 data reveals substantial heterogeneity in the effect of college education on extremely high hourly wages across confounder-defined subpopulations.

stat.ME↗

Analyzing Functional Data with a Mixture of Covariance Structures Using a Curve-Based Sampling Scheme

Motivated by distinct walking patterns in real-world free-living gait data, this paper proposes an innovative curve-based sampling scheme for the analysis of functional data characterized by a mixture of covariance structures. Traditional approaches often fail to adequately capture inherent complexities arising from heterogeneous covariance patterns across distinct subsets of the data. We introduce a unified Bayesian framework that integrates a nonlinear regression function with a continuous-time hidden Markov model, enabling the identification and utilization of varying covariance structures. One of the key contributions is the development of a computationally efficient curve-based sampling scheme for hidden state estimation, addressing the sampling complexities associated with high-dimensional, conditionally dependent data. This paper details the Bayesian inference procedure, examines the asymptotic properties to ensure the structural consistency of the model, and demonstrates its effectiveness through simulated and real-world examples.

stat.ME↗

Intrinsic Gaussian Process Regression Modeling for Manifold-valued Response Variable

Extrinsic Gaussian process regression methods, such as wrapped Gaussian process, have been developed to analyze manifold data. However, there is a lack of intrinsic Gaussian process methods for studying complex data with manifold-valued response variables. In this paper, we first apply the parallel transport operator on Riemannian manifold to propose an intrinsic covariance structure that addresses a critical aspect of constructing a well-defined Gaussian process regression model. We then propose a novel intrinsic Gaussian process regression model for manifold-valued data, which can be applied to data situated not only on Euclidean submanifolds but also on manifolds without a natural ambient space. We establish the asymptotic properties of the proposed models, including information consistency and posterior consistency, and we also show that the posterior distribution of the regression function is invariant to the choice of orthonormal frames for the coordinate representations of the covariance function. Numerical studies, including simulation and real examples, indicate that the proposed methods work well.

stat.ML↗

Bayesian analysis of nonlinear structured latent factor models using a Gaussian Process Prior

Factor analysis models are widely utilized in social and behavioral sciences, such as psychology, education, and marketing, to measure unobservable latent traits. In this article, we introduce a nonlinear structured latent factor analysis model which is more flexible to characterize the relationship between manifest variables and latent factors. The confirmatory identifiability of the latent factor is discussed, ensuring the substantive interpretation of the latent factors. A Bayesian approach with a Gaussian process prior is proposed to estimate the unknown nonlinear function and the unknown parameters. Asymptotic results are established, including structural identifiability of the latent factors, consistency of the estimates of the unknown parameters and the unknown nonlinear function. Simulation studies and a real data analysis are conducted to investigate the performance of the proposed method. Simulation studies show our proposed method performs well in handling nonlinear model and successfully identifies the latent factors. Our analysis incorporates oil flow data, allowing us to uncover the underlying structure of latent nonlinear patterns.

stat.ME↗

Wrapped Gaussian Process Functional Regression Model for Batch Data on Riemannian Manifolds

Regression is an essential and fundamental methodology in statistical analysis. The majority of the literature focuses on linear and nonlinear regression in the context of the Euclidean space. However, regression models in non-Euclidean spaces deserve more attention due to collection of increasing volumes of manifold-valued data. In this context, this paper proposes a concurrent functional regression model for batch data on Riemannian manifolds by estimating both mean structure and covariance structure simultaneously. The response variable is assumed to follow a wrapped Gaussian process distribution. Nonlinear relationships between manifold-valued response variables and multiple Euclidean covariates can be captured by this model in which the covariates can be functional and/or scalar. The performance of our model has been tested on both simulated data and real data, showing it is an effective and efficient tool in conducting functional data regression on Riemannian manifolds.

stat.ME↗

Designing Compact Features for Remote Stroke Rehabilitation Monitoring using Wearable Accelerometers

Stroke is known as a major global health problem, and for stroke survivors it is key to monitor the recovery levels. However, traditional stroke rehabilitation assessment methods (such as the popular clinical assessment) can be subjective and expensive, and it is also less convenient for patients to visit clinics in a high frequency. To address this issue, in this work based on wearable sensing and machine learning techniques, we develop an automated system that can predict the assessment score in an objective manner. With wrist-worn sensors, accelerometer data is collected from 59 stroke survivors in free-living environments for a duration of 8 weeks, and we map the week-wise accelerometer data(3 days per week) to the assessment score by developing signal processing and predictive model pipeline. To achieve this, we propose two types of new features, which can encode the rehabilitation information from both paralysed and non-paralysed sides while suppressing the high level noises such as irrelevant daily activities. Based on the proposed features, we further develop the longitudinal mixed-effects model with Gaussian process prior (LMGP), which can model the random effects caused by different subjects and time slots (during the 8 weeks). Comprehensive experiments are conducted to evaluate our system on both acute and chronic patients, and the promising results suggest its effectiveness.

eess.SP↗

Towards Automated Fatigue Assessment using Wearable Sensing and Mixed-Effects Models

Fatigue is a broad, multifactorial concept that includes the subjective perception of reduced physical and mental energy levels. It is also one of the key factors that strongly affect patients' health-related quality of life. To date, most fatigue assessment methods were based on self-reporting, which may suffer from many factors such as recall bias. To address this issue, in this work, we recorded multi-modal physiological data (including ECG, accelerometer, skin temperature and respiratory rate, as well as demographic information such as age, BMI) in free-living environments and developed automated fatigue assessment models. Specifically, we extracted features from each modality and employed the random forest-based mixed-effects models, which can take advantage of the demographic information for improved performance. We conducted experiments on our collected dataset, and very promising preliminary results were achieved. Our results suggested ECG played an important role in the fatigue assessment tasks.

cs.HC↗

Gaussian Process for Functional Data Analysis: The GPFDA Package for R

We present and describe the GPFDA package for R. The package provides flexible functionalities for dealing with Gaussian process regression (GPR) models for functional data. Multivariate functional data, functional data with multidimensional inputs, and nonseparable and/or nonstationary covariance structures can be modeled. In addition, the package fits functional regression models where the mean function depends on scalar and/or functional covariates and the covariance structure is modeled by a GPR model. In this paper, we present the versatility of GPFDA with respect to mean function and covariance function specifications and illustrate the implementation of estimation and prediction of some models through reproducible numerical examples.

stat.CO↗

Joint Curve Registration and Classification with Two-level Functional Models

Many classification techniques when the data are curves or functions have been recently proposed. However, the presence of misaligned problems in the curves can influence the performance of most of them. In this paper, we propose a model-based approach for simultaneous curve registration and classification. The method is proposed to perform curve classification based on a functional logistic regression model that relies on both scalar variables and functional variables, and to align curves simultaneously via a data registration model. EM-based algorithms are developed to perform maximum likelihood inference of the proposed models. We establish the identifiability results for curve registration model and investigate the asymptotic properties of the proposed estimation procedures. Simulation studies are conducted to demonstrate the finite sample performance of the proposed models. An application of the hyoid bone movement data from stroke patients reveals the effectiveness of the new models.

stat.ME↗

Modeling Function-Valued Processes with Nonseparable and/or Nonstationary Covariance Structure

We discuss a general Bayesian framework on modeling multidimensional function-valued processes by using a Gaussian process or a heavy-tailed process as a prior, enabling us to handle nonseparable and/or nonstationary covariance structure. The nonstationarity is introduced by a convolution-based approach through a varying anisotropy matrix, whose parameters vary along the input space and are estimated via a local empirical Bayesian method. For the varying matrix, we propose to use a spherical parametrization, leading to unconstrained and interpretable parameters. The unconstrained nature allows the parameters to be modeled as a nonparametric function of time, spatial location or other covariates. The interpretation of the parameters is based on closed-form expressions, providing valuable insights into nonseparable covariance structures. Furthermore, to extract important information in data with complex covariance structure, the Bayesian framework can decompose the function-valued processes using the eigenvalues and eigensurfaces calculated from the estimated covariance structure. The results are demonstrated by simulation studies and by an application to wind intensity data. Supplementary materials for this article are available online.

stat.ME↗

A robust estimation for the extended t-process regression model

Robust estimation and variable selection procedure are developed for the extended t-process regression model with functional data. Statistical properties such as consistency of estimators and predictions are obtained. Numerical studies show that the proposed method performs well.

stat.AP↗

Simultaneous Registration and Clustering for Multi-dimensional Functional Data

The clustering for functional data with misaligned problems has drawn much attention in the last decade. Most methods do the clustering after those functional data being registered and there has been little research using both functional and scalar variables. In this paper, we propose a simultaneous registration and clustering (SRC) model via two-level models, allowing the use of both types of variables and also allowing simultaneous registration and clustering. For the data collected from subjects in different unknown groups, a Gaussian process functional regression model with time warping is used as the first level model; an allocation model depending on scalar variables is used as the second level model providing further information over the groups. The former carries out registration and modeling for the multi-dimensional functional data (2D or 3D curves) at the same time. This methodology is implemented using an EM algorithm, and is examined on both simulated data and real data.

stat.ME↗

Regression Analysis for Multivariate Dependent Count Data Using Convolved Gaussian Processes

Research on Poisson regression analysis for dependent data has been developed rapidly in the last decade. One of difficult problems in a multivariate case is how to construct a cross-correlation structure and at the meantime make sure that the covariance matrix is positive definite. To address the issue, we propose to use convolved Gaussian process (CGP) in this paper. The approach provides a semi-parametric model and offers a natural framework for modeling common mean structure and covariance structure simultaneously. The CGP enables the model to define different covariance structure for each component of the response variables. This flexibility ensures the model to cope with data coming from different resources or having different data structures, and thus to provide accurate estimation and prediction. In addition, the model is able to accommodate large-dimensional covariates. The definition of the model, the inference and the implementation, as well as its asymptotic properties, are discussed. Comprehensive numerical examples with both simulation studies and real data are presented.

stat.ME↗

Robust functional regression model for marginal mean and subject-specific inferences

We introduce flexible robust functional regression models, using various heavy-tailed processes, including a Student $t$-process. We propose efficient algorithms in estimating parameters for the marginal mean inferences and in predicting conditional means as well interpolation and extrapolation for the subject-specific inferences. We develop bootstrap prediction intervals for conditional mean curves. Numerical studies show that the proposed model provides robust analysis against data contamination or distribution misspecification, and the proposed prediction intervals maintain the nominal confidence levels. A real data application is presented as an illustrative example.

stat.ME↗

Nonlinear Mixed-effects Scalar-on-function Models and Variable Selection for Kinematic Upper Limb Movement Data

This paper arises from collaborative research the aim of which was to model clinical assessments of upper limb function after stroke using 3D kinematic data. We present a new nonlinear mixed-effects scalar-on-function regression model with a Gaussian process prior focusing on variable selection from large number of candidates including both scalar and function variables. A novel variable selection algorithm has been developed, namely functional least angle regression (fLARS). As they are essential for this algorithm, we studied the representation of functional variables with different methods and the correlation between a scalar and a group of mixed scalar and functional variables. We also propose two new stopping rules for practical usage. This algorithm is able to do variable selection when the number of variables is larger than the sample size. It is efficient and accurate for both variable selection and parameter estimation. Our comprehensive simulation study showed that the method is superior to other existing variable selection methods. When the algorithm was applied to the analysis of the 3D kinetic movement data the use of the non linear random-effects model and the function variables significantly improved the prediction accuracy for the clinical assessment.

stat.AP↗

Automatic Detection of Significant Areas for Functional Data with Directional Error Control

To detect differences between the mean curves of two samples in longitudinal study or functional data analysis, we usually need to partition the temporal or spatial domain into several pre-determined sub-areas. In this paper we apply the idea of large-scale multiple testing to find the significant sub-areas automatically in a general functional data analysis framework. A nonparametric Gaussian process regression model is introduced for two-sided multiple tests. We derive an optimal test which controls directional false discovery rates and propose a procedure by approximating it on a continuum. The proposed procedure controls directional false discovery rates at any specified level asymptotically. In addition, it is computationally inexpensive and able to accommodate different time points for observations across the samples. Simulation studies are presented to demonstrate its finite sample performance. We also apply it to an executive function research in children with Hemiplegic Cerebral Palsy and extend it to the equivalence tests.

stat.ME↗

Simulation-based Sensitivity Analysis for Non-ignorable Missing Data

Sensitivity analysis is popular in dealing with missing data problems particularly for non-ignorable missingness. It analyses how sensitively the conclusions may depend on assumptions about missing data e.g. missing data mechanism (MDM). We called models under certain assumptions sensitivity models. To make sensitivity analysis useful in practice we need to define some simple and interpretable statistical quantities to assess the sensitivity models. However, the assessment is difficult when the missing data mechanism is missing not at random (MNAR). We propose a novel approach in this paper on attempting to investigate those assumptions based on the nearest-neighbour (KNN) distances of simulated datasets from various MNAR models. The method is generic and it has been applied successfully to several specific models in this paper including meta-analysis model with publication bias, analysis of incomplete longitudinal data and regression analysis with non-ignorable missing covariates.

stat.ME↗