SearcharxivSearch

arXiv subjects

Peihua Qiu

Publications and source records attributed to Peihua Qiu.

7 recordsLinked to original sources

Analyzing zero-inflated clustered longitudinal ordinal outcomes using GEE-type models with an application to dental fluorosis studies

Motivated by the Iowa Fluoride Study (IFS), which tracked fluoride intake and dental outcomes from childhood to young adulthood (ages 9, 13, 17, and 23), we analyze dental fluorosis - a condition caused by excessive fluoride exposure during enamel formation. In this context, fluorosis scores across tooth surfaces present as zero-inflated, clustered, and longitudinal ordinal outcomes, prompting the development of a unified modeling framework. Leveraging generalized estimating equations (GEEs), we construct separate models for the presence and severity of fluorosis and propose a combined model that links these components though shared covariates. To improve estimation efficiency and borrowing strength across timepoints, we incorporate James-Stein shrinkage estimators. We compare several working correlation structures, including a data-driven jackknifed structure, and perform model selection via rank aggregation. Simulation studies validate the finite-sample performance of the proposed models, and a bootstrap-based power analysis further confirms the validity of the testing procedure. In our analysis of the IFS data, early-life total daily fluoride intake, average home water fluoride concentration, and specific teeth and zones emerge as significant risk factors for dental fluorosis. Maxillary lateral incisors and zones closer to the gum show protective effects across different ages. These findings reveal novel age-specific associations between early-life exposures and the progression of dental fluorosis through early adulthood.

stat.ME

Conditional-Marginal Nonparametric Estimation for Stage Waiting Times from Multi-Stage Models under Dependent Right Censoring

We investigate two population-level quantities (corresponding to complete data) related to uncensored stage waiting times in a progressive multi-stage model, conditional on a prior stage visit. We show how to estimate these quantities consistently using right-censored data. The first quantity is the stage waiting time distribution (survival function), representing the proportion of individuals who remain in stage j within time t after entering stage j. The second quantity is the cumulative incidence function, representing the proportion of individuals who transition from stage j to stage j' within time t after entering stage j. To estimate these quantities, we present two nonparametric approaches. The first uses an inverse probability of censoring weighting (IPCW) method, which reweights the counting processes and the number of individuals at risk (the at-risk set) to address dependent right censoring. The second method utilizes the notion of fractional observations (FRE) that modifies the at-risk set by incorporating probabilities of individuals (who might have been censored in a prior stage) eventually entering the stage of interest in the uncensored or full data experiment. Neither approach is limited to the assumption of independent censoring or Markovian multi-stage frameworks. Simulation studies demonstrate satisfactory performance for both sets of estimators, though the IPCW estimator generally outperforms the FRE estimator in the setups considered in our simulations. These estimations are further illustrated through applications to two real-world datasets: one from patients undergoing bone marrow transplants and the other from patients diagnosed with breast cancer.

stat.ME

Sequential Adaptive Design for Jump Regression Estimation

Selecting input variables or design points for statistical models has been of great interest in adaptive design and active learning. Motivated by two scientific examples, this paper presents a strategy of selecting the design points for a regression model when the underlying regression function is discontinuous. The first example we undertook was for the purpose of accelerating imaging speed in a high resolution material imaging; the second was use of sequential design for the purpose of mapping a chemical phase diagram. In both examples, the underlying regression functions have discontinuities, so many of the existing design optimization approaches cannot be applied because they mostly assume a continuous regression function. Although some existing adaptive design strategies developed from treed regression models can handle the discontinuities, the Bayesian approaches come with computationally expensive Markov Chain Monte Carlo techniques for posterior inferences and subsequent design point selections, which is not appropriate for the first motivating example that requires computation at least faster than the original imaging speed. In addition, the treed models are based on the domain partitioning that are inefficient when the discontinuities occurs over complex sub-domain boundaries. We propose a simple and effective adaptive design strategy for a regression analysis with discontinuities: some statistical properties with a fixed design will be presented first, and then these properties will be used to propose a new criterion of selecting the design points for the regression analysis. Sequential design with the new criterion will be presented with comprehensive simulated examples, and its application to the two motivating examples will be presented.

stat.ML

Jump detection in generalized error-in-variables regression with an application to Australian health tax policies

Without measurement errors in predictors, discontinuity of a nonparametric regression function at unknown locations could be estimated using a number of existing approaches. However, it becomes a challenging problem when the predictors contain measurement errors. In this paper, an error-in-variables jump point estimator is suggested for a nonparametric generalized error-in-variables regression model. A major feature of our method is that it does not impose any parametric distribution on the measurement error. Its performance is evaluated by both numerical studies and theoretical justifications. The method is applied to studying the impact of Medicare Levy Surcharge on the private health insurance take-up rate in Australia.

stat.AP

Statistical modeling of the time course of tantrum anger

Although anger is an important emotion that underlies much overt aggression at great social cost, little is known about how to quantify anger or to specify the relationship between anger and the overt behaviors that express it. This paper proposes a novel statistical model which provides both a metric for the intensity of anger and an approach to determining the quantitative relationship between anger intensity and the specific behaviors that it controls. From observed angry behaviors, we reconstruct the time course of the latent anger intensity and the linkage between anger intensity and the probability of each angry behavior. The data on which this analysis is based consist of observed tantrums had by 296 children in the Madison WI area during the period 1994--1996. For each tantrum, eight angry behaviors were recorded as occurring or not within each consecutive 30-second unit. So, the data can be characterized as a multivariate, binary, longitudinal (MBL) dataset with a latent variable (anger intensity) involved. Data such as these are common in biomedical, psychological and other areas of the medical and social sciences. Thus, the proposed modeling approach has broad applications.

stat.AP

Distribution-free cumulative sum control charts using bootstrap-based control limits

This paper deals with phase II, univariate, statistical process control when a set of in-control data is available, and when both the in-control and out-of-control distributions of the process are unknown. Existing process control techniques typically require substantial knowledge about the in-control and out-of-control distributions of the process, which is often difficult to obtain in practice. We propose (a) using a sequence of control limits for the cumulative sum (CUSUM) control charts, where the control limits are determined by the conditional distribution of the CUSUM statistic given the last time it was zero, and (b) estimating the control limits by bootstrap. Traditionally, the CUSUM control chart uses a single control limit, which is obtained under the assumption that the in-control and out-of-control distributions of the process are Normal. When the normality assumption is not valid, which is often true in applications, the actual in-control average run length, defined to be the expected time duration before the control chart signals a process change, is quite different from the nominal in-control average run length. This limitation is mostly eliminated in the proposed procedure, which is distribution-free and robust against different choices of the in-control and out-of-control distributions.

stat.AP

Nonparametric estimation of a point-spread function in multivariate problems

The removal of blur from a signal, in the presence of noise, is readily accomplished if the blur can be described in precise mathematical terms. However, there is growing interest in problems where the extent of blur is known only approximately, for example in terms of a blur function which depends on unknown parameters that must be computed from data. More challenging still is the case where no parametric assumptions are made about the blur function. There has been a limited amount of work in this setting, but it invariably relies on iterative methods, sometimes under assumptions that are mathematically convenient but physically unrealistic (e.g., that the operator defined by the blur function has an integrable inverse). In this paper we suggest a direct, noniterative approach to nonparametric, blind restoration of a signal. Our method is based on a new, ridge-based method for deconvolution, and requires only mild restrictions on the blur function. We show that the convergence rate of the method is close to optimal, from some viewpoints, and demonstrate its practical performance by applying it to real images.

math.ST