Searcharxiv⌕ Search

arXiv subjects

Cornelis J. Potgieter

Publications and source records attributed to Cornelis J. Potgieter.

8 recordsLinked to original sources

Tail Estimation via Sample Splitting: A Two-Sample Framework for Stable-like Distributions

Stable distributions provide a flexible framework for modeling heavy-tailed and skewed data, with the stability index $α$ quantifying tail heaviness. We propose a new semiparametric approach that leverages the two-sum closure property of stable distributions within a location-scale framework, \textcolor{black}{where the induced scale parameter satisfies $σ= 2^{1/α}$}. The method constructs two pseudo-independent samples from a single observed sample via repeated random splitting and empirical convolution, and then estimates $σ$ using weighted least squares applied to empirical quantiles. We also propose a bootstrap procedure that uses a reduced number of sample splits together with extrapolation to estimate standard errors, and we demonstrate good finite-sample performance of this procedure. This approach avoids intractable likelihood calculations and offers computational advantages over maximum likelihood estimation. We establish consistency and asymptotic properties of the estimator and assess its finite-sample performance through simulation studies. Beyond the benefit of substantial computational savings, results indicate competitive accuracy, particularly in heavy-tailed settings.

stat.ME↗

Inference for Error-Prone Count Data: Estimation under a Binomial Convolution Framework

Measurement error in count data is common but underexplored in the literature, particularly in contexts where observed scores are bounded and arise from discrete scoring processes. Motivated by applications in oral reading fluency assessment, we propose a binomial convolution framework that extends binary misclassification models to settings where only the aggregate number of correct responses is observed, and errors may involve both overcounting and undercounting the number of events. The model accommodates distinct true positive and true negative accuracy rates and preserves the bounded nature of the data. Assuming the availability of both contaminated and error-free scores on a subset of items, we develop and compare three estimation strategies: maximum likelihood estimation (MLE), linear regression, and generalized method of moments (GMM). Extensive simulations show that MLE is most accurate when the model is correctly specified but is computationally intensive and less robust to misspecification. Regression is simple and stable but less precise, while GMM offers a compromise in model dependence, though it is sensitive to outliers. In practice, this framework supports improved inference in unsupervised settings where contaminated scores serve as inputs to downstream analyses. By quantifying accuracy rates, the model enables score corrections even when no specific outcome is yet defined. We demonstrate its utility using real oral reading fluency data, comparing human and AI-generated scores. Findings highlight the practical implications of estimator choice and underscore the importance of explicitly modeling asymmetric measurement error in count data.

stat.ME↗

A Linear Errors-in-Variables Model with Unknown Heteroscedastic Measurement Errors

In the classic measurement error framework, covariates are contaminated by independent additive noise. This paper considers parameter estimation in such a linear errors-in-variables model where the unknown measurement error distribution is heteroscedastic across observations. We propose a new generalized method of moment (GMM) estimator that combines a moment correction approach and a phase function-based approach. The former requires distributions to have four finite moments, while the latter relies on covariates having asymmetric distributions. The new estimator is shown to be consistent and asymptotically normal under appropriate regularity conditions. The asymptotic covariance of the estimator is derived, and the estimated standard error is computed using a fast bootstrap procedure. The GMM estimator is demonstrated to have strong finite sample performance in numerical studies, especially when the measurement errors follow non-Gaussian distributions.

stat.ME↗

Penalized Likelihood Methods for Modeling Count Data

The paper considers parameter estimation in count data models using penalized likelihood methods. The motivating data consists of multiple independent count variables with a moderate sample size per variable. The data were collected during the assessment of oral reading fluency (ORF) in school-aged children. A sample of fourth-grade students were given one of ten available passages to read with these differing in length and difficulty. The observed number of words read incorrectly (WRI) is used to measure ORF. Three models are considered for WRI scores, namely the binomial, the zero-inflated binomial, and the beta-binomial. We aim to efficiently estimate passage difficulty, a quantity expressed as a function of the underlying model parameters. Two types of penalty functions are considered for penalized likelihood with respective goals of shrinking parameter estimates closer to zero or closer to one another. A simulation study evaluates the efficacy of the shrinkage estimates using Mean Square Error (MSE) as metric. Big reductions in MSE relative to unpenalized maximum likelihood are observed. The paper concludes with an analysis of the motivating ORF data.

stat.ME↗

Phase Function Density Deconvolution with Heteroscedastic Measurement Error of Unknown Type

It is important to properly correct for measurement error when estimating density functions associated with biomedical variables. These estimators that adjust for measurement error are broadly referred to as density deconvolution estimators. While most methods in the literature assume the distribution of the measurement error to be fully known, a recently proposed method based on the empirical phase function (EPF) can deal with the situation when the measurement error distribution is unknown. The EPF density estimator has only been considered in the context of additive and homoscedastic measurement error; however, the measurement error of many biomedical variables is heteroscedastic in nature. In this paper, we developed a phase function approach for density deconvolution when the measurement error has unknown distribution and is heteroscedastic. A weighted empirical phase function (WEPF) is proposed where the weights are used to adjust for heteroscedasticity of measurement error. The asymptotic properties of the WEPF estimator are evaluated. Simulation results show that the weighting can result in large decreases in mean integrated squared error (MISE) when estimating the phase function. The estimation of the weights from replicate observations is also discussed. Finally, the construction of a deconvolution density estimator using the WEPF is compared to an existing deconvolution estimator that adjusts for heteroscedasticity, but assumes the measurement error distribution to be fully known. The WEPF estimator proves to be competitive, especially when considering that it relies on the minimal assumption of the distribution of measurement error.

stat.ME↗

Density Deconvolution for Generalized Skew-Symmetric Distributions

This paper develops a density deconvolution estimator that assumes the density of interest is a member of the generalized skew-symmetric (GSS) family of distributions. Estimation occurs in two parts: a skewing function, as well as location and scale parameters must be estimated. A kernel method is proposed for estimating the skewing function. The mean integrated square error (MISE) of the resulting GSS deconvolution estimator is derived. Based on derivation of the MISE, two bandwidth estimation methods for estimating the skewing function are also proposed. A generalized method of moments (GMM) approach is developed for estimation of the location and scale parameters. The question of multiple solutions in applying the GMM is also considered, and two solution selection criteria are proposed. The GSS deconvolution estimator is further investigated in simulation studies and is compared to the nonparametric deconvolution estimator. For most simulation settings considered, the GSS estimator has performance superior to the nonparametric estimator.

stat.ME↗

A Latent Trait Model for Multivariate Longitudinal Data With Two Sources of Measurement Error

Personality traits are latent variables, and as such, are impossible to measure without the use of an assessment. Responses on the assessments can be influenced by both transient (state-related) error and measurement error, obscuring the true trait levels. Typically, these assessments utilize Likert scales, which yield only discrete data. The loss of information due to the discrete nature of the data represents an additional challenge in assessing the ability of these instruments to measure the latent trait of interest. This paper is concerned with parameter estimation in a model relating a latent variable, as well transient error and measurement error components when data are longitudinal and measured using a Likert scale. Two methods for parameter estimation are detailed: correlation reconstruction, a method that uses polychoric correlations, and maximum likelihood implemented using a Stochastic EM algorithm. These methods are applied to a motivating dataset of 440 college students taking the Big Five inventory twice in a two month period.

stat.AP↗

An EM Algorithm for Estimating an Oral Reading Speed and Accuracy Model

This study proposes a two-part model that includes components for reading accuracy and reading speed. The speed component is a log-normal factor model, for which speed data are measured by reading time for each sentence being assessed. The accuracy component is a binomial-count factor model, where the accuracy data are measured by the number of correctly read words in each sentence. Both underlying latent components are assumed to be Gaussian in nature. In this paper, the theoretical properties of the proposed model are developed and an Monte Carlo EM algorithm for model fitting is outlined. The predictive power of the model is illustrated in a real data application.

stat.AP↗