SearcharxivSearch

arXiv subjects

Dexter Cahoy

Publications and source records attributed to Dexter Cahoy.

10 recordsLinked to original sources

A Benchmark Study of Classical and Dual Polynomial Regression (DPR)-Based Probability Density Estimation Technique

The probability density function (PDF) plays a central role in statistical and machine learning modeling. Real-world data often deviates from Gaussian assumptions, exhibiting skewness and exponential decay. To evaluate how well different density estimation methods capture such irregularities, we generated six unimodal datasets from diverse distributions that reflect real-world anomalies. These were compared using parametric methods (Pearson Type I and Normal distribution) as well as non-parametric approaches, including histograms, kernel density estimation (KDE), and our proposed method. To accelerate computation, we implemented GPU-based versions of KDE (tKDE) and histogram estimation (tHDE) in TensorFlow, both of which outperform Python SciPy's KDE. Prior work demonstrated the use of piecewise modeling for density estimation, such as local polynomial regression; however, these methods are computationally intensive. Based on the concept of piecewise modeling, we developed a computationally efficient model, the Dual Polynomial Regression (DPR) method, which leverages tKDE or tHDE for training. DPR employs the piecewise strategy to split the PDF at its mode and fit polynomial regressions to the left and right halves independently, enabling better capture of the asymmetric shape of the unimodal distribution. We used the Mean Squared Error (MSE), Jensen-Shannon Divergence (JSD), and Pearson's correlation coefficient, with reference to the baseline PDF, to validate accuracy. We verified normalization using Area Under the Curve (AUC) and computational overhead via execution time. Validation on real-world systolic and diastolic data from 300,000 unique patients shows that the DPR of order 4, trained with tKDE, offers the best balance between accuracy and computational overhead.

stat.CO

Combining Data from Surveys and Related Sources

To improve the precision of inferences and reduce costs there is considerable interest in combining data from several sources such as sample surveys and administrative data. Appropriate methodology is required to ensure satisfactory inferences since the target populations and methods for acquiring data may be quite different. To provide improved inferences we use methodology that has a more general structure than the ones in current practice. We start with the case where the analyst has only summary statistics from each of the sources. In our primary method, uncertain pooling, it is assumed that the analyst can regard one source, survey $r$, as the single best choice for inference. This method starts with the data from survey $r$ and adds data from those other sources that are shown to form clusters that include survey $r$. We also consider Dirichlet process mixtures, one of the most popular nonparametric Bayesian methods. We use analytical expressions and the results from numerical studies to show properties of the methodology.

stat.ME

Bayesian inference for asymptomatic COVID-19 infection rates

To strengthen inferences meta analyses are commonly used to summarize information from a set of independent studies. In some cases, though, the data may not satisfy the assumptions underlying the meta analysis. Using three Bayesian methods that have a more general structure than the common meta analytic ones, we can show the extent and nature of the pooling that is justified statistically. In this paper, we re-analyze data from several reviews whose objective is to make inference about the COVID-19 asymptomatic infection rate. When it is unlikely that all of the true effect sizes come from a single source researchers should be cautious about pooling the data from all of the studies. Our findings and methodology are applicable to other COVID-19 outcome variables, and more generally.

stat.AP

Flexible models for overdispersed and underdispersed count data

Within the framework of probability models for overdispersed count data, we propose the generalized fractional Poisson distribution (gfPd), which is a natural generalization of the fractional Poisson distribution (fPd), and the standard Poisson distribution. We derive some properties of gfPd and more specifically we study moments, limiting behavior and other features of fPd. The skewness suggests that fPd can be left-skewed, right-skewed or symmetric; this makes the model flexible and appealing in practice. We apply the model to real big count data and estimate the model parameters using maximum likelihood. Then, we turn to the very general class of weighted Poisson distributions (WPD's) to allow both overdispersion and underdispersion. Similarly to Kemp's generalized hypergeometric probability distribution, which is based on hypergeometric functions, we analyze a class of WPD's related to a generalization of Mittag--Leffler functions. The proposed class of distributions includes the well-known COM-Poisson and the hyper-Poisson models. We characterize conditions on the parameters allowing for overdispersion and underdispersion, and analyze two special cases of interest which have not yet appeared in the literature.

math.PR

A bootstrap test for equality of variances

We introduce a bootstrap procedure to test the hypothesis $H_o$ that $K+1$ variances are homogeneous. The procedure uses a variance-based statistic, and is derived from a normal-theory test for equality of variances. The test equivalently expressed the hypothesis as $H_o: \mathbfη=( η_1,\ldots,η_{K+1})^T=\mathbf{0}$, where $η_i$'s are log contrasts of the population variances. A box-type acceptance region is constructed to test the hypothesis $H_o$. Simulation results indicated that our method is generally superior to the Shoemaker and Levene tests, and the bootstrapped version of Levene test in controlling the Type I and Type II errors.

stat.ME

Parameter estimation for fractional Poisson processes

The paper proposes a formal estimation procedure for parameters of the fractional Poisson process (fPp). Such procedures are needed to make the fPp model usable in applied situations. The basic idea of fPp, motivated by experimental data with long memory is to make the standard Poisson model more flexible by permitting non-exponential, heavy-tailed distributions of interarrival times and different scaling properties. We establish the asymptotic normality of our estimators for the two parameters appearing in our fPp model. This fact permits construction of the corresponding confidence intervals. The properties of the estimators are then tested using simulated data.

stat.ME

Estimation of Mittag-Leffler Parameters

We propose a procedure for estimating the parameters of the Mittag-Leffler (ML) and the generalized Mittag-Leffler (GML) distributions. The algorithm is less restrictive, computationally simple, and necessary to make these models usable in practice. A comparison with the fractional moment estimator indicated favorable results for the proposed method.

stat.ME

Inverse stable prior for exponential models

We consider a class of non-conjugate priors as a mixing family of distributions for a parameter (e.g., Poisson or gamma rate, inverse scale or precision of an inverse-gamma, inverse variance of a normal distribution) of an exponential subclass of discrete and continuous data distributions. The prior class is proper, nonzero at the origin (unlike the gamma and inverted beta priors with shape parameter less than one and Jeffreys prior for a Poisson rate), and is easy to generate random numbers from. The prior class also provides flexibility in capturing a wide array of prior beliefs (right-skewed and left-skewed) as modulated by a bounded parameter $α\in (0, 1).$ The resulting posterior family in the single-parameter case can be expressed in closed-form and is proper, making calibration unnecessary. The mixing induced by the inverse stable family results to a marginal prior distribution in the form of a generalized Mittag-Leffler function, which covers a broad array of distributional shapes. We derive closed-form expressions of some properties like the moment generating function and moments. We propose algorithms to generate samples from the posterior distribution and calculate the Bayes estimators for real data analysis. We formulate the predictive prior and posterior distributions. We test the proposed Bayes estimators using Monte Carlo simulations. The extension to hierarchical modeling and inverse variance components models is straightforward. We can find $α$ (which acts like a smoothing parameter) values for which the inverse stable can provide better shrinkage than the inverted beta prior in many cases. We illustrate the methodology using a real data set, introduce a hyperprior density for the hyperparameters, and extend the model to a heavy-tailed distribution.

stat.ME

Inference for three-parameter M-Wright distributions with applications

We propose point estimators for the three-parameter (location, scale, and the fractional parameter) variant distributions generated by a Wright function. We also provide uncertainty quantification procedures for the proposed point estimators under certain conditions. The class of densities includes the three-parameter one-sided and the three-parameter symmetric bimodal $M$-Wright family of distributions. The one-sided family naturally generalizes the Airy and half-normal models. The symmetric class includes the symmetric Airy and normal or Gaussian densities. The proposed interval estimator for the scale parameter outperformed the estimator derived in \cite{cah12} when the location parameter is zero. We obtain the asymptotic covariance structure for the scale and fractional parameter estimators, which allows estimation of the correlation. The coverage probabilities of the interval estimators slightly depend on the proposed location parameter estimators. For the symmetric case, the sample mean (or median) is favored than the median (or mean) when the fractional parameter is greater (or lesser) than 0.39106 in terms of their asymptotic relative efficiency. The estimation algorithms were tested using synthetic data and were compared with their bootstrap counterparts. The proposed inference procedures were demonstrated on age and height data.

stat.ME

An estimation procedure for the Linnik distribution

We propose estimators for the parameters of the Linnik L$(α,γ)$ distribution. The estimators are derived from the moments of the log-transformed Linnik distributed random variable, and are shown to be asymptotically unbiased. The estimation algorithm is computationally simple and less restrictive. Our procedure is also tested using simulated data.

stat.ME