SearcharxivSearch

arXiv subjects

Landon Hurley

Publications and source records attributed to Landon Hurley.

7 recordsLinked to original sources

An exact unbiased semi-parametric L2 quasi-likelihood framework, complete in the presence of ties

Maximum likelihood style estimators possesses a number of ideal characteristics, but require prior identification of the distribution of errors to ensure exact unbiasedness. Independent of the focus of the primary statistical analysis, the estimation of a covariance matrix \(S^{P \times P}\approx \Sigma^{P \times P}\) must possess a specific structure and regularity constraints. The need to estimate a linear Gaussian covariance models appear in various applications as a formal precondition for scientific investigation and predictive analytics. In this work, we construct an \(\ell_{2}\)-norm based quasi-likelihood framework, identified by binomial comparisons between all pairs \(X_{n},Y_{n}, \forall {n}\). Our work here focuses upon the quasi-likelihood basis for estimation of an exactly unbiased linear regression H\'ajek projection, within which the Kemeny metric space is operationalised via Whitney embedding to obtain exact unbiased minimum variance multivariate covariance estimators upon both discrete and continuous random variables (i.e., exact unbiased identification in the presence of ties upon finite samples). While the covariance estimator is inherently useful, expansion of the Wilcoxon rank-sum testing framework to handle multiple covariates with exact unbiasedness upon finite samples is a currently unresolved research problem, as it maintains identification in the presence of linear surjective mappings onto common points: this model space, by definition, expands our likelihood framework into a consistent non-parametric form of the standard general linear model, which we extend to address both unknown heterogeneity and the problem of weak inferential instruments.

stat.ME

Completing and studentising Spearman's correlation in the presence of ties

Non-parametric correlation coefficients have been widely used for analysing arbitrary random variables upon common populations, when requiring an explicit error distribution to be known is an unacceptable assumption. We examine an \(\ell_{2}\) representation of a correlation coefficient (Emond and Mason, 2002) from the perspective of a statistical estimator upon random variables, and verify a number of interesting and highly desirable mathematical properties, mathematically similar to the Whitney embedding of a Hilbert space into the \(\ell_{2}\)-norm space. In particular, we show here that, in comparison to the traditional Spearman (1904) \(\rho\), the proposed Kemeny \(\rho_{\kappa}\) correlation coefficient satisfies Gauss-Markov conditions in the presence or absence of ties, thereby allowing both discrete and continuous marginal random variables. We also prove under standard regularity conditions a number of desirable scenarios, including the construction of a null hypothesis distribution which is Student-t distributed, parallel to standard practice with Pearson's r, but without requiring either continuous random variables nor particular Gaussian errors. Simulations in particular focus upon highly kurtotic data, with highly nominal empirical coverage consistent with theoretical expectation.

stat.ME

An exact unbiased semi-parametric maximum quasi-likelihood framework which is complete in the presence of ties

This paper introduces a novel quasi-likelihood extension of the generalised Kendall \(\tau_{a}\) estimator, together with an extension of the Kemeny metric and its associated covariance and correlation forms. The central contribution is to show that the U-statistic structure of the proposed coefficient \(\tau_{\kappa}\) naturally induces a quasi-maximum likelihood estimation (QMLE) framework, yielding consistent Wald and likelihood ratio test statistics. The development builds on the uncentred correlation inner-product (Hilbert space) formulation of Emond and Mason (2002) and resolves the associated sub-Gaussian likelihood optimisation problem under the \(\ell_{2}\)-norm via an Edgeworth expansion of higher-order moments. The Kemeny covariance coefficient \(\tau_{\kappa}\) is derived within a novel likelihood framework for pairwise comparison-continuous random variables, enabling direct inference on population-level correlation between ranked or weakly ordered datasets. Unlike existing approaches that focus on marginal or pairwise summaries, the proposed framework supports sample-observed weak orderings and accommodates ties without information loss. Drawing parallels with Thurstone's Case V latent ordering model, we derive a quasi-likelihood-based tie model with analytic standard errors, generalising classical U-statistics. The framework applies to general continuous and discrete random variables and establishes formal equivalence to Bradley-Terry and Thurstone models, yielding a uniquely identified linear representation with both analytic and likelihood-based estimators.

stat.ME

Unbiased analytic non-parametric correlation estimators in the presence of ties

An inner-product Hilbert space formulation is defined over a domain of all permutations with ties upon the extended real line. We demonstrate this work to resolve the common first and second order biases found in the pervasive Kendall and Spearman non-parametric correlation estimators, while presenting as unbiased minimum variance (Gauss-Markov) estimators. We conclude by showing upon finite samples that a strictly sub-Gaussian probability distribution is to be preferred for the Kemeny $τ_κ$ and $ρ_κ$ estimators, allowing for the construction of expected Wald test statistics which are analytically consistent with the Gauss-Markov properties upon finite samples.

stat.ME

Studentising Kendall's Tau: U-Statistic Estimators and Bias Correction for a Generalised Rank Variance-Covariance framework

Kemeny (1959) introduced a topologically complete metric space to study ordinal random variables, particularly in the context of Condorcet's paradox and the measurability of ties. Building on this, Emond & Mason (2002) reformulated Kemeny's framework into a rank correlation coefficient by embedding the metric space into a Hilbert structure. This transformation enables the analysis of data under weak order-preserving transformations (monotonically non-decreasing) within a linear probabilistic framework. However, the statistical properties of this rank correlation estimator, such as bias, estimation variance, and Type I error rates, have not been thoroughly evaluated. In this paper, we derive and prove a complete U-statistic estimator in the presence of ties for Kemeny's \(\tau_{\kappa}\), addressing the positive bias introduced by tied ranks. We also introduce a consistent population standard error estimator. The null distribution of the test statistic is shown to follow a \(t_{(N-2)}\)-distribution. Simulation results demonstrate that the proposed method outperforms Kendall's \(\tau_{b}\), offering a more accurate and robust measure of ordinal association which is topologically complete upon standard linear models.

stat.ME

An unbiased non-parametric correlation estimator in the presence of ties

An inner-product Hilbert space formulation of the Kemeny distance is defined over the domain of all permutations with ties upon the extended real line, and results in an unbiased minimum variance (Gauss-Markov) correlation estimator upon a homogeneous i.i.d. sample. In this work, we construct and prove the necessary requirements to extend this linear topology for both Spearman's \(ρ\) and Kendall's \(τ_{b}\), showing both spaces to be both biased and inefficient upon practical data domains. A probability distribution is defined for the Kemeny \(τ_κ\) estimator, and a Studentisation adjustment for finite samples is provided as well. This work allows for a general purpose linear model duality to be identified as a unique consistent solution to many biased and unbiased estimation scenarios.

stat.ME

An unbiased minimum variance non-parametric analytic and likelihood estimator for discrete and continuous score spaces

This manuscript develops a general purpose inner-product norm for the Kendall \(τ\) and Spearman's \(ρ\), which operates as an unbiased MLE even in the presence of ties. We derive and prove the strict sub-Gaussianity of the Kemeny norm-space, thereby disproving conclusions developed by both \textcite{kendall1948} and \textcite{diaconis1977} as to the nature of the appropriate, finite sample, probability distribution and test statistics. A non-parametric MLE framework for all bivariate pairs is developed, thereby resolving an hypothesis of \textcite{olkin1994} concerning an exponential multivariate distribution for order statistics, by showing that for finite samples, the distribution is non-exponential. Non-parametric linear estimators are also constructed for the polychoric correlations and by extension, a linearly decomposable non-parametric multidimensional linear system of equations for non-parametric Factor Analysis is shown.

stat.ME