SearcharxivSearch

arXiv subjects

Jan Hannig

Publications and source records attributed to Jan Hannig.

At least 19 recordsLinked to original sources

Constrained Fiducial Inference for Gaussian Models

We propose a new fiducial Markov Chain Monte Carlo (MCMC) method for fitting parametric Gaussian models. We utilize the Cayley transform to decompose the parametric covariance matrix, which in turn allows us to formulate a general data generating algorithm for Gaussian data. Leveraging constrained generalized fiducial inference, we are able to create the basis of an MCMC algorithm, which can be specified to parametric models with minimal effort. The appeal of this novel approach is the wide class of models which it permits, ease of implementation and the posterior-like fiducial distribution without the need for a prior. We provide background information for the derivation of the relevant fiducial quantities, and a proof that the proposed MCMC algorithm targets the correct fiducial distribution. We need not assume independence nor identical distribution of the data, which makes the method attractive for application to time series and spatial data. Well-performing simulation results of the MA(1) and Mat\'ern models are presented.

stat.ME

Calibrating Bayesian Inference

Bayesian statistics has gained popularity in psychological research due to its intuitive uncertainty quantification and convenient information-updating rules. In many applications, however, prior distributions are introduced merely as instruments to facilitate computation, rather than as representations of genuine subjective belief. Consequently, relying on standard Bayesian justifications for inferential procedures becomes conceptually ungrounded. In this paper, we recommend evaluating finite-sample performance over repeated sampling of data and parameters as an alternative justification for "pragmatic Bayes." We demonstrate a key vulnerability in the usual posterior-based inference: when analysts' chosen prior distribution mismatches the true parameter-generating process, Bayesian inference can be misleading. Given that this true process is rarely known in practice, we propose a safer alternative: calibrating Bayesian credible regions to achieve frequentist validity. This latter criterion is stronger and guarantees validity of Bayesian inference regardless of the underlying parameter-generating mechanism. To solve the calibration problem in practice, we propose a novel stochastic approximation algorithm. A Monte Carlo experiment is conducted and reported, in which we observe that uncalibrated Bayesian inference can be liberal under certain parameter-generating scenarios, whereas our calibrated solution consistently maintain validity. We also illustrate the proposed calibration procedure using a real-data example involving location-scale regression.

stat.ME

Advanced Distribution Theory for Significance in Scale Space

Smoothing methods find signals in noisy data. A challenge for Statistical inference is the choice of smoothing parameter. SiZer addressed this challenge in one-dimension by detecting significant slopes across multiple scales, but was not a completely valid testing procedure. This was addressed by the development of an advanced distribution theory that ensures fully valid inference in the 1-D setting by applying extreme value theory. A two-dimensional extension of SiZer, known as Significance in Scale Space (SSS), was developed for image data, enabling the detection of both slopes and curvatures across multiple spatial scales. However, fully valid inference for 2-D SSS has remained unavailable, largely due to the more complex dependence structure of random fields. In this paper, we use a completely different probability methodology which gives an advanced distribution theory for SSS, establishing a valid hypothesis testing procedure for both slope and curvature detection. When applied to pure noise images (no true underlying signal), the proposed method controls the Type I error, whereas the original SSS identifies spurious features across scales. When signal is present, the proposed method maintains a high level of statistical power, successfully identifying important true slopes and curvatures in real data such as gamma camera images.

math.ST

Evaluating Gender Wage Inequality in Academia using Causal Inference Methods for Observational Data

Observational studies often present challenges for causal inference due to confounding and heterogeneity. In this paper, we illustrate how modern causal inference methods can be applied to large-scale academic salary data. Using records from 12,039 tenure-track faculty in the University of North Carolina system, linked with bibliometric indicators and institutional classifications, we estimate the causal effect of gender on faculty salaries. Our analysis combines propensity score matching with causal forests to adjust for rank, discipline, research productivity, and career experience. Results indicate that female faculty earn approximately 6% less than comparable male colleagues, with variation in the gap across career stages and levels of research productivity. This case study demonstrates how causal inference methods for observational data can provide insight into structural disparities in complex social systems.

stat.AP

Bayesian Forensic DNA Mixture Deconvolution Using a Novel String Similarity Measure

Mixture interpretation is a central challenge in forensic science, where evidence often contains contributions from multiple sources. In the context of DNA analysis, biological samples recovered from crime scenes may include genetic material from several individuals, necessitating robust statistical tools to assess whether a specific person of interest (POI) is among the contributors. Methods based on capillary electrophoresis (CE) are currently in use worldwide, but offer limited resolution in complex mixtures. Advancements in massively parallel sequencing (MPS) technologies provide a richer, more detailed representation of DNA mixtures, but require new analytical strategies to fully leverage this information. In this work, we present a Bayesian framework for evaluating whether a POIs DNA is present in an MPS-based forensic sample. The model accommodates known contributors, such as the victim, and uses a novel string edit distance to quantify similarity between observed alleles and sequencing artifacts. The resulting Bayes factors enable effective discrimination between samples that do and do not contain the POIs DNA, demonstrating strong performance in both hypothesis testing and classification settings.

stat.ME

Movement Dynamics in Elite Female Soccer Athletes: The Quantile Cube Approach

This paper presents the quantile cube, a novel three-dimensional summary representation designed to analyze external load using GPS-derived movement data. While broadly applicable, we demonstrate its utility through an application to data from elite female soccer athletes across 23 matches. The quantile cube segments athlete movements into discrete quantiles of velocity, acceleration, and movement angle across match halves, providing a structured and interpretable framework to capture complex movement dynamics. Statistical analysis revealed significant differences in movement distributions between the first and second halves for individual athletes across all matches. Principal Component Analysis identified matches with unique movement dynamics, particularly at the start and end of the season. Dirichlet-multinomial regression further explored how factors such as athlete position, playing time, and match characteristics influenced movement profiles. Our analysis reveals external load variations over time and provides insights into performance optimization. The integration of these statistical techniques demonstrates the potential of data-driven strategies to enhance athlete monitoring and workload management in women's soccer.

stat.ME

A Generalized Fiducial Framework for Partially Identified Treatment Effects

Over the past two decades, there has been renewed interest in fiducial inference, a statistical framework originally proposed by R. A. Fisher in the 1930s. Existing contributions, however, have largely focused on point-identified models, where the parameter vector is uniquely determined by the joint distribution of the observables and the model assumptions. This paper develops a unified fiducial inference procedure for partially identified models, thereby extending the applicability of the framework to settings in which the identified set is non-degenerate. The acceptance rate of the proposed sampler naturally serves as a diagnostic: high values support the identifying assumptions, while near-zero values signal their potential violation. We establish a Bernstein-von Mises theorem for the fiducial distribution, which provides theoretical guarantees for the proposed sampler. As two leading examples, we provide uncertainty quantification for the instrumental variable model and the mediation model under a range of causal assumptions and target parameters. The proposed methodology is illustrated through extensive simulations and empirical applications.

stat.ME

Elastic Shape Analysis of Movement Data

Osteoarthritis (OA) is a highly prevalent degenerative joint disease, and the knee is the most commonly affected joint. Biomechanical factors, particularly forces exerted during walking, are often measured in modern studies of knee joint injury and OA, and understanding the relationship among biomechanics, clinical profiles, and OA has high clinical relevance. Biomechanical forces are typically represented as curves over time, but a standard practice in biomechanics research is to summarize these curves by a small number of discrete values (or landmarks). The objective of this work is to demonstrate the added value of analyzing full movement curves over conventional discrete summaries. We developed a shape-based representation of variation in full biomechanical curve data from the Intensive Diet and Exercise for Arthritis (IDEA) study (Messier et al., 2009, 2013), and demonstrated through nested model comparisons that our approach, compared to conventional discrete summaries, yields stronger associations with OA severity and OA-related clinical traits. Notably, our work is among the first to quantitatively evaluate the added value of analyzing full movement curves over conventional discrete summaries.

stat.AP

Multi-faceted Neuroimaging Data Integration via Analysis of Subspaces

Neuroimaging studies, such as the Human Connectome Project (HCP), often collect multi-faceted and multi-block data to study the complex human brain. However, these data are often analyzed in a pairwise fashion, which can hinder our understanding of how different brain-related measures interact with each other. In this study, we comprehensively analyze the multi-block HCP data using the Data Integration via Analysis of Subspaces (DIVAS) method. We integrate structural and functional brain connectivity, substance use, cognition, and genetics in an exhaustive five-block analysis. This gives rise to the important finding that genetics is the single data modality most predictive of brain connectivity, outside of brain connectivity itself. Nearly 14\% of the variation in functional connectivity (FC) and roughly 12\% of the variation in structural connectivity (SC) is attributed to shared spaces with genetics. Moreover, investigations of shared space loadings provide interpretable associations between particular brain regions and drivers of variability, such as alcohol consumption in the substance-use data block. Novel Jackstraw hypothesis tests are developed for the DIVAS framework to establish statistically significant loadings. For example, in the (FC, SC, and Substance Use) shared space, these novel hypothesis tests highlight largely negative functional and structural connections suggesting the brain's role in physiological responses to increased substance use. Furthermore, our findings have been validated using a subset of genetically relevant siblings or twins not studied in the main analysis.

q-bio.NC

Asymptotic Theory for Estimation of the Husler-Reiss Distribution via Block Maxima Method

The H\"usler-Reiss distribution describes the limit of the pointwise maxima of a bivariate normal distribution. This distribution is defined by a single parameter, $\lambda$. We provide asymptotic theory for maximum likelihood estimation of $\lambda$ under a block maxima approach. Our work assumes independent and identically distributed bivariate normal random variables, grouped into blocks where the block size and number of blocks increase simultaneously. With these assumptions our results provide conditions for the asymptotic normality of the Maximum Likelihood Estimator (MLE). We characterize the bias of the MLE, provide conditions under which this bias is asymptotically negligible, and discuss how to choose the block size to minimize a bias-variance trade-off. The proofs are an extension of previous results for choosing the block size in the estimation of univariate extreme value distributions (Dombry and Ferreria 2019), providing a potential basis for extensions to multivariate cases where both the marginal and dependence parameters are unknown. The proofs rely on the Argmax Theorem applied to a localized loglikelihood function, combined with a Lindeberg-Feller Central Limit Theorem argument to establish asymptotic normality. Possible applications of the method include composite likelihood estimation in Brown-Resnick processes, where it is known that the bivariate distributions are of H\"usler-Reiss form.

math.ST

Semiparametric fiducial inference for Cox models

R. A. Fisher introduced the fiducial distribution as a potential replacement for the Bayesian posterior distribution in the 1930s. During the past century, fiducial approaches have been explored in various parametric and nonparametric settings. However, to the best of our knowledge, no fiducial inference has been developed in the realm of semiparametric statistics. In this paper, we propose a novel fiducial approach for semiparametric models. In memory of Sir David Cox who passed away in 2022, we use the Cox proportional hazards model, which is the most popular model for the analysis of survival data, as a running example. Other models and extensions are also discussed. In our experiments, we find that our method performs particularly well in situations where the maximum likelihood estimator fails.

stat.ME

AutoGFI: Streamlined Generalized Fiducial Inference for Modern Inference Problems in Models with Additive Errors

The concept of fiducial inference was introduced by R. A. Fisher in the 1930s to address the perceived limitations of Bayesian inference, particularly the need for subjective prior distributions in cases with limited prior information. However, Fisher's fiducial approach lost favor due to complications, especially in multi-parameter problems. With renewed interest in fiducial inference in the 2000s, generalized fiducial inference (GFI) emerged as a promising extension of Fisher's ideas, offering new solutions for complex inference challenges. Despite its potential, GFI's adoption has been hindered by demanding mathematical derivations and complex implementation requirements, such as Markov Chain Monte Carlo (MCMC) algorithms. This paper introduces AutoGFI, a streamlined variant of GFI designed to simplify its application across various inference problems with additive noise. AutoGFI's accessibility lies in its simplicity-requiring only a fitting routine-making it a feasible option for a wider range of researchers and practitioners. To demonstrate its efficacy, AutoGFI is applied to three challenging problems: tensor regression, matrix completion, and network cohesion regression. These case studies showcase AutoGFI's competitive performance against specialized solutions, highlighting its potential to broaden the application of GFI in practical domains, ultimately enriching the statistical inference toolkit.

stat.ME

Dempster-Shafer P-values: Thoughts on an Alternative Approach for Multinomial Inference

In this paper, we demonstrate that a new measure of evidence we developed called the Dempster-Shafer p-value which allow for insights and interpretations which retain most of the structure of the p-value while covering for some of the disadvantages that traditional p- values face. Moreover, we show through classical large-sample bounds and simulations that there exists a close connection between our form of DS hypothesis testing and the classical frequentist testing paradigm. We also demonstrate how our approach gives unique insights into the dimensionality of a hypothesis test, as well as models the effects of adversarial attacks on multinomial data. Finally, we demonstrate how these insights can be used to analyze text data for public health through an analysis of the Population Health Metrics Research Consortium dataset for verbal autopsies.

stat.ME

A Bernstein-von Mises Theorem for Generalized Fiducial Distributions

An established and growing literature on generalized fiducial inference and related fiducial ideas points to the adoption of fiducial inference as a mainstream perspective among modern statisticians. Like Bayesian posteriors, generalized fiducial distributions (GFDs) are known to satisfy Bernstein-von Mises (BvM)-type results under classical regularity conditions. Existing fiducial BvM results, however, rely on relatively restrictive smoothness assumptions and are limited in scope. In this paper, we establish a Bernstein-von Mises theorem for generalized fiducial inference under the general framework of local asymptotic normality, which accommodates non-i.i.d. data settings and reduces to the familiar differentiability in quadratic mean condition in the i.i.d. case. We apply our result to extend existing fiducial theory for free-knot spline models first developed in Sonderegger and Hannig (2014), and further illustrate its generality in models where classical regularity conditions fail or i.i.d. assumptions are not met.

math.ST

Bayes Watch: Bayesian Change-point Detection for Process Monitoring with Fault Detection

When a predictive model is in production, it must be monitored in real-time to ensure that its performance does not suffer due to drift or abrupt changes to data. Ideally, this is done long before learning that the performance of the model itself has dropped by monitoring outcome data. In this paper we consider the problem of monitoring a predictive model that identifies the need for palliative care currently in production at the Mayo Clinic in Rochester, MN. We introduce a framework, called \textit{Bayes Watch}, for detecting change-points in high-dimensional longitudinal data with mixed variable types and missing values and for determining in which variables the change-point occurred. Bayes Watch fits an array of Gaussian Graphical Mixture Models to groupings of homogeneous data in time, called regimes, which are modeled as the observed states of a Markov process with unknown transition probabilities. In doing so, Bayes Watch defines a posterior distribution on a vector of regime assignments, which gives meaningful expressions on the probability of every possible change-point. Bayes Watch also allows for an effective and efficient fault detection system that assesses what features in the data where the most responsible for a given change-point.

stat.AP

Introduction to Generalized Fiducial Inference

Fiducial inference was introduced in the first half of the 20th century by Fisher (1935) as a means to get a posterior-like distribution for a parameter without having to arbitrarily define a prior. While the method originally fell out of favor due to non-exactness issues in multivariate cases, the method has garnered renewed interest in the last decade. This is partly due to the development of generalized fiducial inference, which is a fiducial perspective on generalized confidence intervals: a method used to find approximate confidence distributions. In this chapter, we illuminate the usefulness of the fiducial philosophy, introduce the definition of a generalized fiducial distribution, and apply it to interesting, non-trivial inferential examples.

stat.ME

A fiducial approach to nonparametric deconvolution problem: discrete case

Fiducial inference, as generalized by Hannig et al. (2016), is applied to nonparametric g-modeling (Efron, 2016) in the discrete case. We propose a computationally efficient algorithm to sample from the fiducial distribution, and use the generated samples to construct point estimates and confidence intervals. We study the theoretical properties of the fiducial distribution and perform extensive simulations in various scenarios. The proposed approach yields good statistical performance in terms of the mean squared error of point estimators and the coverage of confidence intervals. Furthermore, we apply the proposed fiducial method to estimate the probability of each satellite site being malignant using gastric adenocarcinoma data with 844 patients (Efron, 2016).

stat.ME

Generalized Fiducial Inference on Differentiable Manifolds

We introduce a novel approach to inference on parameters that take values in a Riemannian manifold embedded in a Euclidean space. Parameter spaces of this form are ubiquitous across many fields, including chemistry, physics, computer graphics, and geology. This new approach uses generalized fiducial inference to obtain a posterior-like distribution on the manifold, without needing to know a parameterization that maps the constrained space to an unconstrained Euclidean space. The proposed methodology, called the constrained generalized fiducial distribution (CGFD), is obtained by using mathematical tools from Riemannian geometry. A Bernstein-von Mises-type result for the CGFD, which provides intuition for how the desirable asymptotic qualities of the unconstrained generalized fiducial distribution are inherited by the CGFD, is provided. To demonstrate the practical use of the CGFD, we provide three proof-of-concept examples: inference for data from a multivariate normal density with the mean parameters on a sphere, a linear logspline density estimation problem, and a reimagined approach to the AR(1) model, all of which exhibit desirable coverages via simulation. We discuss two Markov chain Monte Carlo algorithms for the exploration of these constrained parameter spaces and adapt them for the CGFD.

stat.ME