SearcharxivSearch

arXiv subjects

Dennis Christensen

Publications and source records attributed to Dennis Christensen.

8 recordsLinked to original sources

Posterior uncertainty for kernel density estimates

Recent work in predictive Bayesian inference has enabled novel Bayesian interpretations of many well-known stochastic one-step-ahead predictive algorithms. In this paper, we study classic kernel density estimation in the predictive Bayesian framework. We prove that their predictive measures converge weakly almost surely $\unicode{x2013}$ meaning that their associated predictive resampling sequences are almost surely asymptotically exchangeable $\unicode{x2013}$ and we provide estimators for moments of the limiting random probability measure. We also show that the resampling sequences do not satisfy standard assumptions like being conditionally identically distributed (c.i.d.) or almost c.i.d. (a.c.i.d.), thus providing a non-trivial example of a predictive sequence which is not a.c.i.d. but nevertheless converges weakly almost surely. For Gaussian kernels, we show that the limiting directing measure is almost surely absolutely continuous with respect to the Lebesgue measure, meaning it emits a probability density. This enables us to derive credibility intervals for kernel density estimates, which we illustrate on two real datasets.

stat.ME

Computing Conditional Shapley Values Using Tabular Foundation Models

Shapley values have become a cornerstone of explainable AI, but they are computationally expensive to use, especially when features are dependent. Evaluating them requires approximating a large number of conditional expectations, either via Monte Carlo integration or regression. Until recently it has not been possible to fully exploit deep learning for the regression approach, because retraining for each conditional expectation takes too long. Tabular foundation models such as TabPFN overcome this computational hurdle by leveraging in-context learning, so each conditional expectation can be approximated without any re-training. In this paper, we compute Shapley values with multiple variants of TabPFN and compare their performance with state-of-the-art methods on both simulated and real datasets. In most cases, TabPFN yields the best performance; where it does not, it is only marginally worse than the best method, at a fraction of the runtime. We discuss further improvements and how tabular foundation models can be better adapted specifically for conditional Shapley value estimation.

cs.AI

Stop using limiting stimuli as a measure of sensitivities of energetic materials

Accurately estimating the sensitivity of explosive materials is a potentially life-saving task which requires standardised protocols across nations. One of the most widely applied procedures worldwide is the so-called '1-In-6' test from the United Nations (UN) Manual of Tests in Criteria, which estimates a 'limiting stimulus' for a material. In this paper we demonstrate that, despite their popularity, limiting stimuli are not a well-defined notion of sensitivity and do not provide reliable information about a material's susceptibility to ignition. In particular, they do not permit construction of confidence intervals to quantify estimation uncertainty. We show that continued reliance on limiting stimuli through the 1-In-6 test has caused needless confusion in energetic materials research, both in theoretical studies and practical safety applications. To remedy this problem, we consider three well-founded alternative approaches to sensitivity testing to replace limiting stimulus estimation. We compare their performance in an extensive simulation study and apply the best-performing approach to real data, estimating the friction sensitivity of pentaerythritol tetranitrate (PETN).

stat.AP

Random irregular histograms

We propose a new method of histogram construction, providing a fully Bayesian approach to irregular histograms. Our procedure applies Bayesian model selection to a piecewise constant model of the underlying distribution, resulting in a method that selects both the number of bins as well as their location based on the data in a fully automatic fashion. We show that the histogram estimate is consistent with respect to the Hellinger metric under mild regularity conditions, and that it attains a convergence rate equal to the minimax rate (up to a logarithmic factor) for H\"{o}lder continuous densities. Simulation studies indicate that the new method performs comparably to other histogram procedures, both for minimizing the estimation error and for identifying modes. A software implementation is included as supplementary material.

stat.ME

Asymptotic properties of adaptive designs through differentiability in quadratic mean

There exist multiple regression applications in engineering, industry and medicine where the outcomes follow an adaptive experimental design in which the next measurement depends on the previous observations, so that the observations are not conditionally independent given the covariates. In the existing literature on such adaptive designs, results asserting asymptotic normality of the maximum likelihood estimator require regularity conditions involving the second or third derivatives of the log-likelihood. Here we instead extend the theory of differentiability in quadratic mean (DQM) to the setting of adaptive designs, which requires strictly fewer regularity assumptions than the classical theory. In doing so, we discover a new DQM assumption, which we call summable differentiability in quadratic mean (S-DQM). As applications, we first verify asymptotic normality for two classical adaptive designs, namely the Bruceton 'up-and-down' design and the Robbins-Monro design. Next, we consider a more complicated problem, namely a Markovian version of the Langlie design.

math.ST

Revisiting the Langlie procedure

Introduced in 1962, the Langlie procedure is one of the most popular approaches to sensitivity testing. It aims to estimate an unknown sensitivity distribution based on the outcomes of binary trials. Officially recognized by the U.S. Department of Defense, the procedure is widely used both in civil and military industry. It first provides an experimental design for how the binary trials should be conducted, and then estimates the sensitivity distribution via maximum likelihood under a simple parametric model like logistic or probit regression. Despite its popularity and longevity, little is known about the statistical properties of the Langlie procedure, but it is well-established that the sequence of inputs tend to narrow in on the median of the sensitivity distribution. For this reason, the U.S. Department of Defense's protocol dictates that the procedure is only appropriate for estimating the median of the distribution, and no other quantiles. This begs the question of whether the parametric model assumption can be disposed of altogether, potentially making the Langlie procedure entirely nonparametric, much like the Robbins-Monro procedure. In this paper we answer this question in the negative by proving that when the Langlie procedure is employed, the sequence of inputs converges with probability zero.

math.ST

perms: Likelihood-free estimation of marginal likelihoods for binary response data in Python and R

In Bayesian statistics, the marginal likelihood (ML) is the key ingredient needed for model comparison and model averaging. Unfortunately, estimating MLs accurately is notoriously difficult, especially for models where posterior simulation is not possible. Recently, Christensen (2023) introduced the concept of permutation counting, which can accurately estimate MLs of models for exchangeable binary responses. Such data arise in a multitude of statistical problems, including binary classification, bioassay and sensitivity testing. Permutation counting is entirely likelihood-free and works for any model from which a random sample can be generated, including nonparametric models. Here we present perms, a package implementing permutation counting. As a result of extensive optimisation efforts, perms is computationally efficient and able to handle large data problems. It is available as both an R package and a Python library. A broad gallery of examples illustrating its usage is provided, which includes both standard parametric binary classification and novel applications of nonparametric models, such as changepoint analysis. We also cover the details of the implementation of perms and illustrate its computational speed via a simple simulation study.

stat.CO

Visualizing thickness-dependent magnetic textures in few-layer $\text{Cr}_2\text{Ge}_2\text{Te}_6$

Magnetic ordering in two-dimensional (2D) materials has recently emerged as a promising platform for data storage, computing, and sensing. To advance these developments, it is vital to gain a detailed understanding of how the magnetic order evolves on the nanometer-scale as a function of the number of atomic layers and applied magnetic field. Here, we image few-layer $\text{Cr}_2\text{Ge}_2\text{Te}_6$ using a combined scanning superconducting quantum interference device and atomic force microscopy probe. Maps of the material's stray magnetic field as a function of applied magnetic field reveal its magnetization per layer as well as the thickness-dependent magnetic texture. Using a micromagnetic model, we correlate measured stray-field patterns with the underlying magnetization configurations, including labyrinth domains and skyrmionic bubbles. Comparison between real-space images and simulations demonstrates that the layer dependence of the material's magnetic texture is a result of the thickness-dependent balance between crystalline and shape anisotropy. These findings represent an important step towards 2D spintronic devices with engineered spin configurations and controlled dependence on external magnetic fields.

cond-mat.mes-hall