SearcharxivSearch

arXiv subjects

Tom Kempton

Publications and source records attributed to Tom Kempton.

At least 19 recordsLinked to original sources

Fairness-Aware Test-Time Prompt Tuning

Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning. However, these models can also exhibit systematic biases that disproportionately affect protected demographic groups and existing approaches to addressing these biases require extensive model retraining and access to demographic attributes. There is a clear need to develop test-time adaptation (TTA) approaches that improve the fairness characteristics of pretrained models under distributional shift. In this paper, we evaluate how episodic TTA affects fairness in CLIP classification under subpopulation shifts and develop FairTPT, a novel fairness-aware episodic TTA method that jointly minimizes target marginal entropy while maximizing spurious marginal entropy through soft-prompt tuning. We find that standard episodic TTA generally exacerbates disparities between majority and minority groups, that blinding a model to spurious attributes without degrading target performance is inherently challenging, and that excessive blinding can lead to catastrophic forgetting. This model collapse can be prevented by monitoring test-time changes in target loss within the linear regime, while still achieving fairness improvements on reactive data and preserving overall performance. FairTPT outperforms all state-of-the-art episodic test-time debiasing methods and establishes a foundation for robust TTA, which is essential for achieving fairness in practice.

cs.LG

Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text

The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text should appear more probable to a detector language model than human-written text. However, we demonstrate that the token-level signal distinguishing human and machine text is non-uniform across the hidden space of the detector model, and naively averaging likelihood-based token scores across regions with fundamentally different statistical structure, as most detectors do, causes a form of Simpson's paradox: a strong local signal is destroyed by inappropriate aggregation. To correct for this, we introduce a learned local calibration step grounded in Bayesian decision theory. Rather than aggregating raw token scores, we first learn lightweight predictors of the score distributions conditioned on position in hidden space, and aggregate calibrated log-likelihood ratios instead. This single intervention dramatically and consistently improves detection performance across all baseline detectors and all datasets we consider. For example, our calibrated variant of Fast-DetectGPT improves AUROC from $0.63$ to $0.85$ on GPT-5.4 text, and a locally-calibrated DMAP detector we introduce achieves state-of-the-art performance across the board. That said, our central contribution is not a new detector, but a precise diagnosis of a significant cause of under-performance of existing detectors and a principled, modular remedy compatible with any token-averaging pipeline. This will serve as a foundation for the community to build upon, with natural avenues including richer distributional models, improved calibration strategies, and principled ensembling with hidden-space geometry signals via the full Bayes-optimal decision rule.

cs.CL

DMAP: A Distribution Map for Text

Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Extracting this signal typically relies on metrics such as perplexity, which do not adequately account for context; how one should interpret a given next-token probability is dependent on the number of reasonable choices encoded by the shape of the conditional distribution. In this work, we present DMAP, a mathematically grounded method that maps a text, via a language model, to a set of samples in the unit interval that jointly encode rank and probability information. This representation enables efficient, model-agnostic analysis and supports a range of applications. We illustrate its utility through three case studies: (i) validation of generation parameters to ensure data integrity, (ii) examining the role of probability curvature in machine-generated text detection, and (iii) a forensic analysis revealing statistical fingerprints left in downstream models that have been subject to post-training on synthetic data. Our results demonstrate that DMAP offers a unified statistical view of text that is simple to compute on consumer hardware, widely applicable, and provides a foundation for further research into text analysis with LLMs.

cs.CL

Emergent Bias and Fairness in Multi-Agent Decision Systems

Multi-agent systems have demonstrated the ability to improve performance on a variety of predictive tasks by leveraging collaborative decision making. However, the lack of effective evaluation methodologies has made it difficult to estimate the risk of bias, making deployment of such systems unsafe in high stakes domains such as consumer finance, where biased decisions can translate directly into regulatory breaches and financial loss. To address this challenge, we need to develop fairness evaluation methodologies for multi-agent predictive systems and measure the fairness characteristics of these systems in the financial tabular domain. Examining fairness metrics using large-scale simulations across diverse multi-agent configurations, with varying communication and collaboration mechanisms, we reveal patterns of emergent bias in financial decision-making that cannot be traced to individual agent components, indicating that multi-agent systems may exhibit genuinely collective behaviors. Our findings highlight that fairness risks in financial multi-agent systems represent a significant component of model risk, with tangible impacts on tasks such as credit scoring and income estimation. We advocate that multi-agent decision systems must be evaluated as holistic entities rather than through reductionist analyses of their constituent components.

cs.LG

Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models

Advances in hardware and language model architecture have spurred a revolution in natural language generation. However, autoregressive models compute probability distributions over next-token choices, and sampling from these distributions, known as decoding, has received significantly less attention than other design choices. Existing decoding strategies are largely based on heuristics, resulting in methods that are difficult to apply or improve in a principled manner. We develop the theory of decoding strategies for language models by expressing popular decoding algorithms as equilibrium states in the language of ergodic theory and stating the objective functions they optimize. Using this, we analyze the effect of the local normalization step required to make probabilities sum to one in top-k, nucleus, and temperature sampling. We argue that local normalization distortion is a fundamental defect of decoding strategies and quantify the size of this distortion and its effect on mathematical proxies for the quality and diversity of generated text. This yields conclusions for the design of decoding algorithms and the detection of machine-generated text.

cs.CL

TempTest: Local Normalization Distortion and the Detection of Machine-generated Text

Existing methods for the zero-shot detection of machine-generated text are dominated by three statistical quantities: log-likelihood, log-rank, and entropy. As language models mimic the distribution of human text ever closer, this will limit our ability to build effective detection algorithms. To combat this, we introduce a method for detecting machine-generated text that is entirely agnostic of the generating language model. This is achieved by targeting a defect in the way that decoding strategies, such as temperature or top-k sampling, normalize conditional probability measures. This method can be rigorously theoretically justified, is easily explainable, and is conceptually distinct from existing methods for detecting machine-generated text. We evaluate our detector in the white and black box settings across various language models, datasets, and passage lengths. We also study the effect of paraphrasing attacks on our detector and the extent to which it is biased against non-native speakers. In each of these settings, the performance of our test is at least comparable to that of other state-of-the-art text detectors, and in some cases, we strongly outperform these baselines.

cs.CL

Local dimension spectrum for dominated planar self-affine sets

The local dimension spectrum provides a framework for quantifying the fractal properties of a measure, and it is well understood for non-overlapping self-similar measures. In this article, we study the local dimension spectrum for dominated self-affine measures. By analyzing exact dimensionality, we obtain deterministic results that extend the scope of the local dimension spectrum beyond the almost-sure setting.

math.DS

Towards Absolutely Continuous Bernoulli Convolutions

We show how to turn the question of the absolute continuity of Bernoulli convolutions into one of counting the growth of the number of overlaps in the system. When the contraction parameter is a hyperbolic algebraic integer, we turn this question of absolute continuity into a question involving the ergodic theory of cocycles over domain exchange transformations.

math.DS

Measures on the Spectra of Algebraic Integers

Given a real number beta > 1, the spectrum of beta is a well studied dynamical object. In this article we show the existence of a certain measure on the spectrum of beta related to the distribution of random polynomials in beta, and discuss the local structure of this measure. We also make links with the question of the Hausdorff dimension of the corresponding Bernoulli Convolution

math.DS

Computing Garsia Entropy for Bernoulli Convolutions with Algebraic Parameters

We introduce a parameter space containing all algebraic integers $\beta\in(1,2]$ that are not Pisot or Salem numbers, and a sequence of increasing piecewise continuous function on this parameter space which gives a lower bound for the Garsia entropy of the Bernoulli convolution $\nu_{\beta}$. This allows us to show that $\mathrm{dim}_\mathrm{H} (\nu_{\beta})=1$ for all $\beta$ with representations in certain open regions of the parameter space.

math.CA

Intermediate dimensions

We introduce a continuum of dimensions which are `intermediate' between the familiar Hausdorff and box dimensions. This is done by restricting the families of allowable covers in the definition of Hausdorff dimension by insisting that $|U| \leq |V|^\theta$ for all sets $U, V$ used in a particular cover, where $\theta \in [0,1]$ is a parameter. Thus, when $\theta=1$ only covers using sets of the same size are allowable, and we recover the box dimensions, and when $\theta=0$ there are no restrictions, and we recover Hausdorff dimension. We investigate many properties of the intermediate dimension (as a function of $\theta$), including proving that it is continuous on $(0,1]$ but not necessarily continuous at $0$, as well as establishing appropriate analogues of the mass distribution principle, Frostman's lemma, and the dimension formulae for products. We also compute, or estimate, the intermediate dimensions of some familiar sets, including sequences formed by negative powers of integers, and Bedford-McMullen carpets.

math.MG

On the Hausdorff Dimension of Bernoulli Convolutions

We give an expression for the Garsia entropy of Bernoulli convolutions in terms of products of matrices. This gives an explicit rate of convergence of the Garsia entropy and shows that one can calculate the Hausdorff dimension of the Bernoulli convolution $\nu_\beta$ to arbitrary given accuracy whenever $\beta$ is algebraic. In particular, if the Garsia entropy $H(\beta)$ is not equal to $\log(\beta)$ then we have a finite time algorithm to determine whether or not $\mathrm{dim}_\mathrm{H} (\nu_\beta)=1$.

math.CA

On the L^q Dimensions of Measures on Hueter-Lalley Type Self-Affine Sets

We study the L^q -dimensions of self-affine measures and the Kaenmaki measure on a class of self-affine sets similar to the class considered by Hueter and Lalley. We give simple, checkable conditions under which the Lq -dimensions are equal to the value predicted by Falconer for a range of q. As a corollary this gives a wider class of self-affine sets for which the Hausdorff dimension can be explicitly calculated. Our proof combines the potential theoretic approach developed by Hunt and Kaloshin with recent advances in the dynamics of self-affine sets.

math.DS

The dimension of projections of self-affine sets and measures

Let E be a plane self-affine set defined by affine transformations with linear parts given by matrices with positive entries. We show that if mu is a Bernoulli measure on E with dim_H mu = dim_L mu, where dim_H and dim_L denote Hausdorff and Lyapunov dimensions, then the projection of mu in all but at most one direction has Hausdorff dimension min{dim_H mu,1}. We transfer this result to sets and show that many self-affine sets have projections of dimension min{dim_H E,1} in all but at most one direction.

math.DS

The random continued fraction transformation

We introduce a random dynamical system related to continued fraction expansions. It uses random combination of the Gauss map and the R\'enyi (or backwards) continued fraction map. We explore the continued fraction expansions that this system produces as well as the dynamical properties of the system.

math.DS

The Scenery Flow for Self-Affine Measures

We describe the scaling scenery associated to Bernoulli measures supported on separated self-affine sets under the condition that certain projections of the measure are absolutely continuous.

math.DS

Planar self-affine sets with equal Hausdorff, box and affinity dimensions

Using methods from ergodic theory along with properties of the Furstenberg measure we obtain conditions under which certain classes of plane self-affine sets have Hausdorff or box-counting dimensions equal to their affinity dimension. We exhibit some new specific classes of self-affine sets for which these dimensions are equal.

math.DS