SearcharxivSearch

arXiv subjects

Nicholas Barnfield

Publications and source records attributed to Nicholas Barnfield.

10 recordsLinked to original sources

Sharp Capacity Thresholds in Linear Associative Memory: From Top-1 Retrieval to Tail-Average Learning

How many key-value associations can a $d\times d$ linear memory store? The answer depends not only on the $d^2$ degrees of freedom in the memory matrix, but also on the retrieval criterion. Under isotropic Gaussian embeddings, we prove a sharp threshold for top-1 retrieval, where every signal must beat its largest distractor: the critical value of $d^2/(n\log n)$ is $2$. Above the threshold, we explicitly construct a linear memory that retrieves all $n$ associations with high probability; below it, no data-dependent linear memory can do so. The $\log n$ factor is therefore the unavoidable extreme-value cost of winner-take-all decoding. Without the logarithmic factor---that is, when $n/d^2\to\alpha\in(0,\infty)$---simultaneous top-1 retrieval is impossible. The matched target can nevertheless remain near the top of the ranking. We capture this weaker retrieval goal with the Tail-Average Margin (TAM), which, for list size $k$, compares each signal with the average of its $k$ strongest competitors; a positive TAM margin certifies that the target belongs to the top-$k$ candidate list. When $k/n\to r\in(0,1)$, we learn the memory by empirical risk minimization with a smoothed TAM objective and derive an exact high-dimensional characterization through a two-parameter scalar variational problem. The result gives limiting laws for signal and competitor scores, margins, and percentile ranks. Sending the ridge parameter to zero after the high-dimensional limit yields a closed-form critical load $\alpha_c(r)$ separating vanishing from positive average loss. The analysis in this work also develops a coupled leave-one-out method for matrix-valued empirical risk problems in which each sample enters many dependent comparisons, a tool that may be useful beyond associative memory.

stat.ML

A Scale-Shape Dual Newton Method for Entropic Least Squares

We give a damped inexact Newton method for entropy-regularized least-squares on the nonnegative orthant that converges globally at a linear rate with $O(\log\epsilon^{-1})$ iteration complexity, locally at a superlinear-to-quadratic rate, and is immune to the finite-precision overflow that limits classical dual solvers. A scale-shape decomposition of the primal -- separating its scale from its direction -- produces a dual with a nonsingular Jacobian. Objectives and Jacobians are evaluated through stable log-sum-exp and softmax primitives. Lambert W bounds on the scale uniformly control the Jacobian's spectrum, from which both rates follow. The solution map is jointly Lipschitz in the data, regularization parameter, and reference measure, and extends continuously to the vanishing-regularization limit. Experiments on a problem from analytic continuation of quantum Monte Carlo data confirm the predicted overflow resilience and convergence behavior.

math.OC

Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning

Recent progress has rapidly advanced our understanding of the mechanisms underlying in-context learning in modern attention-based neural networks. However, existing results focus exclusively on unimodal data; in contrast, the theoretical underpinnings of in-context learning for multi-modal data remain poorly understood. We introduce a mathematically tractable framework for studying multi-modal learning and explore when transformer-like architectures can recover Bayes-optimal performance in-context. To model multi-modal problems, we assume the observed data arises from a latent factor model. Our first result comprises a negative take on expressibility: we prove that single-layer, linear self-attention fails to recover the Bayes-optimal predictor uniformly over the task distribution. To address this limitation, we introduce a novel, linearized cross-attention mechanism, which we study in the regime where both the number of cross-attention layers and the context length are large. We show that this cross-attention mechanism is provably Bayes optimal when optimized using gradient flow. Our results underscore the benefits of depth for in-context learning and establish the provable utility of cross-attention for multi-modal distributions.

stat.ML

The noiseless limit and improved-prior limit of the maximum entropy method and their implications for the analytic continuation problem

Quantum Monte Carlo (QMC) methods are uniquely capable of providing exact simulations of quantum many-body systems. Unfortunately, the applications of a QMC simulation are limited because extracting dynamic properties requires solving the analytic continuation (AC) problem. Across the many fields that use QMC methods, there is no universally accepted analytic continuation algorithm for extracting dynamic properties, but many publications compare to the maximum entropy method. We investigate when entropy maximization is an acceptable approach. We show that stochastic sampling algorithms reduce to entropy maximization when the Bayesian prior is near to the true solution. We investigate when is Bryan's controversial optimization algorithm [Bryan, Eur. Biophys. J. 18, 165-174 (1990)] for entropy maximization (sometimes known as the maximum entropy method) appropriate to use. We show that Bryan's algorithm is appropriate when the noise is near zero or when the Bayesian prior is near to the true solution. We also investigate the mean squared error, finding a better scaling when the Bayesian prior is near the true solution than when the noise is near zero. We point to examples of improved data-driven Bayesian priors that have already leveraged this advantage. We support these results by solving the double Gaussian problem using both Bryan's algorithm and the newly formulated dual approach to entropy maximization [Chuna et al., J. Phys. A: Math. Theor. 58, 335203 (2025)].

physics.comp-ph

High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification

When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these questions theoretically in a sparse-token classification model in which positive samples embed a weak signal vector in a randomly chosen subset of tokens, whereas negative samples are pure noise. In the long-sequence limit, we show that a simple single-layer attention classifier can in principle achieve vanishing test error when the signal strength grows only logarithmically in the sequence length $L$, whereas linear classifiers require $\sqrt{L}$ scaling. Moving from representational power to learnability, we study training at finite $L$ in a high-dimensional regime, where sample size and embedding dimension grow proportionally. We prove that just two gradient updates suffice for the query weight vector of the attention classifier to acquire a nontrivial alignment with the hidden signal, inducing an attention map that selectively amplifies informative tokens. We further derive an exact asymptotic expression for the test error and training loss of the trained attention-based classifier, and quantify its capacity -- the largest dataset size that is typically perfectly separable -- thereby explaining the advantage of adaptive token selection over nonadaptive linear baselines.

cs.LG

Estimates of the dynamic structure factor for the finite temperature electron liquid via analytic continuation of path integral Monte Carlo data

Understanding the dynamic properties of the uniform electron gas (UEG) is important for numerous applications ranging from semiconductor physics to exotic warm dense matter. In this work, we apply the maximum entropy method (MEM), as implemented in Chuna \emph{et al.}~[arXiv:2501.01869], to \emph{ab initio} path integral Monte Carlo (PIMC) results for the imaginary-time correlation function $F(q,\tau)$ to estimate the dynamic structure factor $S(q,\omega)$ over an unprecedented range of densities at the electronic Fermi temperature. To conduct the MEM, we propose to construct the Bayesian prior $\mu$ from the PIMC data. Constructing the static approximation leads to a drastic improvement in $S(q,\omega)$ estimate over using the more simple random phase approximation (RPA) as the Bayesian prior. We find good agreement with existing results by Dornheim \emph{et al.}~[\textit{Phys.~Rev.~Lett.}~\textbf{121}, 255001 (2018)], where they are available. In addition, we present new results for the strongly coupled electron liquid regime with $r_s=50,...,200$, which reveal a pronounced roton-type feature and an incipient double peak structure in $S(q,\omega)$ at intermediate wavenumbers. We also find that our dynamic structure factors satisfy known sum rules, even though these sum rules are not enforced explicitly. An advantage of our set-up is that it is not specific to the UEG, thereby opening up new avenues to study the dynamics of real warm dense matter systems based on cutting-edge PIMC simulations in future works.

cond-mat.str-el

Dual formulation of the maximum entropy method applied to analytic continuation of quantum Monte Carlo data

Many fields of physics use quantum Monte Carlo techniques, but struggle to estimate dynamic spectra via the analytic continuation of imaginary-time quantum Monte Carlo data. One of the most ubiquitous approaches to analytic continuation is the maximum entropy method (MEM). We supply a dual Newton optimization algorithm to be used within the MEM and provide analytic bounds for the algorithm's error. The MEM is typically used with Bryan's controversial algorithm [Rothkopf, "Bryan's Maximum Entropy Method" Data 5.3 (2020)]. We present new theoretical issues that are not yet in the literature. Our algorithm has all the theoretical benefits of Bryan's algorithm without these theoretical issues. We compare the MEM with Bryan's optimization to the MEM with our dual Newton optimization on test problems from lattice quantum chromodynamics and plasma physics. These comparisons show that in the presence of noise the dual Newton algorithm produces better estimates and error bars; this indicates the limits of Bryan's algorithm's applicability. We use the MEM to investigate authentic quantum Monte Carlo data for the uniform electron gas at warm dense matter conditions and further substantiate the roton-type feature in the dispersion relation.

physics.comp-ph

Ziv-Merhav estimation for hidden-Markov processes

We present a proof of strong consistency of a Ziv-Merhav-type estimator of the cross entropy rate for pairs of hidden-Markov processes. Our proof strategy has two novel aspects: the focus on decoupling properties of the laws and the use of tools from the thermodynamic formalism.

cs.IT

On the Ziv-Merhav theorem beyond Markovianity II: leveraging the thermodynamic formalism

We prove asymptotic results for a modification of the cross-entropy estimator originally introduced by Ziv and Merhav in the Markovian setting in 1993. Our results concern a more general class of decoupled measures on shift spaces over a finite alphabet and in particular imply strong asymptotic consistency of the modified estimator for all pairs of functions of stationary, irreducible, finite-state Markov chains satisfying a mild decay condition. Our approach is based on the study of a rescaled cumulant-generating function called the cross-entropic pressure, importing to information theory some techniques from the study of large deviations within the thermodynamic formalism.

math.PR

On the Ziv-Merhav theorem beyond Markovianity

We generalize to a broader class of decoupled measures a result of Ziv and Merhav on universal estimation of the specific cross (or relative) entropy for a pair of multi-level Markov measures. The result covers pairs of suitably regular g-measures and pairs of equilibrium measures arising from the small space of interactions in mathematical statistical mechanics.

cs.IT