SearcharxivSearch

arXiv subjects

Francesco D'Angelo

Publications and source records attributed to Francesco D'Angelo.

At least 19 recordsLinked to original sources

Induction Heads Interpolate N-Grams

Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive. We study transformers trained on order-$k$ Markov chains and identify two complementary smoothing mechanisms. First, at finite attention-weight scale, the circuit implements a soft context-matching estimator: it aggregates contributions from exact and partial context matches, weighted exponentially by their overlap, and induces a data-dependent interpolation across context orders analogous to Jelinek-Mercer smoothing. Second, a beginning-of-sequence (BOS) token induces additive pseudo-counts, recovering Dirichlet-style smoothing. We construct a disentangled transformer implementing both mechanisms and show that trained transformers recover the predicted attention patterns. Across settings where pseudo-count smoothing is optimal or lower-order contexts provide structured evidence, trained transformers match or outperform classical count-based baselines. Our results bridge mechanistic interpretability of induction heads with classical statistical smoothing, revealing that transformers learn to regularize in-context estimation rather than simply count.

cs.LG

Transformers Learn Latent Mixture Models In-Context via Mirror Descent

Sequence modelling requires determining which past tokens are causally relevant from the context and their importance: a process inherent to the attention layers in transformers, yet whose underlying learned mechanisms remain poorly understood. In this work, we formalize the task of estimating token importance as an in-context learning problem by introducing a framework based on Mixture of Transition Distributions, where a latent variable determines the influence of past tokens on the next. The distribution over this latent variable is parameterized by unobserved mixture weights that transformers must learn in-context. We demonstrate that transformers can implement Mirror Descent to learn these weights from the context. Specifically, we give an explicit construction of a three-layer transformer that exactly implements one step of Mirror Descent and prove that the resulting estimator is a first-order approximation of the Bayes-optimal predictor. Corroborating our construction and its learnability via gradient descent, we empirically show that transformers trained from scratch learn solutions consistent with our theory: their predictive distributions, attention patterns, and learned transition matrix closely match the construction, while deeper models achieve performance comparable to multi-step Mirror Descent.

cs.LG

Lattice determination of the QCD low-energy constant $\ell_{\scriptscriptstyle{7}}$

We provide a non-perturbative determination of the scheme- and scale-independent low-energy constant $\ell_{\scriptscriptstyle{7}}$, appearing in the QCD effective chiral Lagrangian at next-to-leading order, by means of lattice QCD simulations with $N_{\scriptscriptstyle{\rm f}}=2+1$ quark flavors. We adopt staggered fermions and extract $\ell_{\scriptscriptstyle{7}}$ from the pion mass splitting by suitably generalizing the method introduced in [Phys. Rev. D 104 (2021) 074513] for the Wilson discretization. Adopting 12 gauge ensembles with 3 different values of the pion mass, and 4 different values of the lattice spacing, we are able to achieve controlled extrapolations towards the continuum, infinite volume, and chiral limits. Our final result $\ell_{\scriptscriptstyle{7}} \,\times \, 10^3 = 2.79(58)_{\scriptscriptstyle{\rm stat}}(19)_{\scriptscriptstyle{\rm syst}} = 2.79(61)_{\scriptscriptstyle{\rm tot}}$ agrees with and substantially improves on previous determinations.

hep-lat

Exact Learning of Arithmetic with Differentiable Agents

We explore the possibility of exact algorithmic learning with gradient-based methods and introduce a differentiable framework capable of strong length generalization on arithmetic tasks. Our approach centers on Differentiable Finite-State Transducers (DFSTs), a Turing-complete model family that avoids the pitfalls of prior architectures by enabling constant-precision, constant-time generation, and end-to-end log-parallel differentiable training. Leveraging policy-trajectory observations from expert agents, we train DFSTs to perform binary and decimal addition and multiplication. Remarkably, models trained on tiny datasets generalize without error to inputs thousands of times longer than the training examples. These results show that training differentiable agents on structured intermediate supervision could pave the way towards exact gradient-based learning of algorithmic skills. Code available at \href{https://github.com/dngfra/differentiable-exact-algorithmic-learner.git}{https://github.com/dngfra/differentiable-exact-algorithmic-learner.git}.

cs.LG

Selective Induction Heads: How Transformers Select Causal Structures In Context

Transformers have exhibited exceptional capabilities in sequence modeling tasks, leveraging self-attention and in-context learning. Critical to this success are induction heads, attention circuits that enable copying tokens based on their previous occurrences. In this work, we introduce a novel framework that showcases transformers' ability to dynamically handle causal structures. Existing works rely on Markov Chains to study the formation of induction heads, revealing how transformers capture causal dependencies and learn transition probabilities in-context. However, they rely on a fixed causal structure that fails to capture the complexity of natural languages, where the relationship between tokens dynamically changes with context. To this end, our framework varies the causal structure through interleaved Markov chains with different lags while keeping the transition probabilities fixed. This setting unveils the formation of Selective Induction Heads, a new circuit that endows transformers with the ability to select the correct causal structure in-context. We empirically demonstrate that transformers learn this mechanism to predict the next token by identifying the correct lag and copying the corresponding token from the past. We provide a detailed construction of a 3-layer transformer to implement the selective induction head, and a theoretical analysis proving that this mechanism asymptotically converges to the maximum likelihood solution. Our findings advance the understanding of how transformers select causal structures, providing new insights into their functioning and interpretability.

cs.LG

An update on the determination of the sphaleron rate in finite temperature QCD

The sphaleron rate is a key phenomenological quantity both for the axion thermal production in the Early Universe and the Chiral Magnetic Effect occurring in the Quark-Gluon Plasma in presence of a background magnetic field. In this talk we present an extension of our recent determination of the sphaleron rate, in the SU(3) gauge theory, based on the determination of the two-point function of the topological charge density at finite temperature.

hep-lat

Why Do We Need Weight Decay in Modern Deep Learning?

Weight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models. Despite its widespread usage and being extensively studied in the classical literature, its role remains poorly understood for deep learning. In this work, we highlight that the role of weight decay in modern deep learning is different from its regularization effect studied in classical learning theory. For deep networks on vision tasks trained with multipass SGD, we show how weight decay modifies the optimization dynamics enhancing the ever-present implicit regularization of SGD via the loss stabilization mechanism. In contrast, for large language models trained with nearly one-epoch training, we describe how weight decay balances the bias-variance tradeoff in stochastic optimization leading to lower training loss and improved training stability. Overall, we present a unifying perspective from ResNets on vision tasks to LLMs: weight decay is never useful as an explicit regularizer but instead changes the training dynamics in a desirable way. The code is available at https://github.com/tml-epfl/why-weight-decay

cs.LG

Non-perturbative determination of the $N_f=2+1$ QCD sphaleron rate

The strong sphaleron rate, i.e., the rate of real time QCD topological transitions, is a key phenomenological quantity, playing a fundamental role in several physical contexts. In heavy-ion collisions, a non-vanishing rate can lead to the so-called Chiral Magnetic Effect. In early-Universe cosmology, instead, it can be related to the rate of thermal production of QCD axions. In this talk, we present the first reliable fully non-perturbative computation of the strong sphaleron rate in $N_f=2+1$ QCD at the physical point by means of lattice simulations, in a range of temperatures going from 200 MeV to 600 MeV. Our strategy is based on the inversion of lattice correlators via a recently-proposed modified version of the Backus-Gilbert method.

hep-lat

Sphaleron rate of $N_f=2+1$ QCD

We compute the sphaleron rate of $N_f=2+1$ QCD at the physical point for a range of temperatures $200$ MeV $\lesssim T \lesssim 600$ MeV. We adopt a strategy recently applied in the quenched case, based on the extraction of the rate via a modified version of the Backus-Gilbert method from finite-lattice-spacing and finite-smoothing-radius Euclidean topological charge density correlators. The physical sphaleron rate is finally computed by performing a continuum limit at fixed physical smoothing radius, followed by a zero-smoothing extrapolation. Dynamical fermions were discretized using the staggered formulation, which is known to yield large lattice artifacts for the topological susceptibility. However, we find them to be rather mild for the sphaleron rate.

hep-lat

Sphaleron rate as an inverse problem: a novel lattice approach

We compute the sphaleron rate on the lattice. We adopt a novel strategy based on the extraction of the spectral density via a modified version of the Backus-Gilbert method from finite-lattice-spacing and finite-smoothing-radius Euclidean topological charge density correlators. The physical sphaleron rate is computed by performing controlled continuum limit and zero-smoothing extrapolations both in pure gauge and, for the first time, in full QCD.

hep-lat

The chiral condensate of $N_f=2+1$ QCD from the spectrum of the staggered Dirac operator

We compute the chiral condensate of $2+1$ QCD from the mode number of the staggered Dirac operator, performing controlled extrapolations to both the continuum and the chiral limit. We consider also alternative strategies, based on the quark mass dependence of the topological susceptibility and of the pion mass, and obtain consistent results within errors. Results are also consistent with phenomenological expectations and with previous numerical determinations obtained with different lattice discretizations.

hep-lat

Sphaleron rate from lattice QCD

We compute the sphaleron rate on the lattice from the inversion of the Euclidean time correlators of the topological charge density, performing also controlled continuum and zero-smoothing extrapolations. The correlator inversion is performed by means of a recently-proposed modification of the Backus-Gilbert method.

hep-lat

Sphaleron rate from a modified Backus-Gilbert inversion method

We compute the sphaleron rate in quenched QCD for a temperature $T \simeq 1.24~T_c$ from the inversion of the Euclidean lattice time correlator of the topological charge density. We explore and compare two different strategies: one follows a new approach proposed in this study and consists in extracting the rate from finite lattice spacing correlators, and then in taking the continuum limit at fixed smoothing radius followed by a zero-smoothing extrapolation; the other follows the traditional approach of extracting the rate after performing such double extrapolation directly on the correlator. In both cases the rate is obtained from a recently-proposed modification of the standard Backus-Gilbert procedure. The two strategies lead to compatible estimates within errors, which are then compared to previous results in the literature at the same or similar temperatures; the new strategy permits to obtain improved results, in terms of statistical and systematic uncertainties.

hep-lat

Repulsive Deep Ensembles are Bayesian

Deep ensembles have recently gained popularity in the deep learning community for their conceptual simplicity and efficiency. However, maintaining functional diversity between ensemble members that are independently trained with gradient descent is challenging. This can lead to pathologies when adding more ensemble members, such as a saturation of the ensemble performance, which converges to the performance of a single model. Moreover, this does not only affect the quality of its predictions, but even more so the uncertainty estimates of the ensemble, and thus its performance on out-of-distribution data. We hypothesize that this limitation can be overcome by discouraging different ensemble members from collapsing to the same function. To this end, we introduce a kernelized repulsive term in the update rule of the deep ensembles. We show that this simple modification not only enforces and maintains diversity among the members but, even more importantly, transforms the maximum a posteriori inference into proper Bayesian inference. Namely, we show that the training dynamics of our proposed repulsive ensembles follow a Wasserstein gradient flow of the KL divergence with the true posterior. We study repulsive terms in weight and function space and empirically compare their performance to standard ensembles and Bayesian baselines on synthetic and real-world prediction tasks.

cs.LG

The QCD topological susceptibility at high temperatures via staggered fermions spectral projectors

The QCD topological observables are essential inputs to obtain theoretical predictions about axion phenomenology, which are of utmost importance for current and future experimental searches for this particle. Among them, we focus on the topological susceptibility, related to the axion mass. We present lattice results for the topological susceptibility in QCD at high temperatures obtained by discretizing this observable via spectral projectors on eigenmodes of the staggered Dirac operator, and we compare them with those obtained with the standard gluonic definition. The adoption of the spectral discretization is motivated by the large lattice artifacts affecting the standard gluonic susceptibility, related to the choice of non-chiral fermions in the lattice action.

hep-lat

Topological susceptibility of $N_f=2+1$ QCD from staggered fermions spectral projectors at high temperatures

We compute the topological susceptibility of $N_f=2+1$ QCD with physical quark masses in the high-temperature phase, using numerical simulations of the theory discretized on a space-time lattice. More precisely we estimate the topological susceptibility for five temperatures in the range from $\sim200$ MeV up to $\sim600$ MeV, adopting the spectral projectors definition of the topological charge based on the staggered Dirac operator. This strategy turns out to be effective in reducing the large lattice artifacts which affect the standard gluonic definition, making it possible to perform a reliable continuum extrapolation. Our results for the susceptibility in the explored temperature range are found to be partially in tension with previous determinations in the literature.

hep-lat

On out-of-distribution detection with Bayesian neural networks

The question whether inputs are valid for the problem a neural network is trying to solve has sparked interest in out-of-distribution (OOD) detection. It is widely assumed that Bayesian neural networks (BNNs) are well suited for this task, as the endowed epistemic uncertainty should lead to disagreement in predictions on outliers. In this paper, we question this assumption and show that proper Bayesian inference with function space priors induced by neural networks does not necessarily lead to good OOD detection. To circumvent the use of approximate inference, we start by studying the infinite-width case, where Bayesian inference can be exact due to the correspondence with Gaussian processes. Strikingly, the kernels derived from common architectural choices lead to function space priors which induce predictive uncertainties that do not reflect the underlying input data distribution and are therefore unsuited for OOD detection. Importantly, we find the OOD behavior in this limiting case to be consistent with the corresponding finite-width case. To overcome this limitation, useful function space properties can also be encoded in the prior in weight space, however, this can currently only be applied to a specified subset of the domain and thus does not inherently extend to OOD data. Finally, we argue that a trade-off between generalization and OOD capabilities might render the application of BNNs for OOD detection undesirable in practice. Overall, our study discloses fundamental problems when naively using BNNs for OOD detection and opens interesting avenues for future research.

cs.LG