SearcharxivSearch

arXiv subjects

Reda Boumasmoud

Publications and source records attributed to Reda Boumasmoud.

13 recordsLinked to original sources

A Compositional Theory of Causally Masked Transformers

What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains and therefore what the model can compute. We develop an algebraic formalization that derives expressivity directly from the model's implemented dynamics. Its central object is its memory; the finite internal state computed by attention that summarizes the information from the prefix available to all future queries. Each attention head updates its own state independently within a layer, while layers compose hierarchically, providing a uniform route from model assumptions to expressivity bounds. Applying this method to transformers without positional embeddings, we obtain an expressivity hierarchy governed by the attention type under specific numerical semantics. Width-one sliding-window attention supports bounded-suffix memory, while a modified form of soft attention supports irreversible, checklist-like state, and combining the two mechanisms provides an interplay of both. Ordinary left-to-right floating-point soft attention can realize more expressive memory operations than any of the above. Algebraically, the four cases correspond to definite, R-trivial, locally R-trivial, and aperiodic semigroups. Under an explicit free-wiring assumption, all four bounds are tight.

cs.FL

Causally Evaluating the Learnability of Formal Language Tasks

Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. Answering this for natural language is difficult: tasks are hard to delineate and can confound one another. To rigorously investigate the relationship between data frequency and learnability, we turn to a controlled setting using formal languages induced from probabilistic finite automata. These serve as a methodological testbed to demonstrate that standard correlational evaluation practices are inherently flawed. To enable causal analysis, we introduce the binning semiring, an algebraic object that lets us control how often a targeted property occurs in a sampled corpus. We formulate the experimental pipeline as a causal graphical model and derive decomposed Kullback-Leibler divergence metrics to measure the learnability of specific sub-tasks. Our experiments show that evaluating learnability without causal intervention leads to incorrect conclusions due to confounders in correlational analysis, and serve as a warning about correlational pitfalls in natural-language settings.

cs.CL

An Algebraic View of the Expressivity of Recurrent Language Models

What formal languages can a recurrent neural language model recognize? Formal results in the literature conflict: some authors report Turing-completeness, while others show equivalence to regular languages. The reason for this discrepancy is that the underlying arithmetic model differs. The paper develops a unified algebraic account of the expressivity of recurrent neural networks, starting with a formal account of various arithmetic models. This account reduces expressivity to an algebraic question, e.g., whether a network's syntactic monoid divides a certain wreath product. As a case study, the paper revisits diagonal state-space models: the same architecture cannot implement an even-modulus counter once floating-point recurrences are enforced, yet realizes every even-modulus counter under unsigned-integer quantization.

cs.FL

Transducing Language Models

Modern language models define distributions over strings, but downstream tasks often require different output formats. For instance, a model that generates byte-pair strings does not directly produce word-level predictions, and a DNA model does not directly produce amino-acid sequences. In such cases, a deterministic string-to-string transformation can convert the model's output to the desired form. This is a familiar pattern in probability theory: applying a function $f$ to a random variable $X\sim p$ yields a transformed random variable $f(X)$ with an induced distribution. While such transformations are occasionally used in language modeling, prior work does not treat them as yielding new, fully functional language models. We formalize this perspective and introduce a general framework for language models derived from deterministic string-to-string transformations. We focus on transformations representable as finite-state transducers -- a commonly used state-machine abstraction for efficient string-to-string mappings. We develop algorithms that compose a language model with an FST to *marginalize* over source strings mapping to a given target, propagating probabilities through the transducer without altering model parameters and enabling *conditioning* on transformed outputs. We present an exact algorithm, an efficient approximation, and a theoretical analysis. We conduct experiments in three domains: converting language models from tokens to bytes, from tokens to words, and from DNA to amino acids. These experiments demonstrate inference-time adaptation of pretrained language models to match application-specific output requirements.

cs.CL

An $\mathbf{L^*}$ Algorithm for Deterministic Weighted Regular Languages

Extracting finite state automata (FSAs) from black-box models offers a powerful approach to gaining interpretable insights into complex model behaviors. To support this pursuit, we present a weighted variant of Angluin's (1987) $\mathbf{L^*}$ algorithm for learning FSAs. We stay faithful to the original algorithm, devising a way to exactly learn deterministic weighted FSAs whose weights support division. Furthermore, we formulate the learning process in a manner that highlights the connection with FSA minimization, showing how $\mathbf{L^*}$ directly learns a minimal automaton for the target language.

cs.CL

On Affine Homotopy between Language Encoders

Pre-trained language encoders -- functions that represent text as vectors -- are an integral component of many NLP tasks. We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar? We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity -- the performance on downstream tasks. It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation. In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity. We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity. We confirm this on datasets of natural language representations. Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them.

cs.CL

The center of Hecke algebras of types

We describe the center of the Hecke algebra of a type attached to a Bernstein block under some hypothesis. When $\bf G$ is a connected reductive group over non-archimedean local field $F$ that splits over a tamely ramified extension of $F$ and the residue characteristic of $F$ does not divide the order of the absolute Weyl group of $\bf G$, the works of Kim-Yu and Fintzen associate a type to each Bernstein block and our hypothesis is satisfied for such types. We use our results to give a description of the Bernstein center of the Hecke algebra $\mathcal{H}({\bf G } (F),K)$ when $K$ belongs to a nice family of compact open subgroups of ${\bf G}(F)$ (which includes all the Moy-Prasad filtrations of an Iwahori subgroup) via the theory of types.

math.RT

A tale of parahoric--Hecke algebras, Bernstein and Satake homomorphisms

Let $\mathbf{G}$ be a connected reductive group over a {non-archimedean local field} $F$. Let $K_\mathcal{F}$ be the parahoric subgroup attached to a facet $\mathcal{F}$ in the Bruhat--Tits building of $\mathbf{G}$. The ultimate goal of the present paper is to describe the center of the parahoric--Hecke algebra $\mathcal{H}(\mathbf{G}(F)//K_{\mathcal{F}}, \mathbb{Z}[q^{-1}])$ with level $K_{\mathcal{F}}$ and prove the compatibility of generalized (twisted) Bernstein and Satake homomorphisms.

math.RT

The ring of U-operators: Definitions and Integrality

In this paper, we define and study the arithmetic of the ring of $\mathbb{U}$-operators for reductive $p$-adic groups. These operators generalise the notion of "successor" operators for trees with a marked end. We show that they are integral over the spherical Hecke algebra. This integrality intervenes crucially in the construction of Euler systems obtained from special cycles of general Shimura varieties and the generalization of the famous Eichler-Shimura relation.

math.NT

Seed Relations for Eichler--Shimura congruences and Euler systems

This paper proves that the $\mathbb{U}$-operator attached to a cocharacter is a right root of the corresponding Hecke polynomial. This result is an important ingredient in the proof of (i) the horizontal norm relations in the context of Gross--Gan--Prasad cycles and of (ii) the generalization of Eichler--Shimura relations.

math.NT

Horizontal Distribution Relations for Special Cycles on Unitary Shimura Varieties: Split Case

We study the local behavior of special cycles on Shimura varieties for $\mathbf{U}(2, 1) \times \mathbf{U}(1, 1)$ in the setting of the Gan-Gross-Prasad conjectures at primes $τ$ of the totally real field of definition of the unitary spaces which are split in the corresponding totally imaginary quadratic extension. We establish a local formula for their fields of definition, and prove a distribution relation between the Galois and Hecke actions on them. This complements work of \cite{jetchev:unitary} at inert primes, where the combinatorics of the formulas are reduced to calculations on the Bruhat--Tits trees, which in the split case must be replaced with higher-dimensional buildings.

math.NT

Vertical Distribution Relations For Special Cycles on Unitary Shimura Varieties

We consider cycles on a 3-dimensional Shimura varieties attached to a unitary group, defined over extensions of a CM field $E$, which appear in the context of the conjectures of Gan, Gross, and Prasad \cite{gan-gross-prasad}. We establish a vertical distribution relation for these cycles over an anticyclotomic extension of $E$, complementing the horizontal distribution relation of \cite{jetchev:unitary}, and use this to define a family of norm-compatible cycles over these fields, thus obtaining a universal norm construction similar to the Heegner $Λ$-module constructed from Heegner points.

math.NT