SearcharxivSearch

arXiv · 2501.06662

The magnitude of categories of texts enriched by language models

Abstract

The purpose of this article is twofold. Firstly, we use the next-token probabilities given by a language model to explicitly define a category of texts in natural language enriched over the unit interval, in the sense of Bradley, Terilla, and Vlassopoulos. We consider explicitly the terminating conditions for text generation and determine when the enrichment itself can be interpreted as a probability over texts. Secondly, we compute the M\"obius function and the magnitude of an associated generalized metric space of texts. The magnitude function of that space is a sum over texts (prompts) of the $t$-logarithmic (Tsallis) entropies of the next-token probability distributions associated with each prompt, plus the cardinality of the model's possible outputs. A suitable evaluation of the magnitude function's derivative recovers a sum of Shannon entropies, which justifies seeing magnitude as a partition function. Following Leinster and Shulman, we also express the magnitude function of the generalized metric space as an Euler characteristic of magnitude homology and provide an explicit description of the zeroeth and first magnitude homology groups.

Explore related subjects

Keep this discovery

BibTeXRIS

Tai-Danae Bradley, Juan Pablo Vigneaux. 2025-01-11. The magnitude of categories of texts enriched by language models. https://arxiv.org/abs/2501.06662

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A model structure for cartesian 2-fibrations

Cartesian 2-fibrations provide a way to understand indexed categories, but their classical ``straightening'' construction requires several layers of weak coherence data. This paper develops a homotopical framework that replaces much of this bookkeeping with a fully strict model. By using marked 2-categories to record the cartesian morphisms and 2-cells, we construct a model structure whose fibrant objects are precisely the cartesian 2-fibrations over a fixed 2-category $\mathcal{C}$. We then show that the marked Grothendieck construction identifies these 2-fibrations, up to weak equivalence, with strict 2-functors from $\mathcal{C}$ into $2\mathrm{Cat}$. As an additional contribution, we construct localizations of 2-categories that simultaneously invert selected morphisms and 2-cells.

math.CT

A Natural Fuzzy Order on Fuzzy Numbers

This paper introduces a natural fuzzy order on fuzzy numbers that extends the natural orders on real numbers and interval numbers. We investigate its completeness properties and show that the space of uniformly bounded fuzzy numbers is conically complete and conically cocomplete, and that it is complete if and only if the underlying continuous t-norm is the G\"odel t-norm. Moreover, it is proved that this space constitutes a \([0,1]\)-enriched domain if and only if the underlying continuous t-norm satisfies the (S) condition. These results provide a foundation for ordering fuzzy numbers.

math.CT

Noetherian forms of free non-symmetric operads

In this paper, we study certain categories of labeled finite rooted ordered trees over a fixed set of labels where each label is equipped with an arity: a fixed number of children that the vertex with the given label must have. Equivalently, these are expression trees for operations in a free non-symmetric operad. A morphism between these trees matches a pruning of one tree (a prefix) with an entire subtree of another (a suffix). We characterize such categories, up to isomorphism, in terms of suitable exactness properties. It turns out that these categories exhibit strong algebraic behavior, in the sense that every such category, when appended with a strict initial object, has a particularly nice noetherian form.

math.CT