SearcharxivSearch

arXiv subjects

Vinoth Nandakumar

Publications and source records attributed to Vinoth Nandakumar.

13 recordsLinked to original sources

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers. In language modeling, \textbf{transformers} have emerged as the dominant architecture, with early layers capturing local syntactic patterns and later layers encoding more complex clause-level dependencies. While this intuition has shaped model design, there remains a lack of rigorous theoretical work demonstrating \textbf{how} deep transformers represent such hierarchical structures. In this work, we analyze the expressiveness of deep transformer models through the formal lens of bounded-depth, non-recursive context-free grammars. For this class of grammars, we explicitly construct transformers with positional attention whose depth grows linearly with grammar depth, while the neuron count scales with the number of derivation-tree shapes and quadratically with the number of production rules. Our theoretical results support the linear representation hypothesis by demonstrating that these architectures possess the structural capacity to encode abstract grammatical states into low-dimensional, linearly separable subspaces within the residual stream.

cs.CL

A theoretical model for task routing in mixture-of-expert transformers

Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical studies of frontier MoE transformer models, existing theoretical work analyzes this using continuous mixture models that cannot be used to model natural language effectively. An important open question is to \textit{theoretically explain task-expert specialization in transformer MoE models using discrete models of language}. To address this, we represent structured knowledge via syntactic templates and finite key-value dictionaries, and prove formally that a single-layer MoE transformer can encode knowledge by using experts that specialize in the corresponding tasks. Our construction shows how queries are routed to unique, task-specific experts whose size depends solely on the intrinsic complexity of the given task (i.e. the combined size of its syntactic templates and factual dictionary). Our construction provides a theoretical support for empirical results on localized knowledge circuits in MoE models. We support our theoretical findings with experiments evaluating model performance under varying MoE loss functions.

cs.LG

Why do CNNs excel at feature extraction? A mathematical explanation

Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification benchmarks. However, a fundamental theoretical questions remain answered: why can they solve discrete image classification tasks that involve feature extraction? We address this question in this paper by introducing a novel mathematical model for image classification, based on feature extraction, that can be used to generate images resembling real-world datasets. We show that convolutional neural network classifiers can solve these image classification tasks with zero error. In our proof, we construct piecewise linear functions that detect the presence of features, and show that they can be realized by a convolutional network.

cs.CV

State space models can express n-gram languages

Recent advancements in recurrent neural networks (RNNs) have reinvigorated interest in their application to natural language processing tasks, particularly with the development of more efficient and parallelizable variants known as state space models (SSMs), which have shown competitive performance against transformer models while maintaining a lower memory footprint. While RNNs and SSMs (e.g., Mamba) have been empirically more successful than rule-based systems based on n-gram models, a rigorous theoretical explanation for this success has not yet been developed, as it is unclear how these models encode the combinatorial rules that govern the next-word prediction task. In this paper, we construct state space language models that can solve the next-word prediction task for languages generated from n-gram rules, thereby showing that the former are more expressive. Our proof shows how SSMs can encode n-gram rules using new theoretical results on their memorization capacity, and demonstrates how their context window can be controlled by restricting the spectrum of the state transition matrix. We conduct experiments with a small dataset generated from n-gram rules to show how our framework can be applied to SSMs and RNNs obtained through gradient-based optimization.

cs.CL

Categorification via blocks of modular representations II

Bernstein, Frenkel and Khovanov have constructed a categorification of tensor products of the standard representation of $\mathfrak{sl}_2$ using singular blocks of category $\mathcal{O}$ for $\mathfrak{sl}_n$. In earlier work, we construct a positive characteristic analogue using blocks of representations of $\mathfrak{sl}_n$ over a field $\textbf{k}$ of characteristic $p > n$, with zero Frobenius character, and singular Harish-Chandra character. In the present paper, we extend these results and construct a categorical $\mathfrak{sl}_k$-action, following Sussan's approach, by considering more singular blocks of modular representations of $\mathfrak{sl}_n$. We consider both zero and non-zero Frobenius central character. In the former setting, we construct a graded lift of these categorifications which are equivalent to a geometric construction of Cautis, Kamnitzer and Licata. We establish a Koszul duality between two geometric categorificatons constructed in their work, and resolve a conjecture of theirs. For non-zero Frobenius central characters, we show that the geometric approach to categorical symmetric Howe duality by Cautis and Kamnitzer can be used to construct a graded lift of our categorification using singular blocks of modular representations of $\mathfrak{sl}_n$.

math.RT

Modular representations in type A with a two-row nilpotent central character

We study the category of representations of $\mathfrak{sl}_{m+2n}$ in positive characteristic, whose p-character is a nilpotent whose Jordan type is the two-row partition (m+n,n). In a previous paper with Anno, we used Bezrukavnikov-Mirkovic-Rumynin's theory of positive characteristic localization and exotic t-structures to give a geometric parametrization of the simples using annular crossingless matchings. Building on this, here we give combinatorial dimension formulae for the simple objects, and compute the Jordan-Holder multiplicities of the simples inside the baby Vermas (in special case where n=1, i.e. that a subregular nilpotent, these were known from work of Jantzen). We use Cautis-Kamnitzer's geometric categorification of the tangle calculus to study the images of the simple objects under the [BMR] equivalence. The dimension formulae may be viewed as a positive characteristic analogue of the combinatorial character formulae for simple objects in parabolic category O for $\mathfrak{sl}_{m+2n}$, due to Lascoux and Schutzenberger.

math.RT

Categorification via blocks of modular representations for sl(n)

Bernstein, Frenkel, and Khovanov have constructed a categorification of tensor products of the standard representation of $\mathfrak{sl}_2$, where they use singular blocks of category $\mathcal{O}$ for $\mathfrak{sl}_n$ and translation functors. Here we construct a positive characteristic analogue using blocks of representations of $\mathfrak{sl}_n$ over a field $\textbf{k}$ of characteristic $p$ with zero Frobenius character, and singular Harish-Chandra character. We show that the aforementioned categorification admits a Koszul graded lift, which is equivalent to a geometric categorification constructed by Cautis, Kamnitzer, and Licata using coherent sheaves on cotangent bundles to Grassmanians. In particular, the latter admits an abelian refinement. With respect to this abelian refinement, the stratified Mukai flop induces a perverse equivalence on the derived categories for complementary Grassmanians. This is part of a larger project to give a combinatorial approach to Lusztig's conjectures for representations of Lie algebras in positive characteristic.

math.RT

Irreducible components of exotic Springer fibres

Kato introduced the exotic nilpotent cone to be a substitute for the ordinary nilpotent cone of type C with cleaner properties. Here we describe the irreducible components of exotic Springer fibres (the fibres of the resolution of the exotic nilpotent cone), and prove that they are naturally in bijection with standard bitableaux. As a result, we deduce the existence of an exotic Robinson-Schensted bijection, which is a variant of the type C Robinson-Schensted bijection between pairs of same-shape standard bitableaux and elements of the Weyl group; this bijection is described explicitly in the sequel to this paper. Note that this is in contrast with ordinary type C Springer fibres, where the parametrisation of irreducible components, and the resulting geometric Robinson-Schensted bijection, are more complicated. As an application, we explicitly describe the structure in the special cases where the irreducible components of the exotic Springer fibre have dimension 2, and show that in those cases one obtains Hirzebruch surfaces.

math.RT

Irreducible Components of Exotic Springer fibres II: The Exotic Robinson-Schensted Algorithm

Kato's exotic nilpotent cone was introduced as a substitute for the ordinary nilpotent cone of type C with cleaner properties. The geometric Robinson-Schensted correspondence is obtained by parametrizing the irreducible components of the Steinberg variety (the conormal variety for the action of a semisimple group on two copies of its flag variety); in type A the bijection coincides with the classical Robinson-Schensted algorithm for the symmetric group. Here we give a combinatorial description of the bijection obtained by using the exotic nilpotent cone instead of ordinary type C nilpotent cone in the geometric Robinson-Schensted correspondence; we refer this as the "exotic Robinson-Schensted bijection". This is interesting from a combinatorial perspective, and not a naive extension of the type A Robinson-Schensted bijection.

math.RT

Quiver varieties and crystals in symmetrizable type via modulated graphs

Kashiwara and Saito have a geometric construction of the infinity crystal for any symmetric Kac-Moody algebra. The underlying set consists of the irreducible components of Lusztig's quiver varieties, which are varieties of nilpotent representations of a pre-projective algebra. We generalize this to symmetrizable Kac-Moody algebras by replacing Lusztig's preprojective algebra with a more general one due to Dlab and Ringel. In non-symmetric types we are forced to work over non-algebraically-closed fields.

math.QA

Exotic t-structures for two-block Springer fibers

We study the exotic t-structure on the derived category of coherent sheaves on two-block Springer fibre (i.e. for a nilpotent matrix of type (m+n,n) in type A). The exotic t-structure has been defined by Bezrukavnikov and Mirkovic for Springer theoretic varieties in order to study representations of Lie algebras in positive characteristic. Using work of Cautis and Kamnitzer, we construct functors indexed by affine tangles, between categories of coherent sheaves on different two-block Springer fibres (i.e. for different values of n). After checking some exactness properties of these functors, we describe the irreducible objects in the heart of the exotic t-structure, and enumerate them by crossingless (m,m+2n) matchings. We compute the Ext's between the irreducible objects, and show that the resulting algebras are an annular variant of Khovanov's arc algebras. In subsequent work we will make a link with annular Khovanov homology, and use these results to give a positive characteristic analogue of some categorification results using two-block parabolic category O (by Bernstein-Frenkel-Khovanov, Brundan, Stroppel, et al).

math.RT

Stability conditions for Gelfand-Kirillov subquotients of category O

Recently, Anno, Bezrukavnikov and Mirkovic have introduced the notion of a "real variation of stability conditions" (which is related to Bridgeland's stability conditions), and construct an example using categories of coherent sheaves on Springer fibers. Here we construct another example of representation theoretic significance, by studying certain sub-quotients of category O with a fixed Gelfand-Kirillov dimension. We use the braid group action on the derived category of category O, and certain leading coefficient polynomials coming from translation functors. Consequently, we use this to explicitly describe a sub-manifold in the space of Bridgeland stability conditions on these sub-quotient categories, which is a covering space of a hyperplane complement in the dual Cartan.

math.RT

Equivariant coherent sheaves on the exotic nilpotent cone

Let $G=Sp_{2n}(\mathbb{C})$, and $\mathfrak{N}$ be Kato's exotic nilpotent cone. Following techniques used by Bezrukavnikov in [5] to establish a bijection between $Λ^+$, the dominant weights for a simple algebraic group $H$, and $\textbf{O}$, the set of pairs consisting of a nilpotent orbit and a finite-dimensional irreducible representation of the isotropy group of the orbit, we prove an analogous statement for the exotic nilpotent cone. First we prove that dominant line bundles on the exotic Springer resolution $\widetilde{\mathfrak{N}}$ have vanishing higher cohomology, and compute their global sections using techniques of Broer. This allows to show that the direct images of these dominant line bundles constitute a quasi-exceptional set generating the category $D^b(Coh^G(\mathfrak{N}))$, and deduce that the resulting $t$-structure on $D^b(Coh^G(\mathfrak{N}))$ coincides with the perverse coherent $t$-structure. The desired result now follows from the bijection between costandard objects and simple objects in the heart of this $t$-structure on $D^b(Coh^G(\mathfrak{N}))$.

math.QA