SearcharxivSearch

arXiv subjects

Ben Blum-Smith

Publications and source records attributed to Ben Blum-Smith.

At least 19 recordsLinked to original sources

If you can distinguish, you can express: Galois theory, Stone--Weierstrass, machine learning, and linguistics

This essay develops a parallel between the Fundamental Theorem of Galois Theory and the Stone--Weierstrass theorem: both can be viewed as assertions that tie the distinguishing power of a class of objects to their expressive power. We provide an elementary theorem connecting the relevant notions of "distinguishing power". We also discuss machine learning and data science contexts in which these theorems, and more generally the theme of links between distinguishing power and expressive power, appear. Finally, we discuss the same theme in the context of linguistics, where it appears as a foundational principle, and illustrate it with several examples.

math.HO

Comparing the face rings of a boolean complex and its barycentric subdivision

We consider the relationship between the Stanley--Reisner ring (a.k.a. face ring) of a simplicial or boolean complex $Δ$ and that of its barycentric subdivision. These rings share a distinguished parameter subring. S. Murai asked if they are isomorphic, equivariantly with respect to the automorphism group $\operatorname{Aut}(Δ)$, as modules over this parameter subring. We show that, in general, the answer is no, but for Cohen--Macaulay complexes in characteristic coprime to $|\operatorname{Aut}(Δ)|$, it is yes, and we give an explicit construction of an isomorphism. To give this construction, we adapt a pair of tools introduced by A. Garsia in 1980. The first one transfers bases from a Stanley--Reisner ring to closely related rings of which it is a Gröbner degeneration, and the second identifies bases to transfer.

math.AC

Generic orbits, normal bases, and generation degree for fields of rational invariants

For a faithful linear representation $V$ of a finite group $G$ in coprime characteristic, we show that if the field Noether number $β_{\mathrm{field}}$ is the minimum $d$ such that the invariant polynomials of degree $\leq d$ generate the field $k(V)^G$ of rational invariants as a field, and the spanning degree $D_\mathrm{span}$ is the minimum $d$ such that the polynomials of degree $\leq d$ span the rational function field $k(V)$ as a vector space over $k(V)^G$, then $β_{\mathrm{field}} \leq 2D_\mathrm{span} + 1$, and this is sharp. This generalizes a recent result of Edidin and Katz. We also study $D_\mathrm{span}$. We show that it is related to various quantities previously studied in invariant and representation theory. Dropping the coprime characteristic hypothesis, we prove several basic inequalities, including that it is monotonically nondecreasing in $G$, nonincreasing in $V$, and satisfies $D_\mathrm{span} \leq |G|-1$. The latter refines a recent result of Kollar and Tiep.

math.AC

Geometry of numbers and degree bounds for rational invariants

We investigate degree bounds for fields of rational invariants of representations of finite groups. We prove many cases of a bound for $\mathbb{Z}/p\mathbb{Z}$ conjectured by Blum-Smith, Garcia, Hidalgo, and Rodriguez. For arbitrary groups, we also prove a new bound on the minimum degree $d$ such that the polynomials of degree $\leq d$ span the field of rational functions as a vector space over the invariant field. This latter quantity also bounds the degree $d$ such that the polynomials of degree $\leq d$ contain a copy of the regular representation of $G$, advancing an inquiry of Kollár and Tiep. The methods involve Euclidean lattices and Minkowski's geometry of numbers.

math.AC

Estimating the Euclidean distortion of an orbit space

Given a finite-dimensional inner product space $V$ and a group $G$ of isometries, we consider the problem of embedding the orbit space $V/G$ into a Hilbert space in a way that preserves the quotient metric as well as possible. This inquiry is motivated by applications to invariant machine learning. We introduce several new theoretical tools before using them to tackle various fundamental instances of this problem.

math.MG

A Galois theorem for machine learning: Functions on symmetric matrices and point clouds via lightweight invariant features

In this work, we present a mathematical formulation for machine learning of (1) functions on symmetric matrices that are invariant with respect to the action of permutations by conjugation, and (2) functions on point clouds that are invariant with respect to rotations, reflections, and permutations of the points. To achieve this, we provide a general construction of generically separating invariant features using ideas inspired by Galois theory. We construct $O(n^2)$ invariant features derived from generators for the field of rational functions on $n\times n$ symmetric matrices that are invariant under joint permutations of rows and columns. We show that these invariant features can separate all distinct orbits of symmetric matrices except for a measure zero set; such features can be used to universally approximate invariant functions on almost all weighted graphs. For point clouds in a fixed dimension, we prove that the number of invariant features can be reduced, generically without losing expressivity, to $O(n)$, where $n$ is the number of points. We combine these invariant features with DeepSets to learn functions on symmetric matrices and point clouds with varying sizes. We empirically demonstrate the feasibility of our approach on molecule property regression and point cloud distance prediction.

cs.LG

Equivariant geometric convolutions for emulation of dynamical systems

Machine learning methods are increasingly being employed as surrogate models in place of computationally expensive and slow numerical integrators for a bevy of applications in the natural sciences. However, while the laws of physics are relationships between scalars, vectors, and tensors that hold regardless of the frame of reference or chosen coordinate system, surrogate machine learning models are not coordinate-free by default. We enforce coordinate freedom by using geometric convolutions in three model architectures: a ResNet, a Dilated ResNet, and a UNet. In numerical experiments emulating 2D compressible Navier-Stokes, we see better accuracy and improved stability compared to baseline surrogate models in almost all cases. The ease of enforcing coordinate freedom without making major changes to the model architecture provides an exciting recipe for any CNN-based method applied to an appropriate class of problems

cs.LG

Degree bounds for rational generators of invariant fields of finite abelian groups

We study degree bounds on rational but not necessarily polynomial generators for the field $\mathbf{k}(V)^G$ of rational invariants of a linear action of a finite abelian group. We show that lattice-theoretic methods used recently by the author and collaborators to study polynomial generators for the same field largely carry over, after minor modifications to the arguments. It then develops that the specific degree bounds found in that setting also carry over.

math.AC

Degree bounds for fields of rational invariants of $\mathbb{Z}/p\mathbb{Z}$ and other finite groups

Degree bounds for algebra generators of invariant rings are a topic of longstanding interest in invariant theory. We study the analogous question for field generators for the field of rational invariants of a representation of a finite group, focusing on abelian groups and especially the case of $\mathbb{Z}/p\mathbb{Z}$. The inquiry is motivated by an application to signal processing. We give new lower and upper bounds depending on the number of distinct nontrivial characters in the representation. We obtain additional detailed information in the case of two distinct nontrivial characters. We conjecture a sharper upper bound in the $\mathbb{Z}/p\mathbb{Z}$ case, and pose questions for further investigation.

math.AC

Estimation under group actions: recovering orbits from invariants

We study a class of orbit recovery problems in which we observe independent copies of an unknown element of $\mathbb{R}^p$, each linearly acted upon by a random element of some group (such as $\mathbb{Z}/p$ or $\mathrm{SO}(3)$) and then corrupted by additive Gaussian noise. We prove matching upper and lower bounds on the number of samples required to approximately recover the group orbit of this unknown element with high probability. These bounds, based on quantitative techniques in invariant theory, give a precise correspondence between the statistical difficulty of the estimation problem and algebraic properties of the group. Furthermore, we give computer-assisted procedures to certify these properties that are computationally efficient in many cases of interest. The model is motivated by geometric problems in signal processing, computer vision, and structural biology, and applies to the reconstruction problem in cryo-electron microscopy (cryo-EM), a problem of significant practical interest. Our results allow us to verify (for a given problem size) that if cryo-EM images are corrupted by noise with variance $σ^2$, the number of images required to recover the molecule structure scales as $σ^6$. We match this bound with a novel (albeit computationally expensive) algorithm for ab initio reconstruction in cryo-EM, based on invariant features of degree at most 3. We further discuss how to recover multiple molecular structures from mixed (or heterogeneous) cryo-EM samples.

math.ST

Machine learning and invariant theory

Inspired by constraints from physical law, equivariant machine learning restricts the learning to a hypothesis class where all the functions are equivariant with respect to some group action. Irreducible representations or invariant theory are typically used to parameterize the space of such functions. In this article, we introduce the topic and explain a couple of methods to explicitly parameterize equivariant functions that are being used in machine learning applications. In particular, we explicate a general procedure, attributed to Malgrange, to express all polynomial maps between linear spaces that are equivariant under the action of a group $G$, given a characterization of the invariant polynomials on a bigger space. The method also parametrizes smooth equivariant maps in the case that $G$ is a compact Lie group.

stat.ML

Scalars are universal: Equivariant machine learning, structured like classical physics

There has been enormous progress in the last few years in designing neural networks that respect the fundamental symmetries and coordinate freedoms of physical law. Some of these frameworks make use of irreducible representations, some make use of high-order tensor objects, and some apply symmetry-enforcing constraints. Different physical laws obey different combinations of fundamental symmetries, but a large fraction (possibly all) of classical physics is equivariant to translation, rotation, reflection (parity), boost (relativity), and permutations. Here we show that it is simple to parameterize universally approximating polynomial functions that are equivariant under these symmetries, or under the Euclidean, Lorentz, and Poincaré groups, at any dimensionality $d$. The key observation is that nonlinear O($d$)-equivariant (and related-group-equivariant) functions can be universally expressed in terms of a lightweight collection of scalars -- scalar products and scalar contractions of the scalar, vector, and tensor inputs. We complement our theory with numerical examples that show that the scalar-based method is simple, efficient, and scalable.

cs.LG

Dimensionless machine learning: Imposing exact units equivariance

Units equivariance (or units covariance) is the exact symmetry that follows from the requirement that relationships among measured quantities of physics relevance must obey self-consistent dimensional scalings. Here, we express this symmetry in terms of a (non-compact) group action, and we employ dimensional analysis and ideas from equivariant machine learning to provide a methodology for exactly units-equivariant machine learning: For any given learning task, we first construct a dimensionless version of its inputs using classic results from dimensional analysis, and then perform inference in the dimensionless space. Our approach can be used to impose units equivariance across a broad range of machine learning methods which are equivariant to rotations and other groups. We discuss the in-sample and out-of-sample prediction accuracy gains one can obtain in contexts like symbolic regression and emulation, where symmetry is important. We illustrate our approach with simple numerical examples involving dynamical systems in physics and ecology.

stat.ML

Chords of an ellipse, Lucas polynomials, and cubic equations

A beautiful theorem of Thomas Price links the Fibonacci numbers and the Lucas polynomials to the plane geometry of an ellipse, generalizing a classic problem about circles. We give a brief history of the circle problem, an account of Price's ellipse proof, and a reorganized proof, with some new ideas, designed to situate the result within a dense web of connections to classical mathematics. It is inspired by Cardano's solution of the cubic equation and Newton's theorem on power sums, and yields an interpretation of generalized Lucas polynomials in terms of the theory of symmetric polynomials. We also develop additional connections that surface along the way; e.g., we give a parallel interpretation of generalized Fibonacci polynomials, and we show that Cardano's method can be used write down the roots of the Lucas polynomials.

math.HO

The Fundamental Theorem on Symmetric Polynomials: History's First Whiff of Galois Theory

We describe the Fundamental Theorem on Symmetric Polynomials (FTSP), exposit a classical proof, and offer a novel proof that arose out of an informal course on group theory. The paper develops this proof in tandem with the pedagogical context that led to it. We also discuss the role of the FTSP both as a lemma in the original historical development of Galois theory and as an early example of the connection between symmetry and expressibility that is described by the theory.

math.HO

Purely noncommuting groups

In this paper we define and investigate a class of groups characterized by a representation-theoretic property we call purely noncommuting or PNC. This property guarantees that the group has an action on a smooth projective variety with mild quotient singularities. It has intrinsic group-theoretic interest as well. The main results are as follows. (i) All supersolvable groups are PNC. (ii) No nonabelian finite simple groups are PNC. (iii) A metabelian group is guaranteed to be PNC if its commutator subgroup's cyclic prime-power-order factors are all distinct, but not in general. We also give a criterion guaranteeing a group is PNC if its nonabelian subgroups are all large, in a suitable sense, and investigate the PNC property for permutations.

math.RT

When are permutation invariants Cohen-Macaulay over all fields?

We prove that the polynomial invariants of a permutation group are Cohen-Macaulay for any choice of coefficient field if and only if the group is generated by transpositions, double transpositions, and 3-cycles. This unites and generalizes several previously known results. The "if" direction of the argument uses Stanley-Reisner theory and a recent result of Christian Lange in orbifold theory. The "only-if" direction uses a local-global result based on a theorem of Raynaud to reduce the problem to an analysis of inertia groups, and a combinatorial argument to identify inertia groups that obstruct Cohen-Macaulayness.

math.AC