SearcharxivSearch

arXiv subjects

Lek-Heng Lim

Publications and source records attributed to Lek-Heng Lim.

At least 19 recordsLinked to original sources

Pierce-Birkhoff conjecture is false

We provide a counterexample to the Pierce-Birkhoff conjecture: a continuous function that is piecewise quadratic on a finite collection of semialgebraic sets that partition $\mathbb{R}^n$ but that cannot be expressed as a finite lattice combination of polynomials. The counterexample was found with the assistance of our multi-agent, multi-model setup that chains together GPT 5.6, GPT 6, Claude Opus 5, and Claude Fable 5.1.

math.AG

The Grassmannian of indefinite subspaces

The homogeneous space $\operatorname{O}_{m,n}(\mathbb{R})/(\operatorname{O}_{p,q}(\mathbb{R}) \times \operatorname{O}_{m-p,n-q}(\mathbb{R}))$ is an object that has received scant attention, with just two brief mentions in existing literature, and christened the indefinite Grassmannian in one of them. In this article, we develop some of its basic properties, building it from ground up. We will see that, aside from its homogeneous space description, the indefinite Grassmannian may be characterized in several other ways: set-theoretically, it is the manifold of indefinite $(p+q)$-dimensional subspaces in $(m +n)$-dimensional space; it is an adjoint orbit of a Lie group; a semialgebraic smooth manifold of matrices; a base space of a principal bundle whose total space is the indefinite Stiefel manifold, a natural corresponding notion. It may also be naturally equipped with various structures, turning the indefinite Grassmannian into a pseudo-Riemannian manifold; a Einstein manifold; a symplectic manifold; and a pseudo-K\"ahler manifold (last two only over $\mathbb{C}$). As a centerpiece of this article, we establish two attributes of the indefinite Grassmannian in relation to the standard Grassmannian: (i) any Grassmannian has a Whitney stratification whose highest dimensional strata are indefinite Grassmannians; (ii) the indefinite Grassmannian is a strong deformation retract of a product of two Grassmannians, thereby allowing us to completely ascertain the topology of the former. We will also determine some of the indefinite Grassmannian's features that are within reach --- both geometric (Riemann, Ricci, sectional, and scalar curvatures; second fundamental form) and topological (cohomology ring, homotopy groups, characteristic classes).

math.DG

Low-dimensional topology of deep neural networks

We study layered models, including feedforward networks, ResNets, and transformers, by limiting each layer to a width of $d = 3$, i.e., $\mathbb{R}^3$ as representation space. This allows us to track how a neural network changes low-dimensional topological invariants through its layers. Just about any topological structure may be simplified or even trivialized by simply increasing dimension; e.g., any knot is equivalent to an unknot in $\mathbb{R}^4$. By restricting to $\mathbb{R}^3$, we not only isolate the effects of activation and depth from that of width, we work in a space that lends itself to easy visualization. We focus on linking number here, deferring other invariants like link groups, Milnor's $\bar{\mu}$-invariants, knot types, ambient cobordisms, to a sequel. We provide full proofs and empirical experiments to justify the following insights: When measured by their power to effect changes in linking numbers, the layer-skipping feature in ResNets is as powerful as the attention mechanism in transformers; both ResNets and transformers are strictly more powerful than feedforward neural networks with monotonic activations, which are in turn more powerful than invertible and flow-based models; but replacing monotonic activation with a nonmonotonic one elevates a feedforward network into the same expressivity class as ResNets and transformers. These results suggest that low-dimensional topology can be a useful tool to guide designs of AI architectures. We also generalize our results from $d = 3$ to arbitrary $d > 3$.

cs.LG

Generalized matrix nearness problems II

Given a matrix $A$, a matrix nearness problem seeks an $X$ that most closely approximates $A$ in the sense of minimizing $\lVert A - X\rVert$ under a variety of constraints on $X$. A generalized matrix nearness problem seeks the same but with three given matrices $A,B,C$ and $\lVert A - BXC\rVert$ in place of $\lVert A - X\rVert$. We extend previous studies of the latter problem in three directions: incorporating an affine term, replacing matrix product by Kronecker product in various manners, and generalizing Frobenius norm to any orthogonally invariant norm. We will solve several of these in closed form. For the rest, we develop an iterative algorithm that works for any Schatten norm, proving that it converges to a global minimizer regardless of the initial point. In addition, the algorithm relies purely on numerical linear algebra, and notably does not compute any explicit gradients or subgradients. Along the way, we will also show that there is no Mirsky-type theorem for rank constrained generalized matrix nearness problems.

math.NA

Linear representations of manifolds

A finite-dimensional linear representation of a group or an algebra may be regarded as a map into a space of matrices, endowing abstract elements with coordinates, and encoding algebraic operations as matrix products. With this in mind, we define a linear representation of a $\mathsf{G}$-manifold $\mathcal{M} $ as a map into a space of matrices, representing points as matrices and the $\mathsf{G}$-action as matrix products. We show that this generalizes group representations to any $\mathsf{G}$-manifold that may not have a group structure, with homogeneous spaces $\mathsf{G}/\mathsf{H}$ an important special case; and in this case it also generalizes Cartan embeddings of symmetric spaces to more general $\mathsf{G}/\mathsf{H}$. To demonstrate the utility of such manifold representations, we use them to provide effective bounds for Mostow-Palais $\mathsf{G}$-equivariant embeddings of $\mathsf{G}$-manifolds into $\mathsf{G}$-modules $\mathbb{V}$. Unlike Whitney and Nash embeddings, Mostow-Palais embeddings have no known effective bounds; before our work, it was only known that $\dim \mathbb{V} < \infty$ if $\mathsf{G}$ is compact. We will give explicit values for $\dim \mathbb{V}$ and show that our bounds are sharp. Furthermore, our method is constructive, giving explicit expressions for these minimal-dimensional Mostow-Palais embeddings.

math.DG

Eigen, singular, cosine-sine, and Autonne--Takagi vectors distributions of random matrix ensembles

We show that some of the best-known matrix decompositions of some of the best-known random matrix ensembles give us the unique $G$-invariant uniform distributions on some of the best-known manifolds. The eigenvectors distributions of the Gaussian, Laguerre, and Jacobi ensembles are all given by the uniform distribution on the complete flag manifold. The singular vectors distributions of Ginibre ensembles are given by the uniform distribution on a product of the complete flag manifold with a Stiefel manifold. Circular ensembles split into two types: The cosine-sine vectors distributions of circular real, unitary, and quaternionic ensembles are given by the uniform distributions on products of a (partial) flag manifold with copies of the orthogonal, unitary, or compact symplectic groups. The Autonne--Takagi vectors distributions of circular orthogonal, Lagrangian, and symplectic ensembles are given by the uniform distributions on Lagrangian Grassmannians.

math.PR

Pierce-Birkhoff conjecture is true for splines

We prove the Pierce--Birkhoff conjecture for splines, i.e., continuous piecewise polynomials of degree $d$ in $n$ variables on a hyperplane partition of $\mathbb{R}^n$, can be written as a finite lattice combination of polynomials. We will provide a purely existential proof, followed by a more in-depth analysis that yields effective bounds.

math.AG

Stiefel optimization is NP-hard

We show that linearly constrained linear optimization over a Stiefel or Grassmann manifold is NP-hard in general. We show that the same is true for unconstrained quadratic optimization over a Stiefel manifold. We will show that unless $\mathrm{P}=\mathrm{NP}$, these optimization problems over a Stiefel manifold do not have $\mathrm{FPTAS}$. As an aside we extend our results to flag manifolds. Combined with earlier findings, this shows that manifold optimization is a difficult endeavor -- even the simplest problems like LP and unconstrained QP are already NP-hard on the most common manifolds.

math.OC

Geometric Programming for 3D Circuits

With the soaring demand for high-performing integrated circuits, 3D integrated circuits (ICs) have emerged as a promising alternative to traditional planar structures. Unlike existing 3D ICs that stack 2D layers, a full 3D IC features cubic circuit elements unrestricted by layers, offering greater design freedom. Design problems such as floorplanning, transistor sizing, and interconnect sizing are highly complex due to the 3D nature of the circuits and unavoidably require systematic approaches. We introduce geometric programming to solve these design optimization problems systematically and efficiently.

math.OC

Glivenko-Cantelli for $f$-divergence

We extend the celebrated Glivenko-Cantelli theorem, sometimes called the fundamental theorem of statistics, from its standard setting of total variation distance to all $f$-divergences. A key obstacle in this endeavor is to define $f$-divergence on a subcollection of a $\sigma$-algebra that forms a $\pi$-system but not a $\sigma$-subalgebra. This is a side contribution of our work. We will show that this notion of $f$-divergence on the $\pi$-system of rays preserves nearly all known properties of standard $f$-divergence, yields a novel integral representation of the Kolmogorov-Smirnov distance, and has a Glivenko-Cantelli theorem. We will also discuss the prospects of a Vapnik-Chervonenkis theory for $f$-divergence.

math.ST

Euclidean distance degree in manifold optimization

We determine the Euclidean distance degrees of the three most common manifolds arising in manifold optimization: flag, Grassmann, and Stiefel manifolds. For the Grassmannian, we will also determine the Euclidean distance degree of an important class of Schubert varieties that often appear in applications. Our technique goes further than furnishing the value of the Euclidean distance degree; it will also yield closed-form expressions for all stationary points of the Euclidean distance function in each instance. We will discuss the implications of these results on the tractability of manifold optimization problems.

math.OC

Special orthogonal, special unitary, and symplectic groups as products of Grassmannians

We describe a curious structure of the special orthogonal, special unitary, and symplectic groups that has not been observed, namely, they can be expressed as matrix products of their corresponding Grassmannians realized as involution matrices. We will show that $\operatorname{SO}(n)$ is a product of two real Grassmannians, $\operatorname{SU}(n)$ a product of four complex Grassmannians, and $\operatorname{Sp}(2n, \mathbb{R})$ or $\operatorname{Sp}(2n, \mathbb{C})$ a product of four symplectic Grassmannians over $\mathbb{R}$ or $\mathbb{C}$ respectively.

math.RT

Summing divergent matrix series

We extend several celebrated methods in classical analysis for summing series of complex numbers to series of complex matrices. These include the summation methods of Abel, Borel, Cesáro, Euler, Lambert, Nörlund, and Mittag-Leffler, which are frequently used to sum scalar series that are divergent in the conventional sense. One feature of our matrix extensions is that they are fully noncommutative generalizations of their scalar counterparts -- not only is the scalar series replaced by a matrix series, positive weights are replaced by positive definite matrix weights, order on $\mathbb{R}$ replaced by Loewner order, exponential function replaced by matrix exponential function, etc. We will establish the regularity of our matrix summation methods, i.e., when applied to a matrix series convergent in the conventional sense, we obtain the same value for the sum. Our second goal is to provide numerical algorithms that work in conjunction with these summation methods. We discuss how the block and mixed-block summation algorithms, the Kahan compensated summation algorithm, may be applied to matrix sums with similar roundoff error bounds. These summation methods and algorithms apply not only to power or Taylor series of matrices but to any general matrix series including matrix Fourier and Dirichlet series. We will demonstrate the utility of these summation methods: establishing a Fejér's theorem and alleviating the Gibbs phenomenon for matrix Fourier series; extending the domains of matrix functions and accurately evaluating them; enhancing the matrix Padé approximation and Schur--Parlett algorithms; and more.

math.NA

Grassmannian optimization is NP-hard

We show that unconstrained quadratic optimization over a Grassmannian $\operatorname{Gr}(k,n)$ is NP-hard. Our results cover all scenarios: (i) when $k$ and $n$ are both allowed to grow; (ii) when $k$ is arbitrary but fixed; (iii) when $k$ is fixed at its lowest possible value $1$. We then deduce the NP-hardness of unconstrained cubic optimization over the Stiefel manifold $\operatorname{V}(k,n)$ and the orthogonal group $\operatorname{O}(n)$. As an addendum we demonstrate the NP-hardness of unconstrained quadratic optimization over the Cartan manifold, i.e., the positive definite cone $\mathbb{S}^n_{\scriptscriptstyle++}$ regarded as a Riemannian manifold, another popular example in manifold optimization. We will also establish the nonexistence of $\mathrm{FPTAS}$ in all cases.

math.OC

Attention is a smoothed cubic spline

We highlight a perhaps important but hitherto unobserved insight: The attention module in a transformer is a smoothed cubic spline. Viewed in this manner, this mysterious but critical component of a transformer becomes a natural development of an old notion deeply entrenched in classical approximation theory. More precisely, we show that with ReLU-activation, attention, masked attention, encoder-decoder attention are all cubic splines. As every component in a transformer is constructed out of compositions of various attention modules (= cubic splines) and feed forward neural networks (= linear splines), all its components -- encoder, decoder, and encoder-decoder blocks; multilayered encoders and decoders; the transformer itself -- are cubic or higher-order splines. If we assume the Pierce-Birkhoff conjecture, then the converse also holds, i.e., every spline is a ReLU-activated encoder. Since a spline is generally just $C^2$, one way to obtain a smoothed $C^\infty$-version is by replacing ReLU with a smooth activation; and if this activation is chosen to be SoftMax, we recover the original transformer as proposed by Vaswani et al. This insight sheds light on the nature of the transformer by casting it entirely in terms of splines, one of the best known and thoroughly understood objects in applied mathematics.

cs.AI

Simple matrix models for the flag, Grassmann, and Stiefel manifolds

We derive three families of orthogonally-equivariant matrix submanifold models for the Grassmann, flag, and Stiefel manifolds respectively. These families are exhaustive -- every orthogonally-equivariant submanifold model of the lowest dimension for any of these manifolds is necessarily a member of the respective family, with a small number of exceptions. They have several computationally desirable features. The orthogonal equivariance allows one to obtain, for various differential geometric objects and operations, closed-form analytic expressions that are readily computable with standard numerical linear algebra. The minimal dimension aspect translates directly to a speed advantage in computations. And having an exhaustive list of all possible matrix models permits one to identify the model with the lowest matrix condition number, which translates to an accuracy advantage in computations. As an interesting aside, we will see that the family of models for the Stiefel manifold is naturally parameterized by the Cartan manifold, i.e., the positive definite cone equipped with its natural Riemannian metric.

math.NA

Minimal equivariant embeddings of the Grassmannian and flag manifold

We show that the flag manifold $\operatorname{Flag}(k_1,\dots, k_p, \mathbb{R}^n)$, with Grassmannian the special case $p=1$, has an $\operatorname{SO}_n(\mathbb{R})$-equivariant embedding in an Euclidean space of dimension $(n-1)(n+2)/2$, two orders of magnitude below the current best known result. We will show that the value $(n-1)(n+2)/2$ is the smallest possible and that any $\operatorname{SO}_n(\mathbb{R})$-equivariant embedding of $\operatorname{Flag}(k_1,\dots, k_p, \mathbb{R}^n)$ in an ambient space of minimal dimension is equivariantly equivalent to the aforementioned one.

math.RT

Simple matrix expressions for the curvatures of Grassmannian

We show that modeling a Grassmannian as symmetric orthogonal matrices $\operatorname{Gr}(k,\mathbb{R}^n) \cong\{Q \in \mathbb{R}^{n \times n} : Q^{\scriptscriptstyle\mathsf{T}} Q = I, \; Q^{\scriptscriptstyle\mathsf{T}} = Q,\; \operatorname{tr}(Q)=2k - n\}$ yields exceedingly simple matrix formulas for various curvatures and curvature-related quantities, both intrinsic and extrinsic. These include Riemann, Ricci, Jacobi, sectional, scalar, mean, principal, and Gaussian curvatures; Schouten, Weyl, Cotton, Bach, Pleba\'nski, cocurvature, nonmetricity, and torsion tensors; first, second, and third fundamental forms; Gauss and Weingarten maps; and upper and lower delta invariants. We will derive explicit, simple expressions for the aforementioned quantities in terms of standard matrix operations that are stably computable with numerical linear algebra. Many of these aforementioned quantities have never before been presented for the Grassmannian.

math.DG