SearcharxivSearch

arXiv subjects

Rustem Takhanov

Publications and source records attributed to Rustem Takhanov.

At least 19 recordsLinked to original sources

A quaternionic construction behind $841$-point kissing arrangement in ${\mathbb R}^{12}$

Recently, a new record kissing arrangement of $841$ points in $\mathbb R^{12}$ was obtained numerically by optimization (Takhanov-Assylbekov-Yun, 2026). The configuration was released as a coordinate file, without a mathematical description of its structure. The purpose of this paper is to provide such a description. The key observation is that the geometry becomes transparent once we regard $\mathbb R^{12}\cong \mathbb H^3$ as the Cartesian product of three copies of the quaternion algebra. We first introduce a new $840$-point kissing arrangement with a certain quaternionic structure. It consists of three mutually orthogonal regular $24$-cells, supported on the three quaternionic coordinate factors $\mathbb H\times\{0\}\times\{0\}$, $\{0\}\times\mathbb H\times\{0\}$, $\{0\}\times\{0\}\times\mathbb H$, together with two $384$-point families obtained by lifting affine sets of the form $$\{(u,v,w)\in (\mathbb F_2^2)^3\mid u+v+w=\eta\},$$ to quaternionic triples (whose components belong to the binary octahedral group $2O$) and then applying suitable component-wise rotations and weightings. A characteristic feature of this construction is a pronounced asymmetry among the three quaternionic factors. For the $816$ vectors obtained after removing the third $24$-cell, most of the squared norm is concentrated in the first two quaternionic coordinates, while the third coordinate carries systematically less mass. Thus, the third four-dimensional factor contains more available space than the first two. We then show that this $840$-point configuration provides a natural structural model for the numerical $841$-point record. Finally, we introduce a notion of the general quaternionic construction in dimensions divisible by $4$, and check that record kissing arrangements in ${\mathbb R}^{4k}$, $k\leq 5$, admit a quaternionic construction.

math.RA

Structure of kissing arrangements in ${\mathbb R}^{12}$ and a place for the $841$st sphere

Most currently known kissing arrangements of size $840$ in $\mathbb R^{12}$ share a common structure. They consist of $60$ vectors supported on $\mathbb R^6\times\{\mathbf 0\}$, another $60$ vectors supported on $\{\mathbf 0\}\times\mathbb R^6$, and $720$ additional \emph{bridge vectors}. The bridge vectors encode the interaction between the two six-dimensional factors and are constructed from the unique $1$-factorization of the complete graph $K_6$. In this paper we investigate kissing arrangements of this type while keeping the bridge vectors fixed. We show that each $60$-point block admits substantial flexibility: $12$ of its vectors may be chosen as the signed coordinate vectors $\pm e_i$, while the remaining $48$ vectors may vary within a positive-dimensional family of configurations, which we call $48$-systems. As a consequence, we obtain infinitely many pairwise non-isometric kissing arrangements of size $840$ in $\mathbb R^{12}$. The geometric freedom revealed by these constructions provides new insight into the local structure of extremal configurations. Exploiting this structure, we develop a specialized initialization scheme for logarithmic Riesz energy optimization. Starting from such structurally informed initial configurations, we numerically construct a kissing arrangement of size $841$ in $\mathbb R^{12}$.

cs.IT

Classification of independent sets in signed Johnson graphs and applications to kissing arrangements

Johnson graph are a family of graphs that play an important role in the theory of constant-weight codes, extremal combinatorics, and combinatorial geometry. We study signed analogues of classical Johnson graphs, denoted by $J_\pm(n,k)$, whose vertices are vectors of the form $\pm e_{i_1}\pm\cdots\pm e_{i_k}$, where two vertices are adjacent whenever their dot product equals $k-1$. We are particularly interested in maximum independent sets in the case $k=4$. An example of such an independent set in $J_\pm(n,4)$, which we call \emph{classical}, is obtained by lifting an arbitrary optimal $(n,4,4)$-code. Such independent sets naturally define kissing arrangements in ${\mathbb R}^n$. We develop an algorithm that is practical for computing all maximum independent sets in $J_\pm(n,4)$ up to signed permutations for $n\le 12$, $n\ne 11$. In addition to obtaining complete lists, we provide structural characterizations of all types of maximum independent sets in these dimensions, excluding $n=5$ and $n=11$. Our most striking results concern the case $n=12$. We identify $1579$ non-isomorphic maximum independent sets in $J_\pm(12,4)$, all corresponding to non-isometric kissing arrangements of size $840$ in ${\mathbb R}^{12}$. Structurally, $1575$ of these independent sets arise from three different constructions, the rest are liftings of one of four $(12,4,4)$-codes. To our knowledge, this is the first dimension in which such a large diversity of potentially optimal kissing arrangements has been observed. Beyond this finite range, we prove that for $n\equiv 2$ or $4 \pmod 6$, every maximum independent set arises from a Steiner quadruple system. We also obtain a characterization of the so-called \emph{nontrivially self-compatible} codes, namely optimal $(n,4,4)$-codes from which non-classical maximum independent sets can be constructed.

cs.IT

Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding

Conditionally positive definite (CPD) kernels are defined with respect to a function class $\mathcal{F}$. It is well known that such a kernel $K$ is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method -- called conditional kernel ridge regression (conditional KRR) due to its analogy with KRR -- where the estimated regression function is penalized by the square of its native space norm. This method is of interest because it can be viewed as classical linear regression, with features specified by $\mathcal{F}$, followed by the application of standard KRR to the residual (unexplained) component of the target variable. Methods of this type have recently attracted increasing attention. We study the statistical properties of this method by reducing its behavior to that of KRR with another fixed kernel, called the residual kernel. Our main theoretical result shows that such a reduction is indeed possible, at the cost of an additional term in the expected test risk, bounded by $\mathcal{O}(1/\sqrt{N})$, where $N$ is the sample size and the hidden constant depends on the class $\mathcal{F}$ and the input distribution. This reduction enables us to analyze conditional KRR in the case where $K$ is positive definite and $\mathcal{F}$ is given by the first $k$ principal eigenfunctions in the Mercer decomposition of $K$. We also consider the setting where $\mathcal{F}$ consists of $k$ random features from a random feature representation of $K$. It turns out that these two settings are closely related. Both our theoretical analysis and experiments confirm that conditional KRR outperforms standard KRR in these cases whenever the $\mathcal{F}$-component of the regression function is more pronounced than the residual part.

cs.LG

On the Intrinsic Dimensions of Data in Kernel Learning

The manifold hypothesis suggests that the generalization performance of machine learning methods improves significantly when the intrinsic dimension of the input distribution's support is low. In the context of KRR, we investigate two alternative notions of intrinsic dimension. The first, denoted $d_\rho$, is the upper Minkowski dimension defined with respect to the canonical metric induced by a kernel function $K$ on a domain $\Omega$. The second, denoted $d_K$, is the effective dimension, derived from the decay rate of Kolmogorov $n$-widths associated with $K$ on $\Omega$. Given a probability measure $\mu$ on $\Omega$, we analyze the relationship between these $n$-widths and eigenvalues of the integral operator $\phi \to \int_\Omega K(\cdot,x)\phi(x)d\mu(x)$. We show that, for a fixed domain $\Omega$, the Kolmogorov $n$-widths characterize the worst-case eigenvalue decay across all probability measures $\mu$ supported on $\Omega$. These eigenvalues are central to understanding the generalization behavior of constrained KRR, enabling us to derive an excess error bound of order $O(n^{-\frac{2+d_K}{2+2d_K} + \epsilon})$ for any $\epsilon > 0$, when the training set size $n$ is large. We also propose an algorithm that estimates upper bounds on the $n$-widths using only a finite sample from $\mu$. For distributions close to uniform, we prove that $\epsilon$-accurate upper bounds on all $n$-widths can be computed with high probability using at most $O\left(\epsilon^{-d_\rho}\log\frac{1}{\epsilon}\right)$ samples, with fewer required for small $n$. Finally, we compute the effective dimension $d_K$ for various fractal sets and present additional numerical experiments. Our results show that, for kernels such as the Laplace kernel, the effective dimension $d_K$ can be significantly smaller than the Minkowski dimension $d_\rho$, even though $d_K = d_\rho$ provably holds on regular domains.

cs.LG

Deep Linear Discriminant Analysis Revisited

We show that for unconstrained Deep Linear Discriminant Analysis (LDA) classifiers, maximum-likelihood training admits pathological solutions in which class means drift together, covariances collapse, and the learned representation becomes almost non-discriminative. Conversely, cross-entropy training yields excellent accuracy but decouples the head from the underlying generative model, leading to highly inconsistent parameter estimates. To reconcile generative structure with discriminative performance, we introduce the \emph{Discriminative Negative Log-Likelihood} (DNLL) loss, which augments the LDA log-likelihood with a simple penalty on the mixture density. DNLL can be interpreted as standard LDA NLL plus a term that explicitly discourages regions where several classes are simultaneously likely. Deep LDA trained with DNLL produces clean, well-separated latent spaces, matches the test accuracy of softmax classifiers on synthetic data and standard image benchmarks, and yields substantially better calibrated predictive probabilities, restoring a coherent probabilistic interpretation to deep discriminant models.

stat.ML

The informativeness of the gradient revisited

In the past decade gradient-based deep learning has revolutionized several applications. However, this rapid advancement has highlighted the need for a deeper theoretical understanding of its limitations. Research has shown that, in many practical learning tasks, the information contained in the gradient is so minimal that gradient-based methods require an exceedingly large number of iterations to achieve success. The informativeness of the gradient is typically measured by its variance with respect to the random selection of a target function from a hypothesis class. We use this framework and give a general bound on the variance in terms of a parameter related to the pairwise independence of the target function class and the collision entropy of the input distribution. Our bound scales as $ \tilde{\mathcal{O}}(\varepsilon+e^{-\frac{1}{2}\mathcal{E}_c}) $, where $ \tilde{\mathcal{O}} $ hides factors related to the regularity of the learning model and the loss function, $ \varepsilon $ measures the pairwise independence of the target function class and $\mathcal{E}_c$ is the collision entropy of the input distribution. To demonstrate the practical utility of our bound, we apply it to the class of Learning with Errors (LWE) mappings and high-frequency functions. In addition to the theoretical analysis, we present experiments to understand better the nature of recent deep learning-based attacks on LWE.

cs.LG

Multi-layer random features and the approximation power of neural networks

A neural architecture with randomly initialized weights, in the infinite width limit, is equivalent to a Gaussian Random Field whose covariance function is the so-called Neural Network Gaussian Process kernel (NNGP). We prove that a reproducing kernel Hilbert space (RKHS) defined by the NNGP contains only functions that can be approximated by the architecture. To achieve a certain approximation error the required number of neurons in each layer is defined by the RKHS norm of the target function. Moreover, the approximation can be constructed from a supervised dataset by a random multi-layer representation of an input vector, together with training of the last layer's weights. For a 2-layer NN and a domain equal to an $n-1$-dimensional sphere in ${\mathbb R}^n$, we compare the number of neurons required by Barron's theorem and by the multi-layer features construction. We show that if eigenvalues of the integral operator of the NNGP decay slower than $k^{-n-\frac{2}{3}}$ where $k$ is an order of an eigenvalue, then our theorem guarantees a more succinct neural network approximation than Barron's theorem. We also make some computational experiments to verify our theoretical findings. Our experiments show that realistic neural networks easily learn target functions even when both theorems do not give any guarantees.

cs.LG

On the induced problem for fixed-template CSPs

The Constraint Satisfaction Problem (CSP) is a problem of computing a homomorphism $\mathbf{R}\to \mathbfΓ$ between two relational structures, where $\mathbf{R}$ is defined over a domain $V$ and $\mathbfΓ$ is defined over a domain $D$. In a fixed template CSP, denoted $\rm{CSP}(\mathbfΓ)$, the right side structure $\mathbfΓ$ is fixed and the left side structure $\mathbf{R}$ is unconstrained. In the last two decades it was discovered that the reasons that make fixed template CSPs polynomially solvable are of algebraic nature, namely, templates that are tractable should be preserved under certain polymorphisms. From this perspective the following problem looks natural: given a prespecified finite set of algebras ${\mathcal B}$ whose domain is $D$, is it possible to present the solution set of a given instance of $\rm{CSP}(\mathbfΓ)$ as a subalgebra of ${\mathbb A}_1\times ... \times {\mathbb A}_{|V|}$ where ${\mathbb A}_i\in {\mathcal B}$? We study this problem and show that it can be reformulated as an instance of a certain fixed-template CSP over another template $\mathbfΓ^{\mathcal B}$. We study conditions under which $\rm{CSP}(\mathbfΓ)$ can be reduced to $\rm{CSP}(\mathbfΓ^{\mathcal B})$. This issue is connected with the so-called CSP with an input prototype, formulated in the following way: given a homomorphism from $\mathbf{R}$ to $\mathbfΓ^{\mathcal B}$ find a homomorphism from $\mathbf{R}$ to $\mathbfΓ$. We prove that if ${\mathcal B}$ contains only tractable algebras, then the latter CSP with an input prototype is tractable. We also prove that $\rm{CSP}(\mathbfΓ^{\mathcal B})$ can be reduced to $\rm{CSP}(\mathbfΓ)$ if the set ${\mathcal B}$, treated as a relation over $D$, can be expressed as a primitive positive formula over $\mathbfΓ$.

cs.CC

Gradient Descent Fails to Learn High-frequency Functions and Modular Arithmetic

Classes of target functions containing a large number of approximately orthogonal elements are known to be hard to learn by the Statistical Query algorithms. Recently this classical fact re-emerged in a theory of gradient-based optimization of neural networks. In the novel framework, the hardness of a class is usually quantified by the variance of the gradient with respect to a random choice of a target function. A set of functions of the form $x\to ax \bmod p$, where $a$ is taken from ${\mathbb Z}_p$, has attracted some attention from deep learning theorists and cryptographers recently. This class can be understood as a subset of $p$-periodic functions on ${\mathbb Z}$ and is tightly connected with a class of high-frequency periodic functions on the real line. We present a mathematical analysis of limitations and challenges associated with using gradient-based learning techniques to train a high-frequency periodic function or modular multiplication from examples. We highlight that the variance of the gradient is negligibly small in both cases when either a frequency or the prime base $p$ is large. This in turn prevents such a learning algorithm from being successful.

cs.LG

Intractability of Learning the Discrete Logarithm with Gradient-Based Methods

The discrete logarithm problem is a fundamental challenge in number theory with significant implications for cryptographic protocols. In this paper, we investigate the limitations of gradient-based methods for learning the parity bit of the discrete logarithm in finite cyclic groups of prime order. Our main result, supported by theoretical analysis and empirical verification, reveals the concentration of the gradient of the loss function around a fixed point, independent of the logarithm's base used. This concentration property leads to a restricted ability to learn the parity bit efficiently using gradient-based methods, irrespective of the complexity of the network architecture being trained. Our proof relies on Boas-Bellman inequality in inner product spaces and it involves establishing approximate orthogonality of discrete logarithm's parity bit functions through the spectral norm of certain matrices. Empirical experiments using a neural network-based approach further verify the limitations of gradient-based learning, demonstrating the decreasing success rate in predicting the parity bit as the group order increases.

cs.LG

Autoencoders for a manifold learning problem with a Jacobian rank constraint

We formulate the manifold learning problem as the problem of finding an operator that maps any point to a close neighbor that lies on a ``hidden'' $k$-dimensional manifold. We call this operator the correcting function. Under this formulation, autoencoders can be viewed as a tool to approximate the correcting function. Given an autoencoder whose Jacobian has rank $k$, we deduce from the classical Constant Rank Theorem that its range has a structure of a $k$-dimensional manifold. A $k$-dimensionality of the range can be forced by the architecture of an autoencoder (by fixing the dimension of the code space), or alternatively, by an additional constraint that the rank of the autoencoder mapping is not greater than $k$. This constraint is included in the objective function as a new term, namely a squared Ky-Fan $k$-antinorm of the Jacobian function. We claim that this constraint is a factor that effectively reduces the dimension of the range of an autoencoder, additionally to the reduction defined by the architecture. We also add a new curvature term into the objective. To conclude, we experimentally compare our approach with the CAE+H method on synthetic and real-world datasets.

cs.LG

Computing a partition function of a generalized pattern-based energy over a semiring

Valued constraint satisfaction problems with ordered variables (VCSPO) are a special case of Valued CSPs in which variables are totally ordered and soft constraints are imposed on tuples of variables that do not violate the order. We study a restriction of VCSPO, in which soft constraints are imposed on a segment of adjacent variables and a constraint language $Γ$ consists of $\{0,1\}$-valued characteristic functions of predicates. This kind of potentials generalizes the so-called pattern-based potentials, which were applied in many tasks of structured prediction. For a constraint language $Γ$ we introduce a closure operator, $ \overline{Γ^{\cap}}\supseteq Γ$, and give examples of constraint languages for which $|\overline{Γ^{\cap}}|$ is small. If all predicates in $Γ$ are cartesian products, we show that the minimization of a generalized pattern-based potential (or, the computation of its partition function) can be made in ${\mathcal O}(|V|\cdot |D|^2 \cdot |\overline{Γ^{\cap}}|^2 )$ time, where $V$ is a set of variables, $D$ is a domain set. If, additionally, only non-positive weights of constraints are allowed, the complexity of the minimization task drops to ${\mathcal O}(|V|\cdot |\overline{Γ^{\cap}}| \cdot |D| \cdot \max_{ρ\in Γ}\|ρ\|^2 )$ where $\|ρ\|$ is the arity of $ρ\in Γ$. For a general language $Γ$ and non-positive weights, the minimization task can be carried out in ${\mathcal O}(|V|\cdot |\overline{Γ^{\cap}}|^2)$ time. We argue that in many natural cases $\overline{Γ^{\cap}}$ is of moderate size, though in the worst case $|\overline{Γ^{\cap}}|$ can blow up and depend exponentially on $\max_{ρ\in Γ}\|ρ\|$.

cs.AI

The algebraic structure of the densification and the sparsification tasks for CSPs

The tractability of certain CSPs for dense or sparse instances is known from the 90s. Recently, the densification and the sparsification of CSPs were formulated as computational tasks and the systematical study of their computational complexity was initiated. We approach this problem by introducing the densification operator, i.e. the closure operator that, given an instance of a CSP, outputs all constraints that are satisfied by all of its solutions. According to the Galois theory of closure operators, any such operator is related to a certain implicational system (or, a functional dependency) $Σ$. We are specifically interested in those classes of fixed-template CSPs, parameterized by constraint languages $Γ$, for which there is an implicational system $Σ$ whose size is a polynomial in the number of variables $n$. We show that in the Boolean case, such implicational systems exist if and only if $Γ$ is of bounded width. For such languages, $Σ$ can be computed in log-space or in a logarithmic time with a polynomial number of processors. Given an implicational system $Σ$, the densification task is equivalent to the computation of the closure of input constraints. The sparsification task is equivalent to the computation of the minimal key.

cs.CC

Reducing the dimensionality of data using tempered distributions

We reformulate unsupervised dimension reduction problem (UDR) in the language of tempered distributions, i.e. as a problem of approximating an empirical probability density function by another tempered distribution, supported in a $k$-dimensional subspace. We show that this task is connected with another classical problem of data science -- the sufficient dimension reduction problem (SDR). In fact, an algorithm for the first problem induces an algorithm for the second and vice versa. In order to reduce an optimization problem over distributions to an optimization problem over ordinary functions we introduce a nonnegative penalty function that ``forces'' the support of the model distribution to be $k$-dimensional. Then we present an algorithm for the minimization of the penalized objective, based on the infinite-dimensional low-rank optimization, which we call the alternating scheme. Also, we design an efficient approximate algorithm for a special case of the problem, where the distance between the empirical distribution and the model distribution is measured by Maximum Mean Discrepancy defined by a Mercer kernel of a certain type. We test our methods on four examples (three UDR and one SDR) using synthetic data and standard datasets.

math.ST

On the speed of uniform convergence in Mercer's theorem

The classical Mercer's theorem claims that a continuous positive definite kernel $K({\mathbf x}, {\mathbf y})$ on a compact set can be represented as $\sum_{i=1}^\infty λ_iϕ_i({\mathbf x})ϕ_i({\mathbf y})$ where $\{(λ_i,ϕ_i)\}$ are eigenvalue-eigenvector pairs of the corresponding integral operator. This infinite representation is known to converge uniformly to the kernel $K$. We estimate the speed of this convergence in terms of the decay rate of eigenvalues and demonstrate that for $2m$ times differentiable kernels the first $N$ terms of the series approximate $K$ as $\mathcal{O}\big((\sum_{i=N+1}^\inftyλ_i)^{\frac{m}{m+n}}\big)$ or $\mathcal{O}\big((\sum_{i=N+1}^\inftyλ^2_i)^{\frac{m}{2m+n}}\big)$. Finally, we demonstrate some applications of our results to a spectral charaterization of integral operators with continuous roots and other powers.

cs.LG

Non-asymptotic spectral bounds on the $\varepsilon$-entropy of kernel classes

Let $K: \boldsymbol{\Omega}\times \boldsymbol{\Omega}$ be a continuous Mercer kernel defined on a compact subset of ${\mathbb R}^n$ and $\mathcal{H}_K$ be the reproducing kernel Hilbert space (RKHS) associated with $K$. Given a finite measure $\nu$ on $\boldsymbol{\Omega}$, we investigate upper and lower bounds on the $\varepsilon$-entropy of the unit ball of $\mathcal{H}_K$ in the space $L_p(\nu)$. This topic is an important direction in the modern statistical theory of kernel-based methods. We prove sharp upper and lower bounds for $p\in [1,+\infty]$. For $p\in [1,2]$, the upper bounds are determined solely by the eigenvalue behaviour of the corresponding integral operator $\phi\to \int_{\boldsymbol{\Omega}} K(\cdot,{\mathbf y})\phi({\mathbf y})d\nu({\mathbf y})$. In constrast, for $p>2$, the bounds additionally depend on the convergence rate of the truncated Mercer series to the kernel $K$ in the $L_p(\nu)$-norm. We discuss a number of consequences of our bounds and show that they are substantially tighter than previous bounds for general kernels. Furthermore, for specific cases, such as zonal kernels and the Gaussian kernel on a box, our bounds are asymptotically tight as $\varepsilon\to +0$.

stat.ML

How many moments does MMD compare?

We present a new way of study of Mercer kernels, by corresponding to a special kernel $K$ a pseudo-differential operator $p({\mathbf x}, D)$ such that $\mathcal{F} p({\mathbf x}, D)^†p({\mathbf x}, D) \mathcal{F}^{-1}$ acts on smooth functions in the same way as an integral operator associated with $K$ (where $\mathcal{F}$ is the Fourier transform). We show that kernels defined by pseudo-differential operators are able to approximate uniformly any continuous Mercer kernel on a compact set. The symbol $p({\mathbf x}, {\mathbf y})$ encapsulates a lot of useful information about the structure of the Maximum Mean Discrepancy distance defined by the kernel $K$. We approximate $p({\mathbf x}, {\mathbf y})$ with the sum of the first $r$ terms of the Singular Value Decomposition of $p$, denoted by $p_r({\mathbf x}, {\mathbf y})$. If ordered singular values of the integral operator associated with $p({\mathbf x}, {\mathbf y})$ die down rapidly, the MMD distance defined by the new symbol $p_r$ differs from the initial one only slightly. Moreover, the new MMD distance can be interpreted as an aggregated result of comparing $r$ local moments of two probability distributions. The latter results holds under the condition that right singular vectors of the integral operator associated with $p$ are uniformly bounded. But even if this is not satisfied we can still hold that the Hilbert-Schmidt distance between $p$ and $p_r$ vanishes. Thus, we report an interesting phenomenon: the MMD distance measures the difference of two probability distributions with respect to a certain number of local moments, $r^\ast$, and this number $r^\ast$ depends on the speed with which singular values of $p$ die down.

cs.LG