SearcharxivSearch

arXiv subjects

Yitzchak Shmalo

Publications and source records attributed to Yitzchak Shmalo.

18 recordsLinked to original sources

Entropy and singularity of spectral and pivotal sets in planar percolation

The spectral sample and pivotal set of a percolation crossing have the same one- and two-coordinate inclusion probabilities under the uniform product measure. We prove that both Shannon entropies are comparable to total influence for critical square-lattice bond and triangular-lattice site crossings, with matching estimates at every retained nonroot spatial resolution. The triangular entropy bounds are uniform over a class of inhomogeneous near-critical product measures. For critical triangular-site square crossings, the two laws are asymptotically mutually singular: an elementary triangular face is forbidden in the pivotal set and occurs with positive density in a large spectral sample. Every fixed face density below an explicit threshold has nonempty spectral probability comparable to $R^2α_4(R)^2$, while the nonempty overlap of the two laws has decay exponent $11/12$. Their minimum expected symmetric-difference distance over all couplings is comparable to total influence. Geometric separation persists after independent thinning whenever the retention probability $ρ_R$ satisfies $ρ_R^3R^2α_4(R)\to\infty$; the face detector fails when this quantity tends to zero. The proofs combine spatial encoding, arm estimates, local Fourier cancellation, and an exact six-cycle coercivity inequality preserved under exterior conditioning.

math.PR

Mesoscopic Rectangular Spikes under Subspace Local Laws: Outlier Values and Singular Subspaces

A rectangular data matrix is often modelled as noise plus a signal of low rank. When that rank is fixed the picture is classical: a signal direction must exceed a critical strength before it produces a singular value outside the noise bulk, and above the threshold the singular vectors of the observation retain a definite, computable fraction of the planted direction. We ask what survives when the number of signal directions grows with the dimension. Everything here rests on one hypothesis, which we call a subspace local law: seen from inside the signal subspace, the resolvent of the noise should look like a scalar. We show that this hypothesis alone locates the outliers and counts them, and, in a holomorphic form, determines the outlier singular subspaces as well. The conclusions are stated through spectral projectors rather than individual singular vectors, so they remain meaningful when spike strengths collide or come closer together than the error of the approximation, which at growing rank they must. We then verify the hypothesis in four genuinely different settings, from deterministic noise viewed through randomly oriented signal directions to independent entries with fixed deterministic ones, and for Marchenko--Pastur noise we compute every constant explicitly.

math.PR

Extreme least singular values of random row submatrices with bounded-density subgaussian entries

Let $ξ$ be a centered real subgaussian random variable with positive variance and a bounded Lebesgue density, and let $A_m\in\mathbb{R}^{N_m\times m}$ have independent entries distributed as $ξ$, where $N_m/m\toγ>1$. For each set $I\subset[N_m]$ with $|I|=m$, let $(A_m)_I$ denote the row submatrix indexed by $I$, and define $M_m(A_m):=\min_{I\subset[N_m],\,|I|=m}σ_{\min}((A_m)_I)$. We determine its exponential scale: $\frac{1}{m}\log M_m(A_m)\xrightarrow{\mathbb{P}}-h(γ)$, where $h(γ):=γ\logγ-(γ-1)\log(γ-1)$. This extends the corresponding real Gaussian result. The main new ingredient is an upper-tail argument that avoids uniform control over exponentially many random hyperplanes. We combine a density-level local central limit theorem for delocalized directions, an averaged delocalization estimate for hyperplane normals, an exponential bound for nearly parallel pairs, and amplification using a linear number of independent probe rows. For every fixed $\varepsilon\in(0,h(γ))$, the probability of an $\varepsilon$-deviation is at most $C\exp(-c\sqrt{m})$ for all sufficiently large $m$. Under the canonical coupling induced by a single infinite i.i.d. array, this summable deviation estimate yields a uniform almost-sure exponential law over every compact range of aspect ratios. In particular, at the real phase-retrieval threshold $N_m=2m-1$, the Balan--Wang stability parameter has exponential base $1/4$ in probability and, under this coupling, almost surely.

math.PR

Randomly initialized autoencoders: fixed points and edge-of-chaos

In this paper we study autoencoders, a special class of deep neural nets (DNNs) whose performance can be characterized via their fixed points. This perspective naturally raises questions of existence, stability, and basins of attraction of these fixed points. These questions are addressed via the contractive properties of autoencoders, and are closely related to the notion of edge-of-chaos. Edge-of-chaos (EoC) is an important notion in the theory of DNNs. It describes the critical regime separating ordered and chaotic signal propagation through a randomly initialized network. Initialization at or near this critical regime offers several theoretical and practical advantages, including stability of the network w.r.t. perturbations of the input. EoC was previously introduced for broad classes of neural networks using mean-field averaging methods. In this paper we modify the notion of EoC for the study of autoencoders. Specifically, we introduce local and global EoC for autoencoders that control local (small) and global (arbitrary) perturbations of the input respectively. The study of stability of autoencoders falls within the scope of nonlinear problems in Random Matrix Theory (RMT). Our analysis of local EoC is based on spectral techniques of RMT, whereas global EoC is studied by employing Sudakov-Fernique inequality for Gaussian processes.

cs.LG

The storage capacity of the Ising perceptron: verification of the outstanding numerical conditions

Krauth and Mézard predicted in 1989 that the storage capacity of the Ising perceptron at zero margin is an explicit constant $α_\star\approx0.8330786$. Let $M_N$ be the largest number of random patterns that can be stored by an $N$-dimensional Ising perceptron. We give a computer-assisted proof that \[ \frac{M_N}{N}\xrightarrow{\mathbb P}α_\star, \qquad α_\star\in[0.833078599,0.833078600]. \] Previous work established matching conditional lower and upper bounds, subject respectively to a one-variable global sign condition of Ding--Sun and a two-variable global sign condition of Huang. We rigorously verify both conditions using Arb ball arithmetic. For Huang's condition, a moment-coordinate reparametrization compresses the unbounded parameter plane onto a compact convex body. Convex duality and certified adaptive sweeps control its bulk, while a ray-concavity argument treats the degenerate maximizer. We also re-establish the shared parameter rectangle and verify the full Ding--Sun condition, including its curvature and endpoint requirements. Combining these verifications with the existing sharp-threshold and universality theorems proves the result for Gaussian disorder and for every fixed i.i.d.\ mean-zero, unit-variance subgaussian disorder law, including Bernoulli disorder. The complete verification programs, certificates, and raw records accompany the paper.

math.PR

A law of robustness for two-layer neural networks with arbitrary weights

Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. The two-layer unbounded-weight case requires a different argument. We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise-linear activation, in particular for ReLU networks. For data drawn uniformly from $\mathbb{S}^{d-1}$, $d\ge3$, or from $N(0,I_d/d)$, labels in $[-1,1]$ with noise level $σ^2>0$, and any width-$m$ two-layer network with arbitrary real weights, biases and affine skip connection, fitting the data $\varepsilon$ below the noise floor forces $\mathrm{Lip}(f)\ge c\,\varepsilon\sqrt{n/(\bar m\log(C\bar m nd/\varepsilon))}$, $\bar m=(K-1)m+1$, with high probability. A realized-kink-count version holds on the same event: every realized two-layer piecewise-linear function with $k(f)\le n$ distinct kink hyperplanes obeys the bound with $\bar m$ replaced by $k(f)+1$, irrespective of how many redundant hidden units parameterize it. The proof replaces parameter-space covering, impossible for unbounded weights, by a function-space covering. The central deterministic ingredient is a rigidity lemma: on $B_2$, and on $\mathbb{S}^{d-1}$ for $d\ge3$, the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks on distinct hyperplanes cannot cancel at generic points. Rigidity genuinely fails at $d=2$, and an explicit two-layer ReLU interpolant with $O(1)$ Lipschitz constant at width $2n$ matches the law at the overparameterized endpoint.

cs.LG

Neural means and kernel corrections for operator learning

We combine neural network means with exact Matérn kernel regressions of their residuals and of their learned features, and evaluate the pairing on two public emulation problems with published baselines: the structural-mechanics benchmark of de Hoop et al. and the OCO-2 radiative-transfer emulator of Lamminpää et al. On structural mechanics the combination reaches 4.55% test error, matching the best published architecture, and 5.38% against a published 6.49% in the low-data regime. On OCO-2 it improves on the published Gaussian-process emulator on that problem's own test points, outright on two of the three spectral bands; the same kernel that trails the network tenfold on the raw state overtakes it on the network's features, and we measure why (the target's squared native-space norm drops about fortyfold at fixed effective dimension) and prove the mechanism. Where the two families tie instead, the residuals of every architecture we train correlate above 0.86 and their shared component is flat in diversity and sample size, which reads the published plateau as a property of the data. Supporting results include a second-moment identity that predicts stacking outcomes from measured correlations, an optimal-recovery certificate, and a distribution-free coverage band, the only uncertainty signal that survives our tests.

cs.LG

Extreme least singular values of Gaussian row submatrices and a phase retrieval stability problem

Let $\mathbb F\in\{\mathbb R,\mathbb C\}$ and $d_{\mathbb F}=\dim_{\mathbb R}\mathbb F$. If $A_m\in\mathbb F^{N_m\times m}$ has independent standard Gaussian entries and $N_m/m\toγ>1$, then \[ \min_{\substack{T\subset[N_m]\\ |T|=m}} σ_{\min}(A_{m,T}) = \left(\frac{γ^γ}{(γ-1)^{γ-1}}\right)^{-m/d_{\mathbb F}+o_P(m)} . \] If $N_m=γm+O(1)$, the convergence of $m^{-1}\log M_m^{\mathbb F}$ has probability error $O(m^{-1})$. In particular, at the real phase-retrieval threshold $N=2m-1$, \[ ω(A_m)=4^{-m+o_P(m)}, \] so the Gaussian Balan--Wang critical exponential base is $1/4$.

math.PR

Pruning Deep Neural Networks via the Marchenko--Pastur Distribution

We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets. The main practical contribution is accuracy retention under short calibration and fine-tuning schedules, rather than a long post-pruning reoptimization pipeline. The theory gives deterministic data-path certificates: if the removed component $R$ has small propagated logit effect $L_s \| R ψ_1(s) \|_\infty$, pruning decreases an elastic-net objective and preserves samples whose dense margin exceeds twice the perturbation. The zero-budget case gives perfect pruning; a prune--restore extension models weight restoration inside a fixed sparse-execution pattern; and an additive $L_2$-regularized model shows admissible random-like components vanish at the training limit, with persistent spikes stabilizing as the MP bulk collapses. Under iid-Gaussian sufficient conditions, the fitted MP edge $σ_+$ gives a high-probability layerwise budget signal. On ImageNet-1k, after only three distillation epochs, ViT-B/16 $2{:}4{+}$ToMe reaches $83.41\%$ top-1 ($-1.70$ pp from dense) at $59.81\%$ sparse-execution MAC reduction, with $1.388\times$ best-observed A40 native-$2{:}4$ backend speedup for the same checkpoint and ToMe graph; a separate no-ToMe A100 endpoint gives $2.705\times$. At structured sparsity, ViT-B/16 $6{:}12$ reaches $83.74\%$, ViT-L/16 $8{:}16$ dense+permutation reaches $85.33\%$ ($-0.51$ pp), and ConvNeXtV2-Base $12{:}16$ reaches $86.35\%$ ($-0.37$ pp). For CNNs, ResNet50 $8{:}16$ dense+permutation reaches $75.87\%$ ($-0.26$ pp), and ResNet152d CAST-conv+permutation reaches $81.33\%$ ($-1.53$ pp) at ${\sim}50\%$ MAC accounting with a $1.62\times$ A40 im2col$+2{:}4$ sparse-GEMM audit.

cs.LG

Pruning Deep Neural Networks via a Combination of the Marchenko-Pastur Distribution and Regularization

Deep neural networks (DNNs) have brought significant advancements in various applications in recent years, such as image recognition, speech recognition, and natural language processing. In particular, Vision Transformers (ViTs) have emerged as a powerful class of models in the field of deep learning for image classification. In this work, we propose a novel Random Matrix Theory (RMT)-based method for pruning pre-trained DNNs, based on the sparsification of weights and singular vectors, and apply it to ViTs. RMT provides a robust framework to analyze the statistical properties of large matrices, which has been shown to be crucial for understanding and optimizing the performance of DNNs. We demonstrate that our RMT-based pruning can be used to reduce the number of parameters of ViT models (trained on ImageNet) by 30-50\% with less than 1\% loss in accuracy. To our knowledge, this represents the state-of-the-art in pruning for these ViT models. Furthermore, we provide a rigorous mathematical underpinning of the above numerical studies, namely we proved a theorem for fully connected DNNs, and other more general DNN structures, describing how the randomness in the weight matrices of a DNN decreases as the weights approach a local or global minimum (during training). We verify this theorem through numerical experiments on fully connected DNNs, providing empirical support for our theoretical findings. Moreover, we prove a theorem that describes how DNN loss decreases as we remove randomness in the weight layers, and show a monotone dependence of the decrease in loss with the amount of randomness that we remove. Our results also provide significant RMT-based insights into the role of regularization during training and pruning.

cs.LG

Enhancing Accuracy in Deep Learning Using Random Matrix Theory

We explore the applications of random matrix theory (RMT) in the training of deep neural networks (DNNs), focusing on layer pruning that is reducing the number of DNN parameters (weights). Our numerical results show that this pruning leads to a drastic reduction of parameters while not reducing the accuracy of DNNs and CNNs. Moreover, pruning the fully connected DNNs actually increases the accuracy and decreases the variance for random initializations. Our numerics indicate that this enhancement in accuracy is due to the simplification of the loss landscape. We next provide rigorous mathematical underpinning of these numerical results by proving the RMT-based Pruning Theorem. Our results offer valuable insights into the practical application of RMT for the creation of more efficient and accurate deep-learning models.

cs.LG

Stability of Accuracy for the Training of DNNs Via the Uniform Doubling Condition

We study the stability of accuracy during the training of deep neural networks (DNNs). In this context, the training of a DNN is performed via the minimization of a cross-entropy loss function, and the performance metric is accuracy (the proportion of objects that are classified correctly). While training results in a decrease of loss, the accuracy does not necessarily increase during the process and may sometimes even decrease. The goal of achieving stability of accuracy is to ensure that if accuracy is high at some initial time, it remains high throughout training. A recent result by Berlyand, Jabin, and Safsten introduces a doubling condition on the training data, which ensures the stability of accuracy during training for DNNs using the absolute value activation function. For training data in $\mathbb{R}^n$, this doubling condition is formulated using slabs in $\mathbb{R}^n$ and depends on the choice of the slabs. The goal of this paper is twofold. First, to make the doubling condition uniform, that is, independent of the choice of slabs. This leads to sufficient conditions for stability in terms of training data only. In other words, for a training set $T$ that satisfies the uniform doubling condition, there exists a family of DNNs such that a DNN from this family with high accuracy on the training set at some training time $t_0$ will have high accuracy for all time $t>t_0$. Moreover, establishing uniformity is necessary for the numerical implementation of the doubling condition. The second goal is to extend the original stability results from the absolute value activation function to a broader class of piecewise linear activation functions with finitely many critical points, such as the popular Leaky ReLU.

cs.LG

Deep Learning Weight Pruning with RMT-SVD: Increasing Accuracy and Reducing Overfitting

In this work, we present some applications of random matrix theory for the training of deep neural networks. Recently, random matrix theory (RMT) has been applied to the overfitting problem in deep learning. Specifically, it has been shown that the spectrum of the weight layers of a deep neural network (DNN) can be studied and understood using techniques from RMT. In this work, these RMT techniques will be used to determine which and how many singular values should be removed from the weight layers of a DNN during training, via singular value decomposition (SVD), so as to reduce overfitting and increase accuracy. We show the results on a simple DNN model trained on MNIST. In general, these techniques may be applied to any fully connected layer of a pretrained DNN to reduce the number of parameters in the layer while preserving and sometimes increasing the accuracy of the DNN.

cs.LG

Algorithm for computing Representations of the Braid Group and Temperley-Lieb algebra

The braid group appears in many scientific fields and its representations are instrumental in understanding topological quantum algorithms, topological entropy, classification of manifolds and so on. In this work, we study planer diagrams which are Kauffman's reduction of the braid group algebra to the Temperley-Lieb algebra. We introduce an algorithm for computing all planer diagrams in a given dimension. The algorithm can also be used to multiply planer diagrams and find their matrix representation.

math.GM

Combinatorial Proof of Kakutani's Fixed Point Theorem

Kakutani's fixed point theorem is a generalization of Brouwer's fixed point theorem to upper semicontinuous multivalued maps and is used extensively in game theory and other areas of economics. Earlier works have shown that Sperner's lemma implies Brouwer's theorem. In this paper, a new combinatorial labeling lemma, generalizing Sperner's original lemma, is given and is used to derive a simple proof for Kakutani's fixed point theorem. The proof is constructive and can be easily applied to numerically approximate the location of fixed points. The main method of the proof is also used to obtain a generalization of Kakutani's theorem for discontinuous maps which are locally gross direction preserving.

math.DS

The K-Theoretic Bulk-Boundary Principle for Dynamically Patterned Resonators

Starting from a dynamical system $(Ω,G)$, with $G$ a generic topological group, we devise algorithms that generate families of patterns in the Euclidean space, which densely embed $G$ and on which $G$ acts continuously by rigid shifts. We refer to such patterns as being dynamically generated. For $G=\mathbb Z^d$, we adopt Bellissard's $C^\ast$-algebraic formalism to analyze the dynamics of coupled resonators arranged in dynamically generated point patterns. We then use the standard connecting maps of $K$-theory to derive precise conditions that assure the existence of topological boundary modes when a sample is halved. We supply four examples for which the calculations can be carried explicitly. The predictions are supported by many numerical experiments.

math-ph

A Proof of Atanassov's Conjecture and Other Generalizations of Sperner's Lemma

A simple proof of Atanassov's Conjecture is presented. Atanassov's Conjecture is a generalization of Sperner's Lemma, a lemma which has been used to prove Brouwer's Fixed Point Theorem, among other fixed point theorems. The proof of Atanassov's Conjecture is based on the Brouwer Degree of maps and is extremely elementary. It is much simpler than the original proofs given for the conjecture and provides some insight into the nature of the conjecture. Furthermore, a generalization of the conjecture is presented and finally a new theorem, similar to the original Sperner Lemma, is proved.

math.CO

Combinatorial approach to detection of fixed points, periodic orbits, and symbolic dynamics

We present a combinatorial approach to rigorously show the existence of fixed points, periodic orbits, and symbolic dynamics in discrete-time dynamical systems, as well as to find numerical approximations of such objects. Our approach relies on the method of `correctly aligned windows'. We subdivide the `windows' into cubical complexes, and we assign to the vertices of the cubes labels determined by the dynamics. In this way we encode the dynamics information into a combinatorial structure. We use a version of the Sperner Lemma saying that if the labeling satisfies certain conditions, then there exist fixed points/periodic orbits/orbits with prescribed itineraries. Our arguments are elementary.

math.DS