SearcharxivSearch

arXiv subjects

Alon Beck

Publications and source records attributed to Alon Beck.

5 recordsLinked to original sources

The Implicit Bias of Logit Regularization

Logit regularization, the addition of a convex penalty directly in logit space, is widely used in modern classifiers, with label smoothing as a prominent example. While such methods often improve calibration and generalization, their mechanism remains under-explored. In this work, we analyze a general class of such logit regularizers in the context of linear classification, and demonstrate that they induce an implicit bias of logit clustering around finite per-sample targets. For Gaussian data, or whenever logits are sufficiently clustered, we prove that logit clustering drives the weight vector to align exactly with Fisher's Linear Discriminant. To demonstrate the consequences, we study a simple signal-plus-noise model in which this transition has dramatic effects: Logit regularization halves the critical sample complexity and induces grokking in the small-noise limit, while making generalization robust to noise. Our results extend the theoretical understanding of label smoothing and highlight the efficacy of a broader class of logit-regularization methods.

stat.ML

Wave-packet dynamics in pseudo-Hermitian lattices: Coexistence of Hermitian and non-Hermitian wavefronts

This paper investigates wave-packet dynamics in non-Hermitian lattice systems and reveals a surprising phenomenon: The simultaneous propagation of two distinct wavefronts, one traveling at the non-Hermitian velocity and the other at the Hermitian velocity. We show that this dual-front behavior arises naturally in systems governed by a pseudo-Hermitian Hamiltonian. Using the paradigmatic Hatano-Nelson model as our primary example, we demonstrate that this coexistence is essential for understanding a wide array of unconventional dynamical effects, including abrupt ``non-Hermitian reflections'', sudden shifts of Gaussian wave-packets, and disorder-induced emergent packets seeded by the small initial tails. We present analytic predictions that closely match numerical simulations. These results may offer new insight into the topology of non-Hermitian systems and point toward measurable experimental consequences.

quant-ph

Grokking at the Edge of Linear Separability

We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "memorizing" and "generalizing" solutions can be strictly defined. Surprisingly, we find that grokking arises naturally even in this minimal model when the parameters of the problem are close to a critical point, and provide both empirical and analytical insights into its mechanism. Concretely, by appealing to the implicit bias of gradient descent, we show that logistic regression can exhibit grokking when the training dataset is nearly linearly separable from the origin and there is strong noise in the perpendicular directions. The underlying reason is that near the critical point, "flat" directions in the loss landscape with nearly zero gradient cause training dynamics to linger for arbitrarily long times near quasi-stable solutions before eventually reaching the global minimum. Finally, we highlight similarities between our findings and the recent literature, strengthening the conjecture that grokking generally occurs in proximity to the interpolation threshold, reminiscent of critical phenomena often observed in physical systems.

stat.ML

Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding

Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a simple teacher-student setup with Gaussian inputs. In this setting, the full training dynamics is derived in terms of the training and generalization data covariance matrix. We present exact predictions on how the grokking time depends on input and output dimensionality, train sample size, regularization, and network initialization. We demonstrate that the sharp increase in generalization accuracy may not imply a transition from "memorization" to "understanding", but can simply be an artifact of the accuracy measure. We provide empirical verification for our calculations, along with preliminary results indicating that some predictions also hold for deeper networks, with non-linear activations.

stat.ML

Disorder in dissipation-induced topological states: Evidence for a different type of localization transition

The quest for nonequilibrium quantum phase transitions is often hampered by the tendency of driving and dissipation to give rise to an effective temperature, resulting in classical behavior. Could this be different when the dissipation is engineered to drive the system into a nontrivial quantum coherent steady state? In this work we shed light on this issue by studying the effect of disorder on recently-introduced dissipation-induced Chern topological states, and examining the eigenmodes of the Hermitian steady state density matrix or entanglement Hamiltonian. We find that, similarly to equilibrium, each Landau band has a single delocalized level near its center. However, using three different finite size scaling methods we show that the critical exponent $\nu$ describing the divergence of the localization length upon approaching the delocalized state is significantly different from equilibrium if disorder is introduced into the non-dissipative part of the dynamics. This indicates a different type of nonequilibrium quantum critical universality class accessible in cold-atom experiments.

cond-mat.dis-nn