SearcharxivSearch

arXiv subjects

Bernd Rosenow

Publications and source records attributed to Bernd Rosenow.

At least 19 recordsLinked to original sources

Schrödinger's real-valued equation revisited

The Schrödinger equation can be rewritten as a second-order equation for a single real-valued scalar field. Taken at face value, however, this one-field reduction obscures several local structures of quantum mechanics: the Born density and current, the role of momentum as a generator, and the usual derivation of the uncertainty relation. We argue that these difficulties do not show that real variables fail. They show that the reduced equation is not, by itself, a complete local representation of the theory. Its coherent real form is obtained by restoring the underlying Hamiltonian phase-space structure, where the missing component reappears as a canonical partner. In this real phase-space formulation, the standard structures reappear locally, and minimal magnetic coupling reveals an internal \(SO(2)\) gauge structure. Thus, the symbol \(i\) can be removed, but the symplectic and complex structure it encodes cannot.

quant-ph

The two-particle density matrix of a Luttinger liquid

Two-particle coherence is the first level of the reduced-density-matrix hierarchy that contains correlations inaccessible to single-particle observables, yet analytic two-particle density matrices are rare even in one dimension. We derive a closed, finite-size expression for the equal-time two-particle reduced density matrix of spinless fermions in a Tomonaga-Luttinger liquid using constructive bosonization with an explicit ultraviolet cutoff. In addition to the familiar Luttinger parameter $K$-dependent exponent $γ^2=(K+K^{-1}-2)/2$ which governs the spatial decay of matrix elements, the result exposes a second exponent, $λ=(K^{-1}-K)/2$, which encodes correlations between opposite chiralities and controls the off-diagonal structure. The diagonal limit of the two-particle reduced density matrix yields density correlations and the static structure factor, while its coherences resolve algebraic $2k_F$ charge-density-wave correlations for repulsion and odd-parity p-wave pairing correlations for attraction. After fixing the cutoff from the one-particle density matrix, the analytic result quantitatively reproduces density matrix renormalization group calculations of the interacting J-V chain within the Luttinger liquid regime. The result connects universal Luttinger liquid scaling with observables in finite microscopic systems.

cond-mat.str-el

Anyon Exchange Phase from Antidot Interferometry

Quasiparticles in fractional quantum Hall systems are anyons, carrying a fraction of the electron charge. Exchanging two of them gives rise to a fractional exchange phase. While the fractional charge and the braiding phase -- twice the exchange phase -- have been measured, the exchange phase itself has remained inaccessible. We study a quantum antidot embedded in a Fabry-Perot interferometer. Within a systematic non-equilibrium Keldysh treatment that consistently includes the occupation and level broadening of the antidot, we find that the transmission phase evolves non-monotonically when a gate voltage tunes the antidot through a resonance, in contrast to the monotonic evolution for electrons. The bare exchange phase can be extracted from the difference between the phase plateaus.

cond-mat.mes-hall

Edge Reconstruction in a Quantum Spin Hall Insulator

We study interaction-driven edge reconstruction in a quantum spin Hall insulator described by the Bernevig-Hughes-Zhang model with Kanamori-Hubbard interactions using the real-space density matrix renormalization group method in both the grand-canonical and canonical ensembles. For a two-dimensional cylinder with a smooth edge, we identify discrete particle-number transitions that lead to a spin-polarized edge state stabilized by an emergent ferromagnetic exchange interaction. The reconstruction is orbital-selective, occurring predominantly in the $s$-orbital channel. Our results reveal a microscopic mechanism for emergent fluctuating moments at the edge that could compromise the topological protection of helical edge states by time reversal symmetry.

cond-mat.mes-hall

Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks

The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive. We argue that the bulk consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In deep linear networks these symmetries are exact: they generate flat directions and hence exact zero modes, whose eigenvectors we construct explicitly. Introducing a ReLU nonlinearity as a perturbation, we show that it breaks these symmetries weakly and explicitly. Resolving the spectrum at the level of eigenvectors, we find that the high-curvature directions are orthogonal to the symmetry subspace, while the bulk lies almost entirely within it. We demonstrate the mechanism in a two-layer ReLU student--teacher model and in a network trained on CIFAR-10. A convolutional example demonstrates that the same diagnostic extends beyond fully connected layers. Together, these results link the Hessian bulk to weakly broken symmetries and clarify the origin of near-zero modes.

cs.LG

BBP Phase Transition for an Extensive Number of Outliers

Random-matrix theory helps disentangle signal from noise in large data sets. We analyze rectangular $p \times q$ matrices $W = W_0 + M$ in which the noise $M$ generates a Marchenko-Pastur bulk, whereas the signal $W_0$ injects an extensive set of degenerate singular values. Keeping $\mathrm{rank}$ $W_0/q$ finite as $p,q \to \infty$, we show that the trace of the resolvent of $W^{\top} W$ obeys a quartic equation for one degenerate signal, yielding an exact spectral density, and derive explicit asymptotics in the strong-signal regime. We map out a detailed generalized Baik-Ben Arous-Péché (BBP) phase diagram and clarify how a finite density of spikes reshapes the bulk edges. We further derive a $1/3$-scaling law for the critical signal strength in terms of the rank ratio for rectangular matrices in the finite-to-extensive-rank crossover. Numerical simulations validate the theory and illustrate its relevance for high-dimensional inference tasks with multiple degenerate signals and more general signal distributions.

cond-mat.dis-nn

Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model

Chain-of-thought generation can turn a multi-step computation into a sequence of locally checkable state updates, but the training dynamics by which transformers acquire such updates remain poorly understood. We study this question in a solvable setting: a simplified one-block transformer trained by supervised next-token prediction on state sequences generated by composing permutations. The architecture separates fixed-lag action retrieval, learned by RoPE attention, from a specialized MLP logic module that applies the retrieved permutation to the current state. Using a statistical-physics mean-field description, we derive dynamics for three order parameters measuring attention retrieval, teacher-matrix alignment, and off-target logic overlap. These equations quantitatively match simulations for the order parameters and, combined with a logit-distribution approximation, qualitatively predict the sharp transition in final rollout accuracy. The analysis reveals staged learning: the logic module first learns a mixed heuristic; attention then locks onto the relevant action, enabling efficient MLP alignment. Together, these results provide a controlled mechanistic account of how attention-based retrieval and MLP-based logic co-develop during chain-of-thought state tracking.

cond-mat.dis-nn

A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification

Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between smooth surrogate and discrete labels produces power-law learning curves in an online teacher-student model. After subtracting the mean logit, the thermodynamic-limit dynamics close in centered variables: a growing centered student-teacher alignment $D$ and the residual student variance $Δ$. At late times, examples away from teacher decision boundaries are already classified confidently and contribute exponentially little. Only boundary layers of width $O(D^{-1})$ remain active, while the noise of fixed-learning-rate online gradient descent maintains a nonzero $Δ$. As a function of the training time $α$ the late-time solution yields a $α^{-1/3}$ power law not only for the test loss but also for the generalization error $ε_g$, i.e., one minus test accuracy. This is much slower than the $α^{-1}$ Bayes-optimal reference for the same model. We further show that learning-rate schedules can improve the generalization error towards a $ε_g \sim α^{-1/2}$ power law. Simulations support the predicted order parameter dynamics and learning curves. Controlled experiments with correlated Gaussian inputs and whitened pretrained features show that data structure can dominate transients. Therefore, our result is an asymptotic, complementary mechanism rather than an alternative to spectral explanations of neural scaling laws.

cs.LG

Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation

We analyze the soft committee machine with Rectified Linear Unit (ReLU) activation by means of the replica method. In a realizable teacher--student setting, we compute the quenched free energy within a replica-symmetric ansatz and obtain the typical generalization behavior from the saddle-point equations for the macroscopic order parameters. The system exhibits a transition from an unspecialized symmetric phase to a specialized phase in which the permutation symmetry among hidden units is broken. We determine the critical training-set size as a function of the inverse training temperature and derive analytic expressions both near the transition and in the asymptotic large-sample regime. Unlike the corresponding model with sigmoidal activations, which undergoes a first-order transition, the ReLU soft committee machine shows a continuous specialization transition. These results show that the activation function plays a decisive role in the phase structure and generalization behavior of multilayer networks.

cond-mat.dis-nn

Extracting the Anyonic Exchange Phase from Hanbury Brown-Twiss Correlations

In recent years, interferometry experiments in fractional quantum Hall devices have reported signatures of a fractional braiding phase for quasiparticles. It was noted, however, that the braiding phase alone does not uniquely determine the exchange phase because of a $π$-ambiguity. Here we analyze a Hanbury Brown-Twiss interferometer in a cross geometry that provides direct access to the fractional exchange phase. Using a non-equilibrium Keldysh calculation in an experimentally relevant regime, we show that the exchange phase can be obtained as the phase shift between Aharonov-Bohm oscillations in a single-particle interference current and those in the current cross-correlation arising from two-particle interference.

cond-mat.mes-hall

Unified Description of Learning Dynamics in the Soft Committee Machine from Finite to Ultra-Wide Regimes

We study the learning dynamics of the soft committee machine (SCM) with Rectified Linear Unit (ReLU) activation using a statistical-mechanics approach within the annealed approximation. The SCM consists of a student network with $N$ input units and $K$ hidden units trained to reproduce the output of a teacher network with $M$ hidden units. We introduce a reduced set of macroscopic order parameters that yields a unified description valid from the conventional regime $K \ll N$ to the ultra-wide limit $K \ge N$. The control parameter $α$, proportional to the ratio of training samples to adjustable weights, serves as an effective measure of dataset size. For small $γ= M/N$, we recover a continuous phase transition at $α_{c} \approx 2π$ from an unspecialized, permutation-symmetric state to a specialized state in which student units align with the teacher. For finite $γ$, the transition disappears and the generalization error decreases smoothly with dataset size, reaching a low plateau when $γ=1$. In the asymptotic limit $α\to \infty$, the error scales as $\varepsilon_{g} \propto 1/α$, independent of $γ$ and $K$. The results highlight the central role of network dimensions in SCM learning and provide a framework extendable to other activations and quenched analyses.

cond-mat.dis-nn

Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions

Stochastic Gradient Descent (SGD) has become a cornerstone of neural network optimization due to its computational efficiency and generalization capabilities. However, the gradient noise introduced by SGD is often assumed to be uncorrelated over time, despite the common practice of epoch-based training where data is sampled without replacement. In this work, we challenge this assumption and investigate the effects of epoch-based noise correlations on the stationary distribution of discrete-time SGD with momentum. Our main contributions are twofold: First, we calculate the exact autocorrelation of the noise during epoch-based training under the assumption that the noise is independent of small fluctuations in the weight vector, revealing that SGD noise is inherently anti-correlated over time. Second, we explore the influence of these anti-correlations on the variance of weight fluctuations. We find that for directions with curvature of the loss greater than a hyperparameter-dependent crossover value, the conventional predictions of isotropic weight variance under stationarity, based on uncorrelated and curvature-proportional noise, are recovered. Anti-correlations have negligible effect here. However, for relatively flat directions, the weight variance is significantly reduced, leading to a considerable decrease in loss fluctuations compared to the constant weight variance assumption. Furthermore, we present a numerical experiment where training with these anti-correlations enhances test performance, suggesting that the inherent noise structure induced by epoch-based training may play a role in finding flatter minima that generalize better.

cs.LG

Small Singular Values Matter: A Random Matrix Analysis of Transformer Models

This work analyzes singular-value spectra of weight matrices in pretrained transformer models to understand how information is stored at both ends of the spectrum. Using Random Matrix Theory (RMT) as a zero information hypothesis, we associate agreement with RMT as evidence of randomness and deviations as evidence for learning. Surprisingly, we observe pronounced departures from RMT not only among the largest singular values -- the usual outliers -- but also among the smallest ones. A comparison of the associated singular vectors with the eigenvectors of the activation covariance matrices shows that there is considerable overlap wherever RMT is violated. Thus, significant directions in the data are captured by small singular values and their vectors as well as by the large ones. We confirm this empirically: zeroing out the singular values that deviate from RMT raises language-model perplexity far more than removing values from the bulk, and after fine-tuning the smallest decile can be the third most influential part of the spectrum. To explain how vectors linked to small singular values can carry more information than those linked to larger values, we propose a linear random-matrix model. Our findings highlight the overlooked importance of the low end of the spectrum and provide theoretical and practical guidance for SVD-based pruning and compression of large language models.

cs.LG

Braids and Beams: Exploring Fractional Statistics with Mesoscopic Anyon Colliders

Anyon colliders -- quantum Hall devices where dilute quasiparticle beams collide at a quantum point contact -- provide an interferometer-free probe of anyonic exchange phases through current cross correlations. Within a non-equilibrium bosonization framework, the normalized cross-correlations take a universal form depending only on the exchange phase and the dynamical exponent, enabling experimental demonstration of anyonic statistics. This result can be interpreted as time-domain interference -- braiding in time rather than spatial exclusion or real-space interferometry. Extension to hierarchical states shows that the semiclassical step-function description of quasiparticles fails at large statistical angles. Introducing a finite soliton width resolves this issue and enables quantitative modeling of charge-$e/5$ quasiparticle collisions.

cond-mat.mes-hall

Berezinskii-Kosterlitz-Thouless Renormalization Group Flow at a Quantum Phase Transition

We present a controlled numerical study of the Berezinskii-Kosterlitz-Thouless (BKT) transition in the one-dimensional Bose-Hubbard model at unit filling, providing evidence of the characteristic logarithmic finite-size scaling of the BKT transition. Employing density matrix renormalization group and quantum Monte Carlo simulations under periodic boundary conditions, together with a systematic finite-size scaling analysis of bipartite particle number fluctuations, we resolve boundary-induced complications that previously obscured critical scaling. We demonstrate that a suitably chosen central region under open boundaries reproduces universal RG signatures, reconciling earlier discrepancies. Finally, leveraging a non-parametric Bayesian analysis, we determine the critical interaction strength with high precision, establishing a benchmark for BKT physics in one-dimensional quantum models.

cond-mat.quant-gas

Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning

Large annotated datasets inevitably contain noisy labels, which poses a major challenge for training deep neural networks as they easily memorize the labels. Noise-robust loss functions have emerged as a notable strategy to counteract this issue, but it remains challenging to create a robust loss function which is not susceptible to underfitting. Through a quantitative approach, this paper explores the limited overlap between the network output at initialization and regions of non-vanishing gradients of bounded loss functions in the initial learning phase. Using these insights, we address underfitting of several noise robust losses with a novel method denoted as logit bias, which adds a real number $ε$ to the logit at the position of the correct class. The logit bias enables these losses to achieve state-of-the-art results, even on datasets like WebVision, consisting of over a million images from 1000 classes. In addition, we demonstrate that our method can be used to determine optimal parameters for several loss functions -- without having to train networks. Remarkably, our method determines the hyperparameters based on the number of classes, resulting in loss functions which require zero dataset or noise-dependent parameters.

cs.LG

Friedel oscillations in one-dimensional 4He

One-dimensional bosonic systems, such as helium confined to nanopores, exhibit Luttinger liquid behavior characterized by density waves as collective excitations. We investigate the impact of a scattering potential on a low dimensional quantum liquid. We consider a microscopic model of $^4$He inside a perturbed nanopore with a localized constriction, and employ quantum Monte Carlo simulations to analyze the density of the core within an effective low-energy framework. Our results reveal the emergence of Friedel oscillations in a bosonic quantum liquid without a Fermi surface. Furthermore, we utilize the Luttinger liquid model to predict experimentally observable signatures of this pinning phenomena in elastic scattering and via the temperature and pressure dependence of mass transport through the deformed nanopore.

cond-mat.mes-hall

Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data Spectra

Neural scaling laws describe how the performance of deep neural networks scales with key factors such as training data size, model complexity, and training time, often following power-law behaviors over multiple orders of magnitude. Despite their empirical observation, the theoretical understanding of these scaling laws remains limited. In this work, we employ techniques from statistical mechanics to analyze one-pass stochastic gradient descent within a student-teacher framework, where both the student and teacher are two-layer neural networks. Our study primarily focuses on the generalization error and its behavior in response to data covariance matrices that exhibit power-law spectra. For linear activation functions, we derive analytical expressions for the generalization error, exploring different learning regimes and identifying conditions under which power-law scaling emerges. Additionally, we extend our analysis to non-linear activation functions in the feature learning regime, investigating how power-law spectra in the data covariance matrix impact learning dynamics. Importantly, we find that the length of the symmetric plateau depends on the number of distinct eigenvalues of the data covariance matrix and the number of hidden units, demonstrating how these plateaus behave under various configurations. In addition, our results reveal a transition from exponential to power-law convergence in the specialized phase when the data covariance matrix possesses a power-law spectrum. This work contributes to the theoretical understanding of neural scaling laws and provides insights into optimizing learning performance in practical scenarios involving complex data structures.

stat.ML