SearcharxivSearch

arXiv subjects

John Sous

Publications and source records attributed to John Sous.

At least 19 recordsLinked to original sources

Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.

cs.LG

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime $n\asymp d$, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.

stat.ML

Self-consistent GW theory for superconductivity in SrTiO3 models

Superconductivity in doped SrTiO$_3$ occurs over a wide range of carrier densities, including those for which the Fermi energy is below the polar longitudinal optical phonon scale. In this regime, the assumptions underpinning conventional implementations of Migdal-Eliashberg theory, including frequency cutoffs at the phonon scale and a Coulomb pseudopotential $\mu^\ast$, are not valid. We solve the finite-temperature $GW$ equations with full momentum and frequency dependence, without cutoffs or $\mu^\ast$, for polar one-band models of SrTiO$_3$, using effective masses and three-phonon dielectric functions parameterized from ab initio calculations. Comparing different self-consistency levels, namely $G_0W_0$, $GW_0$, and fully self-consistent $GW$, we find that the one-shot ($G_0W_0$) kernel overestimates the pairing-onset temperature by one to two orders of magnitude. The dominant suppression comes from replacing $G_0$ by $G$, thereby incorporating the phonon renormalization factor in the electron Green function. Using the self-consistently computed interaction $W$ further lowers and narrows the pairing-onset dome. In the dilute limit, our calculations identify the pairing channel as the Fr\"ohlich phonon interaction screened by the incipient ferroelectricity of the material, with plasmonic and electronic screening effects negligible. The numerical solution of the full equations reveals a pairing-onset scale that remains non-zero as the density tends to zero, whereas Fermi-surface projection or Fermi-energy frequency truncation removes it. This work highlights the relevance of incipient ferroelectricity, the importance of self-consistency, and the need for a full momentum- and frequency-dependent treatment in modeling superconductivity in SrTiO$_3$-like doped polar semiconductors.

cond-mat.supr-con

Assign and Add: A Mechanistic Study of Compositional Arithmetic

Large language models are able to compose skills in order to perform complex tasks, many of which might not have been seen during training. The details of how exactly this composition occurs remain elusive. In this paper, we study a mechanism for compositional generalization in transformers by considering a simple controlled setting involving variable assignment and modular addition. By partitioning our training data into disjoint sets, we observe that small transformers are able to generalize to previously unseen combinations of variables and numbers. Our mechanistic analysis shows that the same ``modular addition'' MLP module is used whether the inputs are given directly or indirectly through a separate variable assignment mechanism. We also analyze the training dynamics from an empirical lens, which reveals three phases of learning: first, modular addition is learned, then the structure required for variable assignment, and finally a refinement phase where the model generalizes to some hard sequences not seen in training. Finally, we provide a theoretical framework to explain how compositionality emerges from training dynamics. These results suggest that compositional generalization can be a natural consequence of the compositionality of internal mechanisms in~transformers.

cs.LG

Asymmetric Scaling Laws from Sparse Features

We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold -- where the number of parameters is just sufficient to fit the training data -- resulting in a loss curve governed by two distinct scaling exponents -- one for the overparameterized regime and one for the underparameterized regime -- with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations.

stat.ML

Bipolaronic High-Temperature Superconductivity from Phonon-Modulated Hopping: A Perspective

Phonon-mediated superconductivity is conventionally thought to be capped at a transition temperature $T_{\mathrm{c}}$ no larger than roughly one-tenth of the phonon frequency $\Omega$, a bound rooted in the breakdown of Migdal-Eliashberg theory at intermediate coupling and in the heaviness of bipolarons formed in standard models with phonons that couple to the electron density. In this review I describe a route to phonon-mediated high-$T_{\mathrm{c}}$ superconductivity that bypasses this bound. The key ingredient is a class of electron-phonon couplings in which lattice distortions modulate the electron hopping and therefore its kinetic energy rather than its potential energy, known as the Peierls model (also known as Su-Schrieffer-Heeger model). In these models phonon exchange generates an interaction that binds two electrons into a small but unusually light bipolaron. Using sign-problem-free quantum Monte Carlo simulations of a bond-Peierls model on the square and cubic lattices, my collaborators and I have shown that a dilute liquid of such bipolarons forms an $s$-wave superconductor with a $T_{\mathrm{c}}/\Omega$ that significantly exceeds the conventional bound, that this conclusion is robust against screened Coulomb repulsion, and that $T_{\mathrm{c}}/\Omega$ -- despite being reduced -- remains above bound in presence of strong long-range Coulomb repulsion. A semi-classical instanton analysis explains why, at strong coupling, bipolarons in models with phonon-modulated hopping are lighter than their density-coupled (Holstein) counterparts. I close with a discussion of materials in which this physics may be operative, in particular the iron-based pnictide superconductors, and of design principles that follow from it.

cond-mat.supr-con

Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization

Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and operational diagnostics. Using a controlled family of decoder-only models adapted from the modded NanoGPT codebase, we introduce an empirical protocol based on activation covariance and per-sample gradient SVD spectra. This dual-view reveals three empirical findings and one mechanistic explanation. First, batch size acts as a latent determinant of representation geometry: runs that reach equal loss settle into systematically distinct activation spectra. Second, the activation covariance tail measured early in training reliably forecasts downstream token efficiency. Third, movement of the activation spectrum head (leading modes), together with gradient spectra, characterizes underlying learning-dynamics changes, separating learning-side architectural improvements from primarily execution-side gains. These predictive and diagnostic signals persist across the 12-, 36-, and 48-layer model tiers. Finally, a mechanistic model proves the main observations and explains how activation covariance spectra correlate with task-aligned feature learning.

stat.ML

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong control, but caches guidance tokens at every layer and can clutter long interactions; activation steering is compact but typically weaker and does not support large structured reminders. We introduce memory inception (MI), a training-free method that steers in latent attention space by inserting text-derived key-value (KV) banks only at selected layers. Rather than materializing reminder content throughout the prompt cache, MI treats steering as selective KV allocation, injecting latent slots only where the model routes to them. On matched personality-steering tasks, MI gives the best overall control--drift trade-off, remaining competitive with prompting while consistently outperforming CAA. On updateable guidance, MI supports mid-conversation behavior shifts without rewriting the visible transcript, achieving the highest post-shift alignment on Qwen3. On structured reasoning, MI outperforms visible prompting on HARDMath and PHYSICS (10/12 subject$\times$mode cells), serving as proxies for structured reasoning in verifiable domains, while cutting content-matched KV storage by up to 118$\times$. These results position MI as a powerful steering method when guidance is persistent, structured, or expensive to keep in the visible transcript.

cs.LG

Ultrafast electronic coherence from slow phonons

Light offers a route to engineer new phases of matter far from equilibrium, including transient states suggestive of superconducting, charge-ordered, and excitonic ordering behavior. Yet it remains unclear how optical excitation can dynamically produce long-range phase coherence-a defining feature of true order such as superconductivity-rather than merely enhancing local pairing. Here we show that impulsively driven low-frequency phonons enhance long-range electronic correlations in a low-dimensional metal. Through numerically exact simulations, we demonstrate that slow phonons suppress dynamical disorder, enabling buildup of coherence and enhancement of charge (and pairing) orders. These findings provide direct evidence that light can mediate enhancement of long-range order and suggest that future experimental strategies-such as the design of selective excitations of narrow phonon distributions to limit dephasing-may offer viable routes to design and stabilize transient superconducting states.

cond-mat.supr-con

(Im)possibility of Automated Hallucination Detection in Large Language Models

Is automated hallucination detection possible? In this work, we introduce a theoretical framework to analyze the feasibility of automatically detecting hallucinations produced by large language models (LLMs). Inspired by the classical Gold-Angluin framework for language identification and its recent adaptation to language generation by Kleinberg and Mullainathan, we investigate whether an algorithm, trained on examples drawn from an unknown target language $K$ (selected from a countable collection) and given access to an LLM, can reliably determine whether the LLM's outputs are correct or constitute hallucinations. First, we establish an equivalence between hallucination detection and the classical task of language identification. We prove that any hallucination detection method can be converted into a language identification method, and conversely, algorithms solving language identification can be adapted for hallucination detection. Given the inherent difficulty of language identification, this implies that hallucination detection is fundamentally impossible for most language collections if the detector is trained using only correct examples from the target language. Second, we show that the use of expert-labeled feedback, i.e., training the detector with both positive examples (correct statements) and negative examples (explicitly labeled incorrect statements), dramatically changes this conclusion. Under this enriched training regime, automated hallucination detection becomes possible for all countable language collections. These results highlight the essential role of expert-labeled examples in training hallucination detectors and provide theoretical support for feedback-based methods, such as reinforcement learning with human feedback (RLHF), which have proven critical for reliable LLM deployment.

cs.LG

PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving

We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. We develop a robust automated evaluation system for precise and reliable validation. Our evaluation of leading foundation models reveals substantial limitations. Even the most advanced model, o3-mini, achieves only 59.9% accuracy, highlighting significant challenges in solving high-level scientific problems. Through comprehensive error analysis, exploration of diverse prompting strategies, and Retrieval-Augmented Generation (RAG)-based knowledge augmentation, we identify key areas for improvement, laying the foundation for future advancements.

cs.AI

Dynamical control in a prethermalized molecular ultracold plasma: Local dissipation drives global relaxation

Prethermalization occurs as an important phase in the dynamics of many-body systems when strong coupling drives a quasi-equilibrium in a subspace separated from the thermodynamic equilibrium by the restriction of a gap in energy or other conserved quantity. Here, we report the signature of an enduring prethermal regime of arrested relaxation in the molecular ultracold plasma that forms following the avalanche of a state-selected Rydberg gas of nitric oxide. Electron collisions mix orbital angular momentum, scattering Rydberg molecules to states of very high-$\ell$. Spontaneous predissociation purifies this non-penetrating character, creating an extraordinary gap between the plasma states of $n \approx \ell$, with measured $n>200$ and penetrating states of $\ell = 0, ~1$ and 2. Evolution to a statistically equilibrated state of N and O atoms cannot occur without Rydberg electron penetration, and this gap blocks relaxation for a millisecond or more. Evolving through the critical phase, electrons that balance the NO$^+$ charge behave as though localized in the prethermal phase and play an ineffective role in bridging this gap. However, the application of a weak radiofrequency (RF) field promotes a dramatic degree of relaxation owing to electron collisions. On an entirely different scale, exciting a quantum-state transition in an exceedingly small fraction of the molecules in the prethermalized ensemble acts with even greater effect to drive the entire system toward equilibrium. We ascribe this to dissipative character added to a small fraction of the states in the prethermally localized ensemble. Using the Lindblad master equation, we illustrate qualitatively similar dynamics for a toy model of an open quantum system that consists of a localized set of spins on which dissipation acts locally at a single site.

cond-mat.quant-gas

Phonon state tomography of electron correlation dynamics in optically excited solids

We introduce phonon state tomography (PST) as a diagnostic probe of electron dynamics in solids whose phonons are optically excited by a laser pulse at initial time. Using a projected-purified matrix-product states algorithm, PST decomposes the exact correlated electron-phonon wavefunction into contributions from purely electronic states corresponding to statistically typical configurations of the optically accessible phononic response, enabling a 'tomographic' reconstruction of the electronic dynamics generated by the phonons. Thus, PST may be used to diagnose electronic behavior in experiments that access only the phonon response, such as thermal diffuse x-ray and electron scattering. We study the dynamics of a metal whose infrared phonons are excited by an optical pulse at initial time and use it to simulate the sample-averaged momentum-resolved phonon occupancy and accurately reconstruct the electronic correlations. We also use PST to analyze the influence of different pulse shapes on the light-induced enhancement and suppression of electronic correlations.

cond-mat.str-el

Ultrafast dynamics of a fermion chain in a terahertz field-driven optical cavity

We study the effect of a terahertz field-driven single cavity mode for ultrafast control of a fermion chain with dissipation-induced nonlinearity and quadratic coupling to an infrared-active phonon mode. Without photon loss from the cavity, we uncover a first-order phase transition in the nonequilibrium steady state only for the lower phonon-polariton, accompanied by polaritons whose frequency response is asymmetric with respect to the photon frequency due to the direct laser-induced dressing effect on the photon. A weak laser field fails to induce the phase transition but renders the polaritons symmetrical. Finally, we show that sufficiently strong photon loss from the cavity eliminates the polaritons and the associated phase transition. The experimental feasibility of these phenomena is also proposed.

cond-mat.str-el

Extrapolation of polaron properties to low phonon frequencies by Bayesian machine learning

Feasibility of accurate quantum calculations is often restricted by the dimensionality of the truncated Hilbert space required for the numerical computations. The present work demonstrates Bayesian machine learning (ML) models that use quantum properties in an effectively lower-dimensional Hilbert space to make predictions for the Hamiltonian parameters that require a larger basis set as applied to a classical problem in quantum statistical mechanics, the polaron problem. We consider two polaron models: the Su-Schrieffer-Heeger (SSH) model and the mixed SSH-Holstein model. We demonstrate ML models that can extrapolate polaron properties in the phonon frequency. We consider the sharp transition in the ground-state momentum of the SSH polaron and examine the evolution of this transition from the anti-adiabatic regime to the adiabatic regime. We also demonstrate Bayesian models that use the posterior distributions of highly approximate quantum calculations as the prior distribution for models of more accurate quantum results. This drastically reduces the number of fully converged quantum calculations required to map out the polaron dispersion relations for the full range of Hamiltonian parameters of interest.

quant-ph

Cooper-Paired Bipolaronic Superconductors

Light-mass bipolarons in off-diagonally coupled electron-phonon systems provide a potential route to bipolaronic high-Tc superconductivity. While there has been numerical progress in the physically relevant limit of slow phonons, more insights are needed to fully understand to what extent this mechanism survives at finite densities and in different regimes. We address these questions using advanced tensor-network methods applied to the Su-Schrieffer-Heeger model. Studying both the spectral properties of isolated bipolarons as well as the pairing correlations at small finite densities, we find evidence that the conventional picture of bipolarons as molecular bound states which undergo Bose-Einstein condensation may need to be reconsidered. Instead, our findings suggest correlation-driven formation of a fragmented condensate with spatially separated polaron pairs, stabilized by strong repulsive electron-electron interactions at moderate values of the electron-phonon coupling. These spatially modulated charge clouds exhibit pairing with a typical length-scale set by the Fermi momentum.

cond-mat.supr-con

Semi-classical theory of bipolaronic superconductivity in a bond-modulated electron-phonon model

We analyze the transition temperature $T_c$ of bipolaronic superconductivity in a bond Su-Schrieffer-Heeger (bond-SSH) model -- also known as a bond Peierls model -- where the electron hoppings are modulated by bond phonons. Using a semiclassical instanton approximation justifiable in the adiabatic limit of slow phonons, we find that the bipolaron mass is only weakly enhanced, in contrast to the typical large mass enhancement found in standard (Holstein) electron-phonon models. Specifically, in the strong coupling limit, the bipolarons can freely slide within a degenerate manifold rather than become self-trapped. A gas of these bipolarons can undergo a superfluid transition at a critical temperature for which we obtain an upper bound. We find that this bound is exponentially larger than that in the Holstein model. Our study provides an analytical understanding of the mechanism behind the high-$T_c$ bipolaronic superconductivity numerically observed in [Phys. Rev. X 13, 011010 (2023)].

cond-mat.supr-con

Weakened Topological Protection of the Quantum Hall Effect in a Cavity

We study the quantum Hall effect in a two-dimensional homogeneous electron gas coupled to a quantum cavity field. As initially pointed out by Kohn, Galilean invariance for a homogeneous quantum Hall system implies that the electronic center of mass (CM) decouples from the electron-electron interaction, and the energy of the CM mode, also known as Kohn mode, is equal to the single particle cyclotron transition. In this work, we point out that strong light-matter hybridization between the Kohn mode and the cavity photons gives rise to collective hybrid modes between the Landau levels and the photons. We provide the exact solution for the collective Landau polaritons and we demonstrate the weakening of topological protection at zero temperature due to the existence of the lower polariton mode which is softer than the Kohn mode. This provides an intrinsic mechanism for the recently observed topological breakdown of the quantum Hall effect in a cavity [Appugliese et al., Science 375, 1030-1034 (2022)]. Importantly, our theory predicts the cavity suppression of the thermal activation gap in the quantum Hall transport. Our work paves the way for future developments in the cavity control of quantum materials.

cond-mat.mes-hall