SearcharxivSearch

arXiv subjects

Daniel J. Korchinski

Publications and source records attributed to Daniel J. Korchinski.

8 recordsLinked to original sources

Symmetry in language statistics shapes the geometry of model representations

The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe. To explain this neural code, we first show that language statistics exhibit translation symmetry (for example, the frequency with which any two months co-occur in text depends only on the time interval between them). We prove that this symmetry governs these geometric structures in high-dimensional word embedding models, and we analytically derive the manifold geometry of word representations. These predictions empirically match large text embedding models and large language models. Moreover, the representational geometry persists at moderate embedding dimension even when the relevant statistics are perturbed (e.g., by removing all sentences in which two months co-occur). We prove that this robustness emerges naturally when the co-occurrence statistics are controlled by an underlying latent variable. Our results indicate that these representational manifolds originate in the statistical symmetries of natural language.

cs.LG

Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?

Generative models have a persistent limitation: their tendency to memorize training data can create legal liabilities and erode creative diversity. Understanding which samples are memorized in whole or in part, and under what conditions, therefore remains an important open problem. Here we answer the question "Are atypical or rare samples memorized first?" in the negative. We train diffusion models on strings generated according to the production rules of the Random Hierarchy Model (RHM), and find that samples composed of common substrings are preferentially memorized. This holds true even if the training data consists of entirely unique samples, indicating that deduplication at the data point level does not provide a meaningful privacy guarantee. Correspondingly we predict, then observe, delayed memorization for fat-tailed datasets (i.e., those with more atypical samples). This effect is amplified when fat-tails are introduced into high-level production rules. These together suggest that dataset diversity, particularly at higher levels of abstraction, plays an important role in staving off memorization. Finally, we identify an intermediate regime of partial memorization in which common substrings are learned first and subsequently overproduced during generation. If training is stopped in this regime, models will exhibit the reversion-to-the-mean blandness often derided as "slop".

cs.LG

Sampling Data with Chains of Forward-Backward Diffusion Steps

Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward steps of a diffusion model, in which each step proposes a move that remains on the learned data manifold and, paired with a Metropolis-Hastings correction, samples from energy-modified targets. For synthetic languages, we show that minimal U-turn dynamics undergoes an ergodicity-breaking phase transition driven by fragmentation of the data manifold; ergodicity is restored at larger U-turn magnitude. In the non-ergodic regime, low-level features relax faster than high-level ones, an ordering that inverts only at sufficiently large U-turn magnitude. We test these predictions on natural language and natural images. In both modalities, minimal U-turns relax slowly, especially for high-level features approximated by deep representations in CNNs or LLMs. The layer-ordering inversion appears only at large noise when mixing is efficient -- signatures consistent with strongly constrained, weakly mixing local dynamics. We discuss the implications of these results for sampling with diffusion models.

cs.LG

Learn from your own latents and not from tokens: A sample-complexity theory

Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which networks are trained to predict their \emph{own} latent representations of related views or masked regions, as in data2vec and JEPA -- an idea related to predictive-coding accounts of the cortex. Despite strong empirical results, the theoretical understanding of these methods remains limited. Central questions include: by how much does latent prediction actually improve data efficiency? Is there a benefit to stacking such methods into multi-scale hierarchies? We answer both using as data a tractable probabilistic context-free grammar that captures the compositional structure of natural language and images. Such a grammar generates strings of visible tokens by recursively applying production rules along a tree of hidden symbols of depth $L$. For such data, supervised or token-level SSL require a number of samples \emph{exponential} in $L$ to recover the latent tree; we prove that latent prediction achieves this with a number of samples \emph{constant} in $L$, up to logarithmic factors. We confirm this bound with (i) a hierarchical clustering algorithm, (ii) an end-to-end neural network whose predictor-clusterer modules predict their own latents at each level via gradient descent, and (iii) the first sample-complexity analysis of data2vec, which we show implicitly performs hierarchical latent prediction. This suggests that explicit stacking such as H-JEPA is largely redundant.

cs.LG

On the Emergence of Linear Analogies in Word Embeddings

Models such as Word2Vec and GloVe construct word embeddings based on the co-occurrence probability $P(i,j)$ of words $i$ and $j$ in text corpora. The resulting vectors $W_i$ not only group semantically similar words but also exhibit a striking linear analogy structure -- for example, $W_{\text{king}} - W_{\text{man}} + W_{\text{woman}} \approx W_{\text{queen}}$ -- whose theoretical origin remains unclear. Previous observations indicate that this analogy structure: (i) already emerges in the top eigenvectors of the matrix $M(i,j) = P(i,j)/P(i)P(j)$, (ii) strengthens and then saturates as more eigenvectors of $M (i, j)$, which controls the dimension of the embeddings, are included, (iii) is enhanced when using $\log M(i,j)$ rather than $M(i,j)$, and (iv) persists even when all word pairs involved in a specific analogy relation (e.g., king-queen, man-woman) are removed from the corpus. To explain these phenomena, we introduce a theoretical generative model in which words are defined by binary semantic attributes, and co-occurrence probabilities are derived from attribute-based interactions. This model analytically reproduces the emergence of linear analogy structure and naturally accounts for properties (i)-(iv). It can be viewed as giving fine-grained resolution into the role of each additional embedding dimension. It is robust to various forms of noise and agrees well with co-occurrence statistics measured on Wikipedia and the analogy benchmark introduced by Mikolov et al.

cs.CL

Microscopic description of the intermittent dynamics driving logarithmic creep

Disordered materials under an imposed forcing can display creep and aging effects, accompanied by intermittent, spatially heterogeneous dynamics. We propose a unifying microscopic description of these phenomena, based on the notion that as the system ages, the density of local barriers that enable relaxation displays a slowly evolving gap. As a result, the relaxation dynamics is dominated by the activation of the lowest, extremal tail of the distribution. This framework predicts logarithmic creep, as well as correlated bursts of slow activated rearrangements, or 'thermal avalanches', whose size grows logarithmically with their duration. The time interval between events within avalanches obeys a universal power-law distribution, with a cut-off that is simply proportional to the age of the system. We show that these predictions hold both in numerical models of amorphous solids, as well as in experiments with thin crumpled sheets. This analysis suggests that the heterogeneous dynamics occurring during logarithmic creep is related to other phenomena, including dynamical heterogeneities characterising the glass transition.

cond-mat.dis-nn

Thermally activated intermittent flow in amorphous solids

Using mean field theory and a mesoscale elastoplastic model, we analyze the steady state shear rheology of thermally activated amorphous solids. At sufficiently high temperature and driving rates, flow is continuous and described by well-established rheological flow laws such as Herschel-Bulkley and logarithmic rate dependence. However, we find that these flow laws change in the regime of intermittent flow, were collective events no longer overlap and serrated flow becomes pronounced. In this regime, we identify a thermal activation stress scale, $x_{a}(T,\dotγ)$, that wholly captures the effect of driving rate $\dotγ$ and temperature $T$ on average flow stress, stress drop (avalanche) size and correlation lengths. Different rheological regimes are summarized in a dynamic phase diagram for the amorphous yielding transition. Theoretical predictions call for a need to re-examine the rheology of very slowly sheared amorphous matter much below the glass transition.

cond-mat.soft

Criticality in spreading processes without time-scale separation and the critical brain hypothesis

Spreading processes on networks are ubiquitous in both human-made and natural systems. Understanding their behavior is of broad interest; from the control of epidemics to understanding brain dynamics. While in some cases there exists a clear separation of time scales between the propagation of a single spreading cascade and the initiation of the next -- such that spreading can be modelled as directed percolation or a branching process -- there are also processes for which this is not the case, such as zoonotic diseases or spiking cascades in neural networks. For a large class of relevant network topologies, we show here that in such a scenario the nature of the overall spreading fundamentally changes. This change manifests itself in a transition between different universality classes of critical spreading, which determines the onset and the properties of an avalanche turning epidemic or neural activity turning epileptic, for example. We present analytical results in the mean-field limit giving the critical line along which scale-free spreading behaviour can be observed. The two limits of this critical line correspond to the universality classes of directed and undirected percolation, respectively. Outside these two limits, this duality manifests itself in the appearance of critical exponents from the universality classes of both directed and undirected percolation. We find that the transition between these exponents is governed by a competition between merging and propagation of activity, and identify an appropriate scaling relationship for the transition point. Finally, we show that commonly used measures, such as the branching ratio and dynamic susceptibility, fail to establish criticality in the absence of time-scale separation calling for a reanalysis of criticality in the brain.

cond-mat.stat-mech