SearcharxivSearch

arXiv subjects

Anton Sugolov

Publications and source records attributed to Anton Sugolov.

3 recordsLinked to original sources

A Framework for Stochastic Differentiable Programming

We introduce Parametrized Stochastic Circuits (PSCs), a gate-based intermediate representation for programmable stochastic dynamics in which typed local stochastic kernels with tunable parameters compose over explicit binary, categorical, and continuous wires, and \texttt{torx}, an open-source JAX framework for constructing, executing, and differentiating them. PSCs' data types and stochastic kernels are chosen to align closely with the native operations exposed by emerging probabilistic hardware. In this way, stochastic algorithms can be designed directly in terms of the operations the hardware executes natively, so that the energy advantage arising at this level is not lost on mappings that introduce substantial decomposition, communication, or control overhead. We demonstrate the framework on a variety of example applications such as random walks on graphs, discrete diffusion, stochastic graph networks, jump diffusion and Ising sampling. We also report a hardware experiment in which probabilistic bits on the X0 subthreshold CMOS test chip, hosted by the XTR-0 desktop platform, provide physical randomness for Metropolis-Hastings and importance-sampling estimators, yielding estimates consistent with a software pseudorandom baseline.

cs.ET

Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization

Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training. Motivated by this observation, we introduce \emph{Depth-wise Gradient Augmentation}, a general optimization paradigm in which the update applied to each layer is obtained by transforming the collection of block-wise optimizer updates along the depth dimension. Within this framework, we study \emph{Gradient Smoothing}, a family of depth-wise smoothing methods, and instantiate it with a simple local \emph{Window Smoothing} operator. The resulting method operates directly on block-wise updates produced by arbitrary base optimizers (e.g., SGD, Adam, Muon), incurs minimal computational overhead, and is compatible with existing optimization pipelines. We evaluate Gradient Smoothing across a diverse set of architectures and training regimes, including language model pretraining, RL post-training of LLMs for reasoning, diffusion modeling, and image classification with Vision Transformers. Across these settings, Gradient Smoothing consistently improves optimization and generalization performance without modifying model architectures or training objectives. We further show that it promotes more structured representation evolution across depth, consistent with its interpretation as a structured depth-wise preconditioning method. Together, these results establish Depth-wise Gradient Augmentation as a promising framework for exploiting cross-depth structure in optimization and demonstrate Gradient Smoothing as a simple and broadly applicable instantiation.

cs.LG

Transformer Block Coupling and its Correlation with Generalization in LLMs

Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential. In this work, we analyze the trajectories of token embeddings as they pass through transformer blocks, linearizing the system along these trajectories through their Jacobian matrices. By examining the relationships between these block Jacobians, we uncover the phenomenon of \textbf{transformer block coupling} in a multitude of LLMs, characterized by the coupling of their top singular vectors across tokens and depth. Our findings reveal that coupling \textit{positively correlates} with model performance, and that this relationship is stronger than with other hyperparameters such as parameter count, model depth, and embedding dimension. We further investigate how these properties emerge during training, observing a progressive development of coupling, increased linearity, and layer-wise exponential growth in token trajectories. Additionally, experiments with Vision Transformers (ViTs) corroborate the emergence of coupling and its relationship with generalization, reinforcing our findings in LLMs. Collectively, these insights offer a novel perspective on token interactions in transformers, opening new directions for studying their mechanisms as well as improving training and generalization.

cs.LG