SearcharxivSearch

arXiv subjects

Yiyang Jia

Publications and source records attributed to Yiyang Jia.

At least 19 recordsLinked to original sources

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine RL with verifiable process rewards and SFT for learning $k$-sparse Boolean functions with a one-layer transformer through intermediate reasoning steps akin to CoT. In particular, we consider Boolean functions that can be recursively decomposed into fixed 2-sparse Boolean functions. We first analyze the learning dynamics of RL fine-tuning with verifiable process rewards and SFT in a unified way, allowing us to identify sufficient conditions under which the transformer provably learns these functions. We then verify that the conditions hold for three examples, including $k$-PARITY, $k$-AND, and $k$-OR, thus demonstrating their learnability via both RL and SFT. Notably, we reveal that RL and SFT exhibit distinct learning behaviors depending on supervision: RL learns the whole CoT chain simultaneously, whereas SFT without teacher forcing learns the CoT step-by-step. Overall, our findings provide insights on the mechanisms underlying RL and SFT and how they differ in triggering the CoT capabilities of transformers, and suggest that the comparison between RL and SFT should consider the intermediate supervision.

cs.LG

Self-Motivated Growing Neural Network for Adaptive Architecture via Local Structural Plasticity

Control policies are often implemented with fixed-capacity multilayer perceptrons trained by backpropagation, which require architecture selection in advance and cannot adapt their capacity during learning. This paper introduces the Self-Motivated Growing Neural Network (SMGrNN), a gradient-trained controller whose topology evolves online through a local Structural Plasticity Module (SPM). The SPM monitors edge-wise weight update statistics over short temporal windows and uses these local signals to trigger neuron insertion and pruning, while synaptic weights are optimized by a standard gradient-based optimizer. This allows network capacity to be adjusted during learning without manual architectural tuning. SMGrNN is evaluated on control benchmarks via policy distillation. Compared with multilayer perceptron baselines, it achieves similar or higher returns, lower variance, and task-appropriate network sizes. Ablation studies with growth disabled and growth-only variants isolate the role of structural plasticity, showing that adaptive growth improves reward stability while pruning prevents uncontrolled expansion and supports compact network formation. These results establish the independent value of local structural plasticity within gradient-trained networks and motivate future investigation of whether similar structural rules can be extended to more local or spike-based learning settings.

cs.NE

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward general-purpose task capability, yet their embodied safety limits remain poorly understood. To address this gap, we introduce ForesightSafety-VLA, a diagnostic benchmark that makes safety the primary evaluation target for VLA systems. We define a 13-category safety taxonomy covering physical interaction safety (Safe-Core), instruction-side safety (Safe-Lang), and perception-side safety (Safe-Vis), and evaluate policies under three controlled dimensions of variation -- scene structure, language command, and visual observation -- so that failure sources can be diagnosed rather than hidden in a single aggregate score. Beyond binary task success, ForesightSafety-VLA measures process-level risk through cumulative safety cost (CC) and risk exposure time (RET), together with a four-quadrant decomposition of safe/unsafe success and failure. We instantiate 66 safety-augmented base scenarios in RoboTwin across 5 embodiments and report results on representative VLA baselines. Across the evaluated baselines, even the strongest policy incurs non-trivial safety cost and unsafe nominal success, while structure and visual variation induce substantially stronger safety degradation than ordinary language variation. These results suggest that embodied safety is tightly coupled to perception, grounding, and control competence rather than being reducible to post-hoc safety filtering alone.

cs.RO

When Autoregressive Consistency Hurts Safety Alignment

Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens. We argue that this phenomenon can be understood through autoregressive consistency, the tendency of next-token prediction to preserve and extend the current response trajectory consistently. By analyzing the learning dynamics of safety alignment, we show that autoregressive consistency can concentrate alignment updates on early tokens, offering a mechanistic explanation for shallow safety alignment. The same mechanism also predicts a broader class of attacks on LLMs: attacks that induce harmful continuation states at arbitrary positions in the output trajectory. As a concrete example, we introduce random insertion attack, which inserts a short harmful span into an otherwise safe refusal trajectory and exploits autoregressive consistency to sustain the resulting harmful branch, thereby bypassing safety alignment. Notably, a short harmful span can redirect the generation to be harmful even after a long refusal prefix, highlighting autoregressive consistency as a potential broader failure mechanism. This suggests that safety alignment should also break harmful autoregressive consistency throughout the output trajectory. We therefore propose adversarial safety alignment, an initial framework based on worst-case harmful continuation states, and instantiate it with random worst-insertion training. Overall, our results suggest that autoregressive consistency should be treated as a central consideration in both safety alignment and attack design.

cs.LG

Modeling GRNs with a Probabilistic Categorical Framework

Understanding the complex and stochastic nature of Gene Regulatory Networks (GRNs) remains a central challenge in systems biology. Existing modeling paradigms often struggle to effectively capture the intricate, multi-factor regulatory logic and to rigorously manage the dual uncertainties of network structure and kinetic parameters. In response, this work introduces the Probabilistic Categorical GRN(PC-GRN) framework. It is a novel theoretical approach founded on the synergistic integration of three core methodologies. Firstly, category theory provides a formal language for the modularity and composition of regulatory pathways. Secondly, Bayesian Typed Petri Nets (BTPNs) serve as an interpretable,mechanistic substrate for modeling stochastic cellular processes, with kinetic parameters themselves represented as probability distributions. The central innovation of PC-GRN is its end-to-end generative Bayesian inference engine, which learns a full posterior distribution over BTPN models (P (G, Θ|D)) directly from data. This is achieved by the novel interplay of a GFlowNet, which learns a policy to sample network topologies, and a HyperNetwork, which performs amortized inference to predict their corresponding parameter distributions. The resulting framework provides a mathematically rigorous, biologically interpretable, and uncertainty-aware representation of GRNs, advancing predictive modeling and systems-level analysis.

q-bio.MN

Contact diagrams for chords

We demonstrate how contact chord diagrams can arise from certain Fock-space models and compute the corresponding correlation functions using the chord path integral technique. In particular, our three-point functions are in the right form dictated by conformal symmetry, and some of our four-point functions match the results of some AdS$_2$ contact Witten diagrams.

hep-th

Category-Theoretical and Topos-Theoretical Frameworks in Machine Learning: A Survey

In this survey, we provide an overview of category theory-derived machine learning from four mainstream perspectives: gradient-based learning, probability-based learning, invariance and equivalence-based learning, and topos-based learning. For the first three topics, we primarily review research in the past five years, updating and expanding on the previous survey by Shiebler et al.. The fourth topic, which delves into higher category theory, particularly topos theory, is surveyed for the first time in this paper. In certain machine learning methods, the compositionality of functors plays a vital role, prompting the development of specific categorical frameworks. However, when considering how the global properties of a network reflect in local structures and how geometric properties are expressed with logic, the topos structure becomes particularly significant and profound.

cs.LG

From Chaos to Integrability in Double Scaled SYK via a Chord Path Integral

We study thermodynamic phase transitions between integrable and chaotic dynamics. We do so by analyzing models that interpolate between the chaotic double scaled Sachdev-Ye-Kitaev (SYK) and the integrable $p$-spin systems, in a limit where they are described by chord diagrams. We develop a path integral formalism by coarse graining over the diagrams, which we use to argue that the system has two distinct phases: one is continuously connected to the chaotic system, and the other to the integrable. They are separated by a line of first order transition that ends at some finite temperature.

hep-th

A Path Integral for Chord Diagrams and Chaotic-Integrable Transitions in Double Scaled SYK

We study transitions from chaotic to integrable Hamiltonians in the double scaled SYK and $p$-spin systems. The dynamics of our models is described by chord diagrams with two species. We begin by developing a path integral formalism of coarse graining chord diagrams with a single species of chords, which has the same equations of motion as the bi-local ($GΣ$) Liouville action, yet appears otherwise to be different and in particular well defined. We then develop a similar formalism for two types of chords, allowing us to study different types of deformations of double scaled SYK and in particular a deformation by an integrable Hamiltonian. The system has two distinct thermodynamic phases: one is continuously connected to the chaotic SYK Hamiltonian, the other is continuously connected to the integrable Hamiltonian, separated at low temperature by a first order phase transition. We also analyze the phase diagram for generic deformations, which in some cases includes a zero-temperature phase transition.

hep-th

Bosonic near-CFT$_1$ models from Fock-space fluxes

We construct a family of near-CFT$_1$ models with a conserved U(1) charge, whose basic degrees of freedom are canonical bosons. The Sachdev-Ye-Kitaev (SYK) model -- the first microscopic model that realizes the near-CFT$_1$ dynamics -- is based on random $p$-local interactions among fermions. However, a bosonic near-CFT$_1$ model has remained elusive in the $p$-local approach because such constructions generally suffer from unwanted orderings at low temperatures. Our construction is based on a recent insight that near-CFT$_1$ dynamics can quite generally arise if we place a large amount of random fluxes in a many-body Fock space and $p$-locality is not essential. All such models are essentially solved by chord diagrams regardless of the nature of the underlying degrees of freedom. We further argue that such bosonic models do not suffer from energetic instablities or unwanted low-temperature orderings. For comparison we also consider a second class of charge-conserving models which are based on qubits. The thermodynamic scalings of these models are very similar to those of the double-scaled complex SYK model but are free of certain singularities the latter suffers from. We also show the level statistics of both models are described by random matrix theory universality down to very low energies.

hep-th

Parisi's hypercube, Fock-space fluxes, and the microscopics of near-AdS$_2$/near-CFT$_1$ duality

Parisi's hypercube model describes a charged particle hopping on a $d$-dimensional hypercube with disordered background fluxes in the large $d$ limit. It was noted previously [Jia and Verbaarschot, J. High Energy Phys. 11 (2020) 154] that the hypercube model at leading order in $1/d$ has the same spectral density as the double-scaled Sachdev-Ye-Kitaev (DS-SYK) model. In this work we identify the set of observables that have the same correlation functions as the DS-SYK model, demonstrating that the hypercube model is an equally good microscopic model for near-AdS$_2$/near-CFT$_1$ holography. Unlike the SYK model, the hypercube model is not $p$-local. Rather, we note that the shared feature between the two models is that they both have a large amount of disordered but uniform fluxes on their Fock-space graphs, and we propose this is a broader characterization of near-CFT$_1$ microscopics. Moreover, we suggest that the hypercube model can be viewed as the operator growth model of the DS-SYK model. We explain some universality in subleading corrections and relate them to bulk vertices. Finally, we revise a claim made the aforementioned reference about the existence of a spectral gap.

hep-th

Parisi's hypercube, Fock-space frustration and near-AdS$_2$/near-CFT$_1$ holography

We consider a model of Parisi where a single particle hops on an infinite-dimensional hypercube, under the influence of a uniform but disordered magnetic flux. We reinterpret the hypercube as the Fock-space graph of a many-body Hamiltonian and the flux as a frustration of the return amplitudes in Fock space. We will identify the set of observables that have the same correlation functions as the double-scaled Sachdev-Ye-Kitaev (DS-SYK) model, and hence the hypercube model is an equally good quantum model for near-AdS$_2$/near-CFT$_{1}$ (NAdS$_2$/NCFT$_1$) holography. Unlike the SYK model, the hypercube Hamiltonian is not $p$ local. Instead, the SYK model can be understood as a Fock-space model with similar frustrations. Hence we propose this type of Fock-space frustration as the broader characterization for NAdS$_2$/NCFT$_1$ microscopics, which encompasses the hypercube and the DS-SYK models as two specific examples. We then speculate on the possible origin of such frustrations.

hep-th

Replica Symmetry Breaking in Random Non-Hermitian Systems

Recent studies have revealed intriguing similarities between the contribution of wormholes to the gravitational path integral and the phenomenon of replica symmetry breaking observed in spin glasses and other disordered systems. Interestingly, these configurations may also be important for the explanation of the information paradox of quantum black holes. Motivated by these developments, we investigate the thermodynamic properties of a $PT$-symmetric system composed of two random non-Hermitian Hamiltonians with no explicit coupling between them. After performing ensemble averaging, we identify numerically and analytically a robust first-order phase transition in the free energy of two models with quantum chaotic dynamics: the elliptic Ginibre ensemble of random matrices and a non-Hermitian Sachdev-Ye-Kitaev (SYK) model. The free energy of the Ginibre model is temperature-independent in the low-temperature phase. The SYK model has a similar behavior for sufficiently low temperature, then it experiences a possible continuous phase transition to a phase with a temperature-dependent free energy before the first-order transition takes place at a higher temperature. We identify the order parameter of the first-order phase transition and obtain analytical expressions for the critical temperature. The mechanism behind the transition is the existence of replica symmetry breaking configurations coupling Left and Right replicas that control the low-temperature limit of the partition function. We speculate that quantum chaos may be necessary for the observed dominance of off-diagonal replica symmetry breaking configurations in the low-temperature limit.

hep-th

Replica Symmetry Breaking for the Integrable Two-Site Sachdev-Ye-Kitaev Model

We analyze a two-body nonhermitian two-site Sachdev-Ye-Kitaev model with the couplings of one site complex conjugated to the other site. This model, with no explicit coupling between the sites, shows an infinite number of phase transitions which is a consequence of the partition function factorizing into a product over Matsubara frequencies. We calculate the quenched free energy in two different ways, first in terms of the single-particle energies, and second by solving the Schwinger-Dyson equations. The first calculation can be done entirely in terms of a one-site model. The conjugate replica enters due to non-analyticities when Matsubara frequencies enter the spectral support of the coupling matrix. The second calculation is based on the replica trick of the two-site partition function. Both methods give the same result. The free-fermion partition function can be rephrased as a matrix model for the coupling matrix. Up to minor details, this model is the random matrix model that describes the chiral phase transition of QCD, and the order parameter of the two-body model corresponds to the chiral condensate of QCD. Comparing to the corresponding four-body model, we are able to determine which features of the free energy are due to chaotic nature of the four-body model. The high-temperature phase of both models is entropy dominated, and in both cases is determined by the spectral density. The four-body SYK model has a low-temperature phase whose free energy is almost temperature-independent, signaling an effective gap of the theory even though the actual spectrum does not exhibit a gap. However the low-temperature free energy of the two-body SYK model is not flat, in fact it oscillates to arbitrarily low temperature. This indicates a less desirable feature that the entropy of the two-body model is not always positive, which most likely is a consequence of the nonhermiticity.

hep-th

Replica Symmetry Breaking and Phase Transitions in a PT Symmetric Sachdev-Ye-Kitaev Model

We show that the low temperature phase of a conjugate pair of uncoupled, quantum chaotic, nonhermitian systems such as the Sachdev-Ye-Kitaev (SYK) model or the Ginibre ensemble of random matrices are dominated by replica symmetry breaking configurations with a nearly flat free energy that terminates in a first order phase transition. In the case of the SYK model, we show explicitly that the spectrum of the effective replica theory has a gap. These features are strikingly similar to those induced by wormholes in the gravity path integral which suggests a close relation between both configurations. For a non-chaotic SYK, the results are qualitatively different: the spectrum is gapless in the low temperature phase and there is an infinite number of second order phase transitions unrelated to the restoration of replica symmetry.

hep-th

Chaos on the hypercube

We analyze the spectral properties of a $d$-dimensional HyperCubic (HC) lattice model originally introduced by Parisi. The U(1) gauge links of this model give rise to a magnetic flux of constant magnitude $ϕ$ but random orientation through the faces of the hypercube. The HC model, which also can be written as a model of $2d$ interacting Majorana fermions, has a spectral flow that is reminiscent of the Maldacena-Qi (MQ) model, and its spectrum at $ϕ=0$, actually coincides with the coupling term of the MQ model. As was already shown by Parisi, at leading order in $1/d$ , the spectral density of this model is given by the density function of the Q-Hermite polynomials, which is also the spectral density of the double-scaled Sachdev-Ye-Kitaev model. Parisi demonstrated this by mapping the moments of the HC model to $Q$-weighted sums on chord diagrams. We point out that the subleading moments of the HC model can also be mapped to weighted sums on chord diagrams, in a manner that descends from the leading moments. The HC model has a magnetic inversion symmetry that depends on both the magnitude and the orientation of the magnetic flux through the faces of the hypercube. The spectrum for fixed quantum number of this symmetry exhibits a transition from regular spectra at $ϕ=0$ to chaotic spectra with spectral statistics given by the Gaussian Unitary Ensembles (GUE) for larger values of $ϕ$. For small magnetic flux, the ground state is gapped and is close to a Thermofield Double (TFD) state.

hep-th

Sparse Sachdev-Ye-Kitaev model, quantum chaos and gravity duals

We study a sparse Sachdev-Ye-Kitaev (SYK) model with $N$ Majoranas where only $\sim k N$ independent matrix elements are non-zero. We identify a minimum $k \gtrsim 1$ for quantum chaos to occur by a level statistics analysis. The spectral density in this region, and for a larger $k$, is still given by the Schwarzian prediction of the dense SYK model, though with renormalized parameters. Similar results are obtained for a beyond linear scaling with $N$ of the number of non-zero matrix elements. This is a strong indication that this is the minimum connectivity for the sparse SYK model to still have a quantum gravity dual. We also find an intriguing exact relation between the leading correction to moments of the spectral density due to sparsity and the leading $1/d$ correction of Parisi's U(1) lattice gauge theory in a $d$ dimensional hypercube. In the $k \to 1$ limit, different disorder realizations of the sparse SYK model show emergent random matrix statistics that for fixed $N$ can be in any universality class of the ten-fold way. The agreement with random matrix statistics is restricted to short range correlations, no more than a few level spacings, in particular in the tail of the spectrum. In addition, emergent discrete global symmetries in most of the disorder realizations for $k$ slightly below one give rise to $2^m$-fold degenerate spectra, with $m$ being a positive integer. For $k =3/4$, we observe a large number of such emergent global symmetries with a maximum $2^8$-fold degenerate spectra for $N = 26$.

hep-th