SearcharxivSearch

arXiv subjects

Srivatsan Rajagopal

Publications and source records attributed to Srivatsan Rajagopal.

8 recordsLinked to original sources

ZAYA1-8B Technical Report

We present ZAYA1-8B, a reasoning-focused mixture-of-experts (MoE) model with 700M active and 8B total parameters, built on Zyphra's MoE++ architecture. ZAYA1-8B's core pretraining, midtraining, and supervised fine-tuning (SFT) were performed on a full-stack AMD compute, networking, and software platform. With under 1B active parameters, ZAYA1-8B matches or exceeds DeepSeek-R1-0528 on several challenging mathematics and coding benchmarks, and remains competitive with substantially larger open-weight reasoning models. ZAYA1-8B was trained from scratch for reasoning, with reasoning data included from pretraining onward using an answer-preserving trimming scheme. Post-training uses a four-stage RL cascade: reasoning warmup on math and puzzles; a 400-task RLVE-Gym curriculum; math and code RL with test-time compute traces and synthetic code environments built from competitive-programming references; and behavioral RL for chat and instruction following. We also introduce Markovian RSA, a test-time compute method that recursively aggregates parallel reasoning traces while carrying forward only bounded-length reasoning tails between rounds. In TTC evaluation, Markovian RSA raises ZAYA1-8B to 91.9\% on AIME'25 and 89.6\% on HMMT'25 while carrying forward only a 4K-token tail, narrowing the gap to much larger reasoning models including Gemini-2.5 Pro, DeepSeek-V3.2, and GPT-5-High.

cs.AI

Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design

We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance for both systems and model design. On the systems side, we deliver a comprehensive cluster and networking characterization: microbenchmarks for all core collectives (all-reduce, reduce-scatter, all-gather, broadcast) across message sizes and GPU counts over Pollara. To our knowledge, this is the first at this scale. We further provide MI300X microbenchmarks on kernel sizing and memory bandwidth to inform model design. On the modeling side, we introduce and apply MI300X-aware transformer sizing rules for attention and MLP blocks and justify MoE widths that jointly optimize training throughput and inference latency. We describe our training stack in depth, including often-ignored utilities such as fault-tolerance and checkpoint-reshaping, as well as detailed information on our training recipe. We also provide a preview of our model architecture and base model - ZAYA1 (760M active, 8.3B total parameters MoE, available at https://huggingface.co/Zyphra/ZAYA1-base) - which will be further improved upon in forthcoming papers. ZAYA1-base achieves performance comparable to leading base models such as Qwen3-4B and Gemma3-12B at its scale and larger, and outperforms models including Llama-3-8B and OLMoE across reasoning, mathematics, and coding benchmarks. Together, these results demonstrate that the AMD hardware, network, and software stack are mature and optimized enough for competitive large-scale pretraining.

cs.CL

Modular Flow of Excited States

We develop new techniques for studying the modular and the relative modular flows of general excited states. We show that the class of states obtained by acting on the vacuum (or any cyclic and separating state) with invertible operators from the algebra of a region is dense in the Hilbert space. This enables us to express the modular and the relative modular operators, as well as the relative entropies of generic excited states in terms of the vacuum modular operator and the operator that creates the state. In particular, the modular and the relative modular flows of any state can be expanded in terms of the modular flow of operators in vacuum. We illustrate the formalism with simple examples including states close to the vacuum, and coherent and squeezed states in generalized free field theory.

hep-th

Perturbation Theory for the Logarithm of a Positive Operator

In various contexts in mathematical physics one needs to compute the logarithm of a positive unbounded operator. Examples include the von Neumann entropy of a density matrix and the flow of operators with the modular Hamiltonian in the Tomita-Takesaki theory. Often, one encounters the situation where the operator under consideration, that we denote by $Δ$, can be related by a perturbative series to another operator $Δ_0$, whose logarithm is known. We set up a perturbation theory for the logarithm $\log Δ$. It turns out that the terms in the series possess remarkable algebraic structure, which enable us to write them in the form of nested commutators plus some "contact terms."

hep-th

Global Anomalies, Discrete Symmetries, and Hydrodynamic Effective Actions

We derive effective actions for parity-violating fluids in both $(3+1)$ and $(2+1)$ dimensions, including those with anomalies. As a corollary we confirm the most general constitutive relations for such systems derived previously using other methods. We discuss in detail connections between parity-odd transport and underlying discrete symmetries. In (3+1) dimensions we elucidate connections between anomalous transport coefficients and global anomalies, and clarify a previous puzzle concerning transports and local gravitational anomalies.

hep-th

Holographic Trace Anomaly and Local Renormalization Group

The Hamilton-Jacobi method in holography has produced important results both at a renormalization group (RG) fixed point and away from it. In this paper we use the Hamilton-Jacobi method to compute the holographic trace anomaly for four- and six-dimensional boundary conformal field theories (CFTs), assuming higher-derivative gravity and interactions of scalar fields in the bulk. The scalar field contributions to the anomaly appear in CFTs with exactly marginal operators. Moving away from the fixed point, we show that the Hamilton-Jacobi formalism provides a deep connection between the holographic and the local RG. We derive the local RG equation holographically, and verify explicitly that it satisfies Weyl consistency conditions stemming from the commutativity of Weyl scalings. We also consider massive scalar fields in the bulk corresponding to boundary relevant operators, and comment on their effects to the local RG equation.

hep-th

Quantization of B-I electrodynamics and B-I modified gravity using Faddeev-Popov gauge-fixing procedure

We investigate the quantized versions of Born Infeld electrodynamics and Born Infeld Gravity. We derive Feynman rules for B-I electrodynamics by deriving an effective Lagrangian with the square root removed using the Faddeev-Popov method. In the case of B-I gravity, the square root in the Lagrangian is removed by the introduction of the Vierbien fields. This approach has the advantage that SO(3,1) can be consistently regarded to be the gauge group of gravity. Finally, using a rough argument, the quantum fluctuations of the radii of spatial hypersurfaces in flat space are shown to undergo accelerated increase with time.

gr-qc

Ostrogradsky instability and Born-Infeld modified cosmology in Palatini formalism

The Ostrogradsky instability of higher derivative Lagrangians is derived from first principles using Control theory and Lyapunov Stability Analysis. This result is then used to argue that Born-Infeld Lagrangians are viable modifications of the Einstein-Hilbert action provided the action is varied in accordance with the Palatini formalism, in contrast to the metric formalism. Finally, the Born-Infeld version of the FRW equations are derived and the cosmological dynamics is studied for matter dominated closed and open universes, and the results are compared with the usual cosmology.

gr-qc