SearcharxivSearch

arXiv subjects

Brandon Soubasis

Publications and source records attributed to Brandon Soubasis.

3 recordsLinked to original sources

NVIDIA Nemotron 3: Efficient and Open Intelligence

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to 1M tokens. Super and Ultra models are trained with NVFP4 and incorporate LatentMoE, a novel approach that improves model quality. The two larger models also include MTP layers for faster text generation. All Nemotron 3 models are post-trained using multi-environment reinforcement learning enabling reasoning, multi-step tool use, and support granular reasoning budget control. Nano, the smallest model, outperforms comparable models in accuracy while remaining extremely cost-efficient for inference. Super is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Ultra, the largest model, provides state-of-the-art accuracy and reasoning performance. Nano is released together with its technical report and this white paper, while Super and Ultra will follow in the coming months. We will openly release the model weights, pre- and post-training software, recipes, and all data for which we hold redistribution rights.

cs.CL

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.

cs.CL

Probing axion-like particles with $γγ$ final states from vector boson fusion processes at the LHC

We perform a feasibility study to search for axion-like particles (ALPs) using vector boson fusion (VBF) processes at the LHC. We work in an effective field theory framework with cutoff scale $Λ$ and ALP mass $m_{a}$, and assume that ALPs couple to photons with strength $\propto 1/Λ$. Assuming proton-proton collisions at $\sqrt{s} = 13$ TeV, we present the total VBF ALP production cross sections, ALP decay widths and lifetimes, and relevant kinematic distributions as a function of $m_{a}$ and $Λ$. We consider the $a\toγγ$ decay mode to show that the requirement of an energetic diphoton pair combined with two forward jets with large dijet mass and pseudorapidity separation can significantly reduce the Standard Model backgrounds, leading to a $5σ$ discovery reach for $10 \text{ MeV} \lesssim m_{a} \lesssim 1$ TeV with $Λ\lesssim 2$ TeV, assuming an integrated luminosity of 3000 fb$^{-1}$. In particular, this extends the LHC sensitivity to a previously unstudied region of the ALP parameter space.

hep-ph