SearcharxivSearch

arXiv subjects

Dusan Stosic

Publications and source records attributed to Dusan Stosic.

16 recordsLinked to original sources

Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery

This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying distillation to quantized models is not a new idea, we observe key advantages of QAD for today's LLMs: 1. It shows remarkable effectiveness and stability for models trained through multi-stage post-training pipelines, including supervised fine-tuning (SFT), reinforcement learning (RL), and model merging, where traditional quantization-aware training (QAT) suffers from engineering complexity and training instability; 2. It is robust to data quality and coverage, enabling accuracy recovery without full training data. We evaluate QAD across multiple post-trained models including AceReason Nemotron, Nemotron 3 Nano, Nemotron Nano V2, Nemotron Nano V2 VL (VLM), and Llama Nemotron Super v1, showing consistent recovery to near-BF16 accuracy.

cs.LG

NVIDIA Nemotron 3: Efficient and Open Intelligence

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to 1M tokens. Super and Ultra models are trained with NVFP4 and incorporate LatentMoE, a novel approach that improves model quality. The two larger models also include MTP layers for faster text generation. All Nemotron 3 models are post-trained using multi-environment reinforcement learning enabling reasoning, multi-step tool use, and support granular reasoning budget control. Nano, the smallest model, outperforms comparable models in accuracy while remaining extremely cost-efficient for inference. Super is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Ultra, the largest model, provides state-of-the-art accuracy and reasoning performance. Nano is released together with its technical report and this white paper, while Super and Ultra will follow in the coming months. We will openly release the model weights, pre- and post-training software, recipes, and all data for which we hold redistribution rights.

cs.CL

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.

cs.CL

Pretraining Large Language Models with NVFP4

Large Language Models (LLMs) today are powerful problem solvers across many domains, and they continue to get stronger as they scale in model size, training set size, and training set quality, as shown by extensive research and experimentation across the industry. Training a frontier model today requires on the order of tens to hundreds of yottaflops, which is a massive investment of time, compute, and energy. Improving pretraining efficiency is therefore essential to enable the next generation of even more capable LLMs. While 8-bit floating point (FP8) training is now widely adopted, transitioning to even narrower precision, such as 4-bit floating point (FP4), could unlock additional improvements in computational speed and resource utilization. However, quantization at this level poses challenges to training stability, convergence, and implementation, notably for large-scale models trained on long token horizons. In this study, we introduce a novel approach for stable and accurate training of large language models (LLMs) using the NVFP4 format. Our method integrates Random Hadamard transforms (RHT) to bound block-level outliers, employs a two-dimensional quantization scheme for consistent representations across both the forward and backward passes, utilizes stochastic rounding for unbiased gradient estimation, and incorporates selective high-precision layers. We validate our approach by training a 12-billion-parameter model on 10 trillion tokens -- the longest publicly documented training run in 4-bit precision to date. Our results show that the model trained with our NVFP4-based pretraining technique achieves training loss and downstream task accuracies comparable to an FP8 baseline. These findings highlight that NVFP4, when combined with our training approach, represents a major step forward in narrow-precision LLM training algorithms.

cs.CL

Recipes for Pre-training LLMs with MXFP8

Using fewer bits to represent model parameters and related tensors during pre-training has become a required technique for improving GPU efficiency without sacrificing accuracy. Microscaling (MX) formats introduced in NVIDIA Blackwell generation of GPUs represent a major advancement of this technique, making it practical to combine narrow floating-point data types with finer granularity per-block scaling factors. In turn, this enables both quantization of more tensors than previous approaches and more efficient execution of operations on those tensors. Effective use of MX-formats requires careful choices of various parameters. In this paper we review these choices and show how MXFP8-E4M3 datatype and a specific number conversion algorithm result in training sessions that match those carried out in BF16. We present results using models with up to 8B parameters, trained on high-quality datasets of up to 15T tokens.

cs.LG

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer models designed to reduce inference cost for a given accuracy level. To achieve this goal, we replace the majority of self-attention layers in the common Transformer model architecture with Mamba layers that perform constant computation and require constant memory per generated token. We show that Nemotron-H models offer either better or on-par accuracy compared to other similarly-sized state-of-the-art open-sourced Transformer models (e.g., Qwen-2.5-7B/72B and Llama-3.1-8B/70B), while being up to 3$\times$ faster at inference. To further increase inference speed and reduce the memory required at inference time, we created Nemotron-H-47B-Base from the 56B model using a new compression via pruning and distillation technique called MiniPuzzle. Nemotron-H-47B-Base achieves similar accuracy to the 56B model, but is 20% faster to infer. In addition, we introduce an FP8-based training recipe and show that it can achieve on par results with BF16-based training. This recipe is used to train the 56B model. We are releasing Nemotron-H base model checkpoints with support in Hugging Face and NeMo.

cs.CL

Anomalous optical response of graphene on hexagonal boron nitride substrates

Graphene/hBN heterostructures can be considered as one of the basic building blocks for the next-generation optoelectronics mostly owing to the record-high electron mobilities. However, currently, the studies of the intrinsic optical properties of graphene are limited to the standard substrates (SiO2/Si, glass, quartz) despite the growing interest in graphene/hBN heterostructures. This can be attributed to a challenging task of the determination of hBN's strongly anisotropic dielectric tensor in the total optical response. In this study, we overcome this issue through imaging spectroscopic ellipsometry utilizing simultaneous analysis of hBN's optical response with and without graphene monolayers. Our technique allowed us to retrieve the optical constants of graphene from graphene/hBN heterostructures in a broad spectral range of 250-950 nm. Our results suggest that graphene's absorption on hBN may exceed the one of graphene on SiO2/Si by about 60 %.

cond-mat.mes-hall

FP8 Formats for Deep Learning

FP8 is a natural progression for accelerating deep learning training inference beyond the 16-bit formats common in modern processors. In this paper we propose an 8-bit floating point (FP8) binary interchange format consisting of two encodings - E4M3 (4-bit exponent and 3-bit mantissa) and E5M2 (5-bit exponent and 2-bit mantissa). While E5M2 follows IEEE 754 conventions for representatio of special values, E4M3's dynamic range is extended by not representing infinities and having only one mantissa bit-pattern for NaNs. We demonstrate the efficacy of the FP8 format on a variety of image and language tasks, effectively matching the result quality achieved by 16-bit training sessions. Our study covers the main modern neural network architectures - CNNs, RNNs, and Transformer-based models, leaving all the hyperparameters unchanged from the 16-bit baseline training sessions. Our training experiments include large, up to 175B parameter, language models. We also examine FP8 post-training-quantization of language models trained using 16-bit formats that resisted fixed point int8 quantization.

cs.LG

Ising models of deep neural networks

This work maps deep neural networks to classical Ising spin models, allowing them to be described using statistical thermodynamics. The density of states shows that structures emerge in the weights after they have been trained -- well-trained networks span a much wider range of realizable energies compared to poorly trained ones. These structures propagate throughout the entire network and are not observed in individual layers. The energy values correlate to performance on tasks, making it possible to distinguish networks based on quality without access to data. Thermodynamic properties such as specific heat are also studied, revealing a higher critical temperature in trained networks.

cond-mat.stat-mech

Generalized Weighted Permutation Entropy

A novel heuristic approach is proposed here for time series data analysis, dubbed Generalized weighted permutation entropy, which amalgamates and generalizes beyond their original scope two well established data analysis methods: Permutation entropy, and Weighted permutation entropy. The method introduces a scaling parameter to discern the disorder and complexity of ordinal patterns with small and large fluctuations. Using this scaling parameter, the complexity-entropy causality plane is generalized to the complexity-entropy-scale causality box. Simulations conducted on synthetic series generated by stochastic, chaotic, and random processes, as well as real world data, are shown to produce unique signatures in this three dimensional representation.

cond-mat.stat-mech

A new look at calendar anomalies: Multifractality and day of the week effect

Stock markets can become inefficient due to calendar anomalies known as day-of-the-week effect. Calendar anomalies are well-known in financial literature, but the phenomena remain to be explored in econophysics. In this paper we use multifractal analysis to evaluate if the temporal dynamics of market returns also exhibits calendar anomalies such as day-of-the-week effects. We apply the multifractal detrended fluctuation analysis (MF-DFA) to daily returns of market indices around the world for each day of the week. Our results indicate that individual days of the week are characterized by distinct multifractal properties. Monday returns tend to exhibit more persistent behavior and richer multifractal structures than other day-resolved returns. Shuffling the series reveals that multifractality arises both from a broad probability density function and from long-term correlations. From the time-dependent multifractal analysis we find that multifractal spectra for Monday returns are much wider than for other days during periods of financial crises. The presence of day-of-the-week effects in multifractal dynamics of market returns motivates further research on calendar anomalies from an econophysics perspective.

q-fin.ST

Distance Metric Learning through Minimization of the Free Energy

Distance metric learning has attracted a lot of interest for solving machine learning and pattern recognition problems over the last decades. In this work we present a simple approach based on concepts from statistical physics to learn optimal distance metric for a given problem. We formulate the task as a typical statistical physics problem: distances between patterns represent constituents of a physical system and the objective function corresponds to energy. Then we express the problem as a minimization of the free energy of a complex system, which is equivalent to distance metric learning. Much like for many problems in physics, we propose an approach based on Metropolis Monte Carlo to find the best distance metric. This provides a natural way to learn the distance metric, where the learning process can be intuitively seen as stretching and rotating the metric space until some heuristic is satisfied. Our proposed method can handle a wide variety of constraints including those with spurious local minima. The approach works surprisingly well with stochastic nearest neighbors from neighborhood component analysis (NCA). Experimental results on artificial and real-world data sets reveal a clear superiority over a number of state-of-the-art distance metric learning methods for nearest neighbors classification.

cs.LG

Search Spaces for Neural Model Training

While larger neural models are pushing the boundaries of what deep learning can do, often more weights are needed to train models rather than to run inference for tasks. This paper seeks to understand this behavior using search spaces -- adding weights creates extra degrees of freedom that form new paths for optimization (or wider search spaces) rendering neural model training more effective. We then show how we can augment search spaces to train sparse models attaining competitive scores across dozens of deep learning workloads. They are also are tolerant of structures targeting current hardware, opening avenues for training and inference acceleration. Our work encourages research to explore beyond massive neural models being used today.

cs.LG

Accelerating Sparse Deep Neural Networks

As neural network model sizes have dramatically increased, so has the interest in various techniques to reduce their parameter counts and accelerate their execution. An active area of research in this field is sparsity - encouraging zero values in parameters that can then be discarded from storage or computations. While most research focuses on high levels of sparsity, there are challenges in universally maintaining model accuracy as well as achieving significant speedups over modern matrix-math hardware. To make sparsity adoption practical, the NVIDIA Ampere GPU architecture introduces sparsity support in its matrix-math units, Tensor Cores. We present the design and behavior of Sparse Tensor Cores, which exploit a 2:4 (50%) sparsity pattern that leads to twice the math throughput of dense matrix units. We also describe a simple workflow for training networks that both satisfy 2:4 sparsity pattern requirements and maintain accuracy, verifying it on a wide range of common tasks and model architectures. This workflow makes it easy to prepare accurate models for efficient deployment on Sparse Tensor Cores.

cs.LG

Nonextensive triplets in stock market indices

Stock market indices are one of the most investigated complex systems in econophysics. Here we extend the existing literature on stock markets in connection with nonextensive statistical mechanics. We explore the nonextensivity of price volatilities for 34 major stock market indices between 2010 and 2019. We discover that stock markets follow nonextensive statistics regarding equilibrium, relaxation and sensitivity. We find nonextensive behavior in stock markets for developed countries, but not for developing countries. Distances between nonextensive triplets suggest that some stock markets might share similar nonextensive dynamics, while others are widely different. The current findings strongly indicate that the stock market represents a system whose physics is properly described by nonextensive statistical mechanics. Our results shed light on the complex nature of stock market indices, and establish another formal link with the nonextensive theory.

q-fin.ST

Paths to collapse for isolated skyrmions in few-monolayer ferromagnetic films

Magnetic skyrmions are topological spin configurations in materials with chiral Dzyaloshinskii-Moriya interaction (DMI), that are potentially useful for storing or processing information. To date, DMI has been found in few bulk materials, but can also be induced in atomically thin magnetic films in contact with surfaces with large spin-orbit interactions. Recent experiments have reported that isolated magnetic skyrmions can be stabilized even near room temperature in few-atom thick magnetic layers sandwiched between materials that provide asymmetric spin-orbit coupling. Here we present the minimum-energy path analysis of three distinct mechanisms for the skyrmion collapse, based on ab initio input and the performed atomic-spin simulations. We focus on the stability of a skyrmion in three atomic layers of Co, either epitaxial on the Pt(111) surface, or within a hybrid multilayer where DMI nontrivially varies per monolayer due to competition between different symmetry-breaking from two sides of the Co film. In laterally finite systems, their constrained geometry causes poor thermal stability of the skyrmion toward collapse at the boundary, which we show to be resolved by designing the high-DMI structure within an extended film with lower or no DMI.

cond-mat.mes-hall