SearcharxivSearch

arXiv subjects

Gil Tabak

Publications and source records attributed to Gil Tabak.

8 recordsLinked to original sources

Finer is Better (with the Right Scaling)

Microscaling is a critical technique for preserving the quality of Large Language Models (LLMs) quantized to ultra-low precision formats. Intuitively, finer block sizes should yield lower quantization error; however, a paradox recently identified by Fasoli et al. (2026) demonstrates that standard abs-max scaling can actually result in degraded model quality as block sizes shrink. In this work, we investigate the underlying mechanics of this phenomenon. We demonstrate that this degradation is not an inherent limitation of finer granularity, but is primarily driven by how elements in smaller blocks statistically cluster closer to their local block maximum, interacting poorly with the coarse subnormal E4M3 values used as scaling factors. Specifically, we show that i) preventing the scaling factor from underflowing to zero mitigates errors caused by extreme underflow, ii) targeted algorithmic interventions like the 4-over-6 methodology that give more flexibility to the choice of scaling factor resolve the paradox for larger values, and iii) a brute-force search establishes an optimal baseline, confirming that the theoretical Mean Squared Error (MSE) strictly improves with finer block sizes. Ultimately, our findings highlight a critical insight for hardware-software co-design: the block-size paradox is partially an artifact of naive scale selection. While using hierarchical scaling factors or wider formats like UE5M3 interchangeably resolves much of the quality loss, we found the 4-over-6 scale selection heuristic can even further improve quality, especially for very small block sizes. Consequently, maximizing the performance of next-generation ML accelerators will require treating silicon format specifications and software scaling algorithms as tightly coupled design choices.

cs.LG

EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration

While Large Language Models (LLMs) have become highly influential, their enormous scale presents significant deployment challenges. Efficiently serving these models typically requires distributing them across numerous accelerator devices, which introduces substantial performance overhead from inter-device communication (collectives). While model quantization has been widely adopted to reduce the memory and compute requirements of LLM weights and activations with minimal quality impact, applying quantization directly to collectives like AllReduce is inherently difficult due to the inter-device summation involved, which can lead to numerical instability or significant error accumulation. In this work, we present a native dynamic block-wise efficient quantized AllReduce within the XLA compiler for TPUs (EQuARX). By using TPU-friendly quantization and deep pipelining of communication and compute, EQuARX with int8 precision achieves a 1.8X speedup over baseline BF16 AllReduce across various network topologies. Furthermore, EQuARX accelerates the prefill stage of Gemma 3 27B by 1.25X and Gemma 3 12B by 1.1X, respectively, with small to negligible impact on quality.

cs.LG

A Metric Driven Approach to Mixed Precision Training

As deep learning methodologies have developed, it has been generally agreed that increasing neural network size improves model quality. However, this is at the expense of memory and compute requirements, which also need to be increased. Various efficiency techniques have been proposed to rein in hardware costs, one being the use of low precision numerics. Recent accelerators have introduced several different 8-bit data types to help accommodate DNNs in terms of numerics. In this paper, we identify a metric driven methodology to aid in the choice of numerics. We demonstrate how such a methodology can help scale training of a language representation model. The technique can be generalized to other model architectures.

cs.LG

Factorization of Linear Quantum Systems with Delayed Feedback

We consider the transfer functions describing the input-output relation for a class of linear open quantum systems involving feedback with nonzero time delays. We show how such transfer functions can be factorized into a product of terms which are transfer functions of canonical physically realizable components. We prove under certain conditions that this product converges, and can be approximated on compact sets. Thus our factorization can be interpreted as a (possibly infinite) cascade. Our result extends past work where linear open quantum systems with a state-space realization have been shown to have a pure cascade realization [Nurdin, H. I., Grivopoulos, S., & Petersen, I. R. (2016). The transfer function of generic linear quantum stochastic systems has a pure cascade realization. Automatica, 69, 324-333.]. The functions we consider are inherently non-Markovian, which is why in our case the resulting product may have infinitely many terms.

quant-ph

Correcting Nuisance Variation using Wasserstein Distance

Profiling cellular phenotypes from microscopic imaging can provide meaningful biological information resulting from various factors affecting the cells. One motivating application is drug development: morphological cell features can be captured from images, from which similarities between different drug compounds applied at different doses can be quantified. The general approach is to find a function mapping the images to an embedding space of manageable dimensionality whose geometry captures relevant features of the input images. An important known issue for such methods is separating relevant biological signal from nuisance variation. For example, the embedding vectors tend to be more correlated for cells that were cultured and imaged during the same week than for those from different weeks, despite having identical drug compounds applied in both cases. In this case, the particular batch in which a set of experiments were conducted constitutes the domain of the data; an ideal set of image embeddings should contain only the relevant biological information (e.g. drug effects). We develop a general framework for adjusting the image embeddings in order to `forget' domain-specific information while preserving relevant biological information. To achieve this, we minimize a loss function based on distances between marginal distributions (such as the Wasserstein distance) of embeddings across domains for each replicated treatment. For the dataset we present results with, the only replicated treatment happens to be the negative control treatment, for which we do not expect any treatment-induced cell morphology changes. We find that for our transformed embeddings (i) the underlying geometric structure is not only preserved but the embeddings also carry improved biological signal; and (ii) less domain-specific information is present.

stat.ML

Trapped Modes in Linear Quantum Stochastic Networks with Delays

Networks of open quantum systems with feedback have become an active area of research for applications such as quantum control, quantum communication and coherent information processing. A canonical formalism for the interconnection of open quantum systems using quantum stochastic differential equations (QSDEs) has been developed by Gough, James and co-workers and has been used to develop practical modeling approaches for complex quantum optical, microwave and optomechanical circuits/networks. In this paper we fill a significant gap in existing methodology by showing how trapped modes resulting from feedback via coupled channels with finite propagation delays can be identified systematically in a given passive linear network. Our method is based on the Blaschke-Potapov multiplicative factorization theorem for inner matrix-valued functions, which has been applied in the past to analog electronic networks. Our results provide a basis for extending the Quantum Hardware Description Language (QHDL) framework for automated quantum network model construction (Tezak \textit{et al.} in Philos. Trans. R. Soc. A, Math. Phys. Eng. Sci. 370(1979):5270-5290, to efficiently treat scenarios in which each interconnection of components has an associated signal propagation time delay.

quant-ph

Systematic Stochastic Reduction of Inertial Fluid-Structure Interactions subject to Thermal Fluctuations

We investigate the dynamics of elastic microstructures within a fluid that are subjected to thermal fluctuations. We perform analysis to obtain systematically simplified descriptions of the mechanics in the limiting regimes when (i) the coupling forces that transfer momentum between the fluid and microstructures is strong, (ii) the mass of the microstructures is small relative to the displaced mass of the fluid, and (iii) the response to stresses results in hydrodynamics that relax rapidly to a quasi-steady-state relative to the motions of the microstructure. We derive effective equations using a singular perturbation analysis of the Backward Kolmogorov equations of the stochastic process. Our continuum mechanics description is based on the Stochastic Eulerian Lagrangian Method (SELM) which provides a framework for approximation of the fluid-structure interactions when subject to thermal fluctuations.

cond-mat.soft

The First Three Rungs of the Cosmological Distance Ladder

It is straightforward to determine the size of the Earth and the distance to the Moon without making use of a telescope. The methods have been known since the 3rd century BC. However, few amateur or professional astronomers have worked this out from data they themselves have taken. Here we use a gnomon to determine the latitude and longitude of South Bend, Indiana, and College Station, Texas, and determine a value of the radius of the Earth of 6290 km, only 1.4 percent smaller than the true value. We use the method of Aristarchus and the size of the Earth's shadow during the lunar eclipse of 2011 June 15 to derive an estimate of the distance to the Moon (62.3 R_Earth), some 3.3 percent greater than the true mean value. We use measurements of the angular motion of the Moon against the background stars over the course of two nights, using a simple cross staff device, to estimate the Moon's distance at perigee and apogee. Finally, we use simultaneous CCD observations of asteroid 1996 HW1 obtained with small telescopes in Socorro, New Mexico, and Ojai, California, to derive a value of the Astronomical Unit of (1.59 +/- 0.19) X 10^8 km, about 6 percent too large. The data and methods presented here can easily become part of a beginning astronomy lab class.

astro-ph.IM