Searcharxiv⌕ Search

arXiv · 2610.12336

asdex: Automatic Sparse Differentiation in JAX

Abstract

Many tasks in scientific computing and machine learning require the Jacobian or Hessian matrix of a function. Automatic differentiation (AD) computes these derivatives to machine precision, but materializing a dense $m \times n$ Jacobian requires $n$ forward-mode or $m$ reverse-mode AD passes, one per column or row. For a large class of functions, each output depends on only a few inputs, making the derivative matrix sparse. Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern. The number of colors, and hence of AD passes, is often independent of the problem dimension: a banded Jacobian with $b$ contiguous bands, for instance, only ever requires $b$ colors, regardless of its size. asdex offers the first standalone ASD toolkit in the popular JAX ecosystem. With asdex.jacobian and asdex.hessian, it provides sparse drop-in replacements for jax.jacobian and jax.hessian.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Adrian Hill, Guillaume Dalle. 2026-10-08. asdex: Automatic Sparse Differentiation in JAX. https://arxiv.org/abs/2610.12336

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture

We present a device-resident, multi-GPU architecture for the large-scale computational verification of Goldbach's conjecture. In prior work, a segmented double-sieve eliminated monolithic VRAM bottlenecks but remained constrained by host-side sieve construction and PCIe transfer latency. Here we migrate the entire segment generation pipeline to the GPU using shared-memory tiling, leaving per-segment host-device communication small and independent of segment size, and schedule work across devices through a lock-free atomic counter, scaling near-linearly to two GPUs (speedup 2.03 at $N = 10^{12}$). On the same hardware, the architecture is $13.1\times$ faster than its host-coupled predecessor at $N = 10^{10}$. It verifies Goldbach's conjecture to $10^{12}$ in 145 seconds of computation and to $10^{13}$ in 3246 seconds on a single NVIDIA RTX 5090, and to $10^{13}$ in 1577 seconds on a second machine with two. Sieve cost grows as $Nπ(\sqrt{N})$, which we analyse and which sets the practical limit on the design and identifies where further optimisation must act. This version corrects two concurrency defects in the implementation described in version~1: the affected timings are re-measured with the corrected release, the four-GPU results are withdrawn, and the range is re-verified with the corrected release. The source code is open-source, the measurement logs and the tests for both defects are archived with it, and the experiments are reproducible on commercially available hardware.

cs.MS↗

Scen-Opt: A Scenario Optimization Toolbox for Data-Driven Convex Programming

The scenario approach is a well-established statistical framework for data-driven decision-making. In particular, in data-driven optimization, the scenario approach unveils how the problem structure governs out-of-sample generalization, and offers a principled basis for assessing and certifying the reliability of the optimal solution as per constraint satisfaction. Despite its strong theoretical development and wide applicability, no software toolbox has been available to date that enables user-friendly, data-driven convex optimization within the scenario-approach framework. In this paper, we introduce Scen-Opt, an open-source software tool that integrates convex programming with data samples while providing statistical guarantees grounded in scenario theory. Scen-Opt is implemented in Python, supporting data-driven linear, quadratic, and semidefinite programming, and offers a Python-based web application with an intuitive and reactive graphical user interface (GUI) built using modern web technologies. Scen-Opt can be used directly through its online interface or installed locally, accommodating both manual input and data-file uploads (CSV, JSON, TXT, TSV, MAT, Excel, NPY, NPZ, Parquet). Built on a Python backend with a modern JavaScript frontend, Scen-Opt offers a highly user-friendly experience and efficient usability across desktops, laptops, tablets, and mobile devices. In this paper, Scen-Opt is applied to a set of representative benchmarks, demonstrating its practical effectiveness for data-driven convex optimization with guaranteed performance.

cs.MS↗

A performance enhancement of the Payne-Hanek range reduction algorithm

Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, many existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs ($|x| \ge 2^{16}$) and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the scaled reduced argument has absolute error below $2^{-110}$ and relative error below $2^{-60}$ for every input, including worst cases. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, achieving higher throughput than existing implementations and lower latency than those returning a double-double reduced argument. The algorithm is currently implemented in the LLVM libc project.

cs.MS↗