SearcharxivSearch

arXiv subjects

Jerry Wu

Publications and source records attributed to Jerry Wu.

10 recordsLinked to original sources

Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models

Mixture-of-Experts (MoE) models route each token to a few of many expert networks, and that routing is data-dependent in a way standard inference optimizations do not expect. This paper measures what three of them actually deliver on OLMoE-1B-7B, DeepSeek-V2-Lite, and Qwen3-30B-A3B. Fused Triton kernels reach 5.6x to 9.0x in isolation but 0.999x end to end against a measured 1.07x ceiling, because the model spends its time waiting on roughly a thousand kernel launches per forward pass rather than on the arithmetic those kernels improve. INT4 quantization changes on average 0.53 of the eight selected experts per token position, yet replaying exactly those changed routes through full-precision weights reproduces only 2.7% of the quality loss, which makes the experts substitutable rather than specialized. Removing all 23 graph breaks from PyTorch's compiler, the step prior work treats as the structural fix, makes the model three times slower. A fourth result ties the three together: leaving the routers in FP16 lowers drift by 20% while raising loss, so routing fidelity and output quality are separable objectives. Every number recomputes from committed per-token route dumps.

cs.PF

Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving

As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting factor. Our study shows that using dynamic voltage and frequency scaling (DVFS) to exploit the service-level objective (SLO) slacks is a promising way to improve NPU energy efficiency for LLM services. And as tensor operators in LLMs exhibit diverse bottlenecks across NPU components, it is desirable to configure the frequency separately for each component to maximize their energy efficiency. In this paper, we develop eNPU that enables hardware and software support for spatially fine-grained, component-level DVFS on NPUs. eNPU refactors the NPU core pipeline to partition components into separate V/$f$ domains. It introduces lightweight cross-domain communication mechanisms to mitigate synchronization overheads across components, and extends the NPU ISA for sub-$\mu$s DVFS control. eNPU uses a compiler-driven two-level greedy search to co-optimize instruction scheduling and per-component V/$f$ selection under SLO constraints. We implement eNPU's pipeline design on an open-source NPU core to verify its functionality and evaluate the energy savings with a production-level NPU simulator with various LLMs using production traces. eNPU reduces energy consumption of LLM services by 25.8%--35.2% with 3.45% area overhead on a TPUv4 chip, while preserving strict SLO guarantees.

cs.AR

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-category failure taxonomy validated by two independent human annotators (Cohen's $\kappa$ = 0.906), we classify 30,000 chain-of-thought outputs from five instruction-tuned LLMs (3B--14B parameters) across three quantization precisions (FP32, FP16, NF4) and four reasoning benchmarks. We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable reasoning) shows a significant size-dependent shift under NF4, dropping sharply for the two smallest models tested but remaining invariant for models at 12B parameters and above. This effect is also benchmark-specific: GSM8K is categorically immune while LogiQA and ARC-Challenge show the largest shifts. Furthermore, under NF4, Shortcut Collapse rises from 44% to 78% of wrong-answer failures in LLaMA 3.2-3B while Confidence Snowballing collapses from 15.8% to near zero, a qualitative shift invisible to accuracy metrics. Finally, we show Hollow Convergence cannot be reliably detected from surface-level text features (best F1 = 0.53), establishing it as a deployment-relevant failure mode that standard evaluation pipelines cannot catch.

cs.CL

Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior

Large automatic speech recognition (ASR) models such as Whisper must be deployed across hardware with widely varying memory and inference-speed constraints. We present a compression framework that jointly parametrizes Whisper deployment along \emph{six} dimensions: model size $x_N$, temporal resolution $x_T$, encoder token stride $x_V$, low-rank adaptation capacity $x_R$, weight precision $x_Q$ and sparsity pattern $x_P$. All axes are jointly optimized against three deployment objectives (word error rate, inference FLOPs, and memory footprint) using a non-dominated sorting genetic evolutionary search (NSGA). Across 50 of the 1,680 candidate configurations evaluated, we measure the marginal effect of each axis on the three objectives and identify compression combinations that dominate naive single-axis scaling, and report a consistent negative result: 1:4 structured sparsity fails to recover acceptable accuracy under any tested recovery budget. We report real, measured memory and accuracy figures for genuinely quantized deployment artifacts, and provide a lookup table mapping deployment scenarios (cloud, server, edge, ultra-constrained) to specific axis configurations with their measured accuracy/memory/compute trade-offs

cs.SD

Thermodynamically Stable Phases of Asymptotically Flat Lovelock Black Holes

We present the first examples of phase transitions in asymptotically flat black hole solutions. We analyze the thermodynamic properties of black holes in order $N\ge 3$ Lovelock gravity, with zero cosmological constant. We find a new type of "inverted" swallowtail indicative of stable temperature regions for an otherwise unstable neutral black hole, and demonstrate multiple such stable phases can exist and coexist at multi-critical points. We also find that for charged black holes, ordinary swallowtails can exist on the stable Gibbs free energy branch, allowing for multiple first order phase transitions as seen for AdS black holes. A triple point for $N=5$ and a quadruple point for $N=7$ are presented explicitly. We investigate changes in the Gibbs free energy as the lowest order Lovelock constant is varied, and draw comparisons to pressure changes for AdS black hole systems.

hep-th

Multicritical Phase Transitions in Lovelock AdS Black Holes

We demonstrate that black holes in order $N\ge 4$ Lovelock gravity can exhibit multicritical phase behaviour. We show an explicit example of a quadruple point in $d=10$ fourth-order Lovelock gravity and a quintuple point in $d=14$ sixth-order Lovelock gravity. We also demonstrate that multi-criticality can be realized for uncharged, non-rotating black holes by highlighting a new type of multi-critical point between black holes and thermal radiation. We discuss the methodology used and make comparisons to other black hole multi-critical points in terms of the Gibbs phase rule.

hep-th

Packing Densities of Delzant and Semitoric Polygons

Exploiting the relationship between 4-dimensional toric and semitoric integrable systems with Delzant and semitoric polygons, respectively, we develop techniques to compute certain equivariant packing densities and equivariant capacities of these systems by working exclusively with the polygons. This expands on results of Pelayo and Pelayo-Schmidt. We compute the densities of several important examples and we also use our techniques to solve the equivariant semitoric perfect packing problem, i.e., we list all semitoric polygons for which the associated semitoric system admits an equivariant packing which fills all but a set of measure zero of the manifold. This paper also serves as a concise and accessible introduction to Delzant and semitoric polygons in dimension four.

math.SG

Multicritical Phase Transitions in Multiply Rotating Black Holes

We show that multi-critical points in which more than three phases coalesce are present in multiply rotating Kerr-AdS black holes in $d$-dimensions. We explicitly present a quadruple point for a triply rotating black hole in $d=8$ and a quintuple point for a quadruply rotating black hole in $d=10$. The maximal number of distinct phases $n$ is one larger than the maximal number of independent rotations, and we outline a method for obtaining the associated $n$-tuple point. Situations also exist where more than three phases merge at sub-maximal multi-critical points. Our results show that multi-critical points in black hole thermodynamics are more common than previously thought, with systems potentially supporting many phases as long as a sufficient number of thermodynamic variables are present.

gr-qc

Multi-critical Points in Black Hole Phase Transitions

We present the first examples in black hole thermodynamics of multicritical phase transitions, in which more than three distinct black hole phases merge at a critical point. Working in the context of non-linear electrodynamics, we explicitly present examples of black hole quadruple and quintuple points, and demonstrate how $n$-tuple critical points can be obtained. Our results indicate that black holes can have multiple phases beyond the three types observed so far, resembling the behaviour of multicomponent chemical systems. We discuss the interpretation of our results in the context of the Gibbs Phase Rule.

hep-th

On BBW parabolics for simple classical Lie superalgebras

In this paper the authors introduce a class of parabolic subalgebras for classical simple Lie superalgebras associated to the detecting subalgebras introduced by Boe, Kujawa and Nakano. These parabolic subalgebras are shown to have good cohomological properties governed by the Bott-Borel-Weil theorem involving the zero component of the Lie superalgebra in conjunction with the odd roots. These results are later used to verify an open conjecture given by Boe, Kujawa and Nakano pertaining to the equality of various support varieties.

math.RT