SearcharxivSearch

arXiv subjects

Zhuang Wang

Publications and source records attributed to Zhuang Wang.

At least 19 recordsLinked to original sources

An Event is Worth One Token: Event Tokenization for Industrial-scale LLM Recommendation

LLM-based recommendation has scaled along model capacity and sequence length, yet each position encodes only text, semantic IDs, or a few categorical features, discarding rich user, item, context, and outcome signals available at each event. Under autoregressive modeling, this yields weak queries at each position and, since each position becomes context for the next, the degradation compounds across the sequence. We propose an event-centric paradigm that represents each interaction by its full temporal snapshot, and identify a new scaling dimension we term snapshot resolution: the amount of information encoded per event. To efficiently scale snapshot resolution, we introduce AMBER (Autoregressive Modeling via Bottlenecked Event Representation), which compresses each temporal snapshot into a compact Event Token, a new LLM input modality. The representation is learned end-to-end, while Event Tokens are pre-computed and cached for serving, decoupling snapshot resolution from real-time serving compute. On industrial-scale ranking and retrieval benchmarks, AMBER advances the compute-quality Pareto frontier relative to alternative recommendation paradigms. At sufficient capacity, a single unified tokenizer even outperforms dedicated per-entity tokenizers, demonstrating positive transfer across structurally different entity types. AMBER's Event Tokens also transfer across model architectures: when integrated into a heavily optimized non-LLM ranker as serving-time historical features, they yield statistically significant improvements. Further scaling Event Tokenizer capacity provides additional improvements.

cs.IR

SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training

In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure-localization framework built on one design principle: identify outliers through strict-majority consensus among equivalent replicas. SCOUT aligns replica progress, timing, and numerical evidence, then uses its Consensus Collective Communication (C3) abstraction to identify ranks whose compact signatures disagree with their peers. An out-of-band CPU observer remains responsive when training hangs, whereas in-situ replay exercises recurring stragglers and silent data corruption (SDC) beside the live job with its model state, kernels, allocations, communication path, and thermal and memory pressure present. Collective fingerprints expose rank-local protocol divergence. Clean replay coverage certifies checkpoint numerical integrity, preventing recovery from selecting state corrupted by SDC. SCOUT integrates with PyTorch, TorchTitan, Megatron-Core, and DeepSpeed without training-loop or framework-source modifications. SCOUT is open source at https://github.com/LMResiliency/lm-resiliency.

cs.DC

Equivalent characterizations of John and uniform domains in doubling metric spaces

In this paper, we characterize John and uniform domains in doubling metric spaces. Specifically, we show that a locally quasiconvex domain in a doubling metric space is length John if and only if it is diameter John. For uniform domains, we prove that a domain in a doubling metric space is length uniform if and only if it is diameter uniform (or distance uniform) and locally quasiconvex. Moreover, in a doubling length metric space, we refine this result by showing that a domain is length uniform (resp. John) if and only if it is diameter uniform (resp. John).

math.CV

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers

LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by dynamically selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, pairwise ranking, and confidence-based approaches. Our extensive study of 21 routing methods across five benchmarks reveals a consistent phenomenon that we call the routing plateau: many methods, including kNN, achieve very similar accuracy and converge to a narrow performance range that remains far below the oracle router. Our investigation shows that the plateau is largely caused by a predictability bottleneck: current routers mainly learn global averaged model-performance trends rather than fine-grained query-specific routing signals. As a result, they solve overlapping easy queries but collectively fail on hard queries that require instance-specific routing decisions. We further study how to move beyond the plateau and find that larger training datasets, stronger encoders, and end-to-end fine-tuning can further improve routing accuracy. These findings characterize the common limits of current routing methods and provide insights and actionable directions for the community to build more effective routing systems.

cs.LG

Boxing inequalities for relative fractional perimeter and fractional Poincar\'e-type inequalities on John domains with the BBM factor

For $0<\delta,\tau<1$ and $1\le s\le \frac{n}{n-\delta}$, we prove that for a given $s$-John domain $\Omega\subset \mathbb{R}^n$, the following Boxing inequality holds for every Lebesgue measurable set $U\subset\Omega$ with $|U|/|\Omega|\le\gamma<1$: \[ \mathcal{H}^{s(n-\delta)}_{\infty}(U\setminus\mathcal{N}_U)\le C(1-\delta)\int_\Omega\int_{|x-y|<\tau\operatorname{dist}(y,\partial\Omega)}\frac{|\chi_U(x)-\chi_U(y)|}{|x-y|^{n+\delta}}\,dx\,dy, \] where $\mathcal{H}^{s(n-\delta)}_{\infty}(U)$ denotes the $s(n-\delta)$-dimensional Hausdorff content of $U$, $\mathcal{N}_U$ is a set of Lebesgue measure zero and the constant $C$ depends only on $n,\tau,s,\gamma$, the John constant and the diameter of $\Omega$. Moreover, we establish the functional formulation of the above Boxing inequality and discuss the equivalence between these two formulations. Based on the Boxing inequality, we prove the fractional Poincar\'e--Wirtinger trace inequality on $s$-John domains, of which the fractional Sobolev--Poincar\'e inequality and fractional Hardy-type inequality are special cases. Notably, we prove all of the aforementioned inequalities with the Bourgain--Brezis--Mironescu (BBM) factor $1-\delta$. Furthermore, with the aid of the Bourgain--Brezis--Mironescu formula, we recover the Poincar\'e--Wirtinger trace inequality. Finally, by showing that, under the separation property, any domain supporting the Boxing inequality is necessarily a John domain, we conclude that the John domain condition is essentially sharp for the above inequalities. All the above inequalities with the BBM factor are new even for Lipschitz domains.

math.FA

Geometric properties of Euclidean domains supporting trace inequalities

We investigate the geometric behavior of $\tau(E)$ for bounded finite-perimeter sets $E \subset \mathbb R^n$, where $\tau(E)$ is the trace constant introduced by Figalli--Maggi--Pratelli [Invent. Math. 2010]. This quantity is a key ingredient in proving a quantitative isoperimetric inequality with the optimal exponent. We first show that for every $\epsilon>0$ one can find a bounded open set $\Omega \subset \mathbb R^n$ that is very close to the unit ball $\mathbb B^n$ in the sense that $$ \tau(\mathbb B^n)>\tau(\Omega)>\tau(\mathbb B^n)-\epsilon \quad \text{and} \quad P(\Omega \Delta \mathbb B^n)\le C(n)\epsilon, $$ while at the same time the complement of $\Omega$ has infinitely many connected components. Thus, $\tau(\Omega)$ can be made arbitrarily close to $\tau(\mathbb B^n)$ even when $\Omega$ has highly intricate geometry. We then establish, under a mild additional hypothesis, the equivalence between a condition formulated in terms of $\tau$ and two classical criteria from the literature for open sets that admit trace inequalities. As a consequence, we obtain the John-type characterization of domains that support a trace inequality, assuming the ball separation property.

math.FA

UCCL-Zip: Lossless Compression Supercharged GPU Communication

The rapid growth of large language models (LLMs) has made GPU communication a critical bottleneck. While prior work reduces communication volume via quantization or lossy compression, these approaches introduce numerical errors that can degrade convergence, accuracy, and stability. We present UCCL-Zip, a unified design that integrates lossless compression directly into GPU communication primitives. UCCL-Zip supports both point-to-point (P2P) and collective communication without modifying user-facing APIs or compromising numerical correctness. For P2P communication, Uzip-P2P employs a split-send pipeline that exposes transmissible data early and overlaps compression with communication, while preserving high GPU efficiency by operating on large data blocks. For collective communication, Uzip-NCCL integrates compression into NCCL's persistent kernel model via fused execution, eliminating redundant memory traffic and kernel launches. In real workloads, UCCL-Zip accelerates RL weight synchronization by up to 47.5% and reduces vLLM end-to-end inference latency by up to 10%, all without application changes.

cs.DC

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads

Low-Rank Adaptation (LoRA) is now the dominant method for parameter-efficient fine-tuning of large language models, but achieving a high-quality adapter often requires systematic hyperparameter tuning because LoRA performance is highly sensitive to configuration choices. In practice, this leads to many concurrent LoRA jobs, often spanning heterogeneous tasks in multi-tenant environments. Existing systems largely handle these jobs independently, which both wastes computation on weak candidates and leaves GPUs underutilized. We present ALTO (Adaptive LoRA Tuning and Orchestration), a co-designed training system that accelerates LoRA hyperparameter tuning while enabling efficient cluster sharing across heterogeneous tasks. The central insight behind ALTO is that when multiple tuning jobs run concurrently over a shared frozen backbone, they expose optimization opportunities that single-job designs cannot exploit. Building on this, ALTO monitors loss trajectories to terminate unpromising configurations early, uses fused grouped GEMM together with a new rank-local adapter parallelism to co-locate surviving adapters and reclaim freed GPU capacity, and combines intra-task and inter-task scheduling to improve multi-task placement by leveraging the predictable duration of LoRA jobs. Extensive evaluation shows that ALTO achieves up to $13.8\times$ speedup over state-of-the-art without sacrificing adapter quality.

cs.LG

ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management

LLM-based multi-agent simulations are increasingly adopted across application domains, but remain difficult to scale due to GPU memory pressure. Each agent maintains private GPU-resident states, including models, prefix caches, and adapters, which quickly exhaust device memory as the agent count grows. We identify two key properties of these workloads: sparse agent activation and an estimable agent invocation order. Based on an analysis of representative workload classes, we introduce invocation distance, a unified abstraction that estimates the relative order in which agents will issue future LLM requests. Leveraging this abstraction, we present ScaleSim, a memory-efficient LLM serving system for large-scale multi-agent simulations. ScaleSim enables proactive prefetching and priority-based eviction, supports diverse agent-specific memory through a modular interface, and achieves up to 1.74x speedup over SGLang on simulation benchmarks.

cs.AI

Quantitative correspondence between quasi-symmetric mappings on complete metric spaces and rough quasi-isometric mappings on their hyperbolic fillings

In this paper, we establish a quantitative correspondence between power quasi-symmetric mappings on complete metric spaces and rough quasi-isometric mappings on their hyperbolic fillings. In particular, we prove that the exponents in the power quasi-symmetric mappings coincide with the coefficients in the rough quasi-isometric mappings. This shows that the obtained correspondence is both sharp and consistent. In this way, we generalize the corresponding result by Bj\"orn, Bj\"orn, Gill, and Shanmugalingam (J. Reine Angew. Math., 2017) from the setting of rooted trees to that of hyperbolic fillings.

math.CV

PLoRA: Efficient Concurrent LoRA Training for Large Language Models

Low-Rank Adaptation (LoRA) has gained popularity as a fine-tuning approach for Large Language Models (LLMs) due to its low resource requirements and good performance. While numerous studies have investigated ways to improve LoRA serving efficiency by serving multiple LoRAs concurrently, existing methods assume that a wide range of LoRA adapters are available for serving. In our work, we conduct extensive empirical studies to show that current LoRA training paradigms do not efficiently utilize hardware resources and incur high overhead to obtain a performant LoRA adapter. Leveraging these insights, we propose PLoRA, which automatically orchestrates concurrent LoRA fine-tuning jobs under given hardware and model constraints and develops performant kernels to improve training efficiency. Across a range of LLMs and LoRA configurations, PLoRA improves training throughput by up to 12.8x and reduces the overall fine-tuning makespan by up to 7.52x compared to existing approaches.

cs.LG

Marconi: Prefix Caching for the Era of Hybrid LLMs

Hybrid models that combine the language modeling capabilities of Attention layers with the efficiency of Recurrent layers (e.g., State Space Models) have gained traction in practically supporting long contexts in Large Language Model serving. Yet, the unique properties of these models complicate the usage of complementary efficiency optimizations such as prefix caching that skip redundant computations across requests. Most notably, their use of in-place state updates for recurrent layers precludes rolling back cache entries for partial sequence overlaps, and instead mandates only exact-match cache hits; the effect is a deluge of (large) cache entries per sequence, most of which yield minimal reuse opportunities. We present Marconi, the first system that supports efficient prefix caching with Hybrid LLMs. Key to Marconi are its novel admission and eviction policies that more judiciously assess potential cache entries based not only on recency, but also on (1) forecasts of their reuse likelihood across a taxonomy of different hit scenarios, and (2) the compute savings that hits deliver relative to memory footprints. Across diverse workloads and Hybrid models, Marconi achieves up to 34.4$\times$ higher token hit rates (71.1% or 617 ms lower TTFT) compared to state-of-the-art prefix caching systems.

cs.DC

Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models

Sparsely-activated Mixture-of-Experts (MoE) architecture has increasingly been adopted to further scale large language models (LLMs). However, frequent failures still pose significant challenges as training scales. The cost of even a single failure is significant, as all GPUs need to idle wait until the failure is resolved, potentially losing considerable training progress as training has to restart from checkpoints. This problem is exacerbated by the growing use of spot instances on public clouds for model training, which despite offering substantial cost savings, introduce frequent preemptions-essentially failures that regularly occur throughout the training process. Existing solutions for efficient fault-tolerant training either lack elasticity or rely on building resiliency into pipeline parallelism, which cannot be applied to MoE models due to the expert parallelism strategy adopted by the MoE architecture. We present Lazarus, a system for resilient and elastic training of MoE models. Lazarus adaptively allocates expert replicas to address the inherent imbalance in expert workload and speeds up training, while a provably optimal expert placement algorithm is developed to maximize the probability of recovery upon failures. Through adaptive expert placement and a flexible token dispatcher, Lazarus can also fully utilize all available nodes after failures, leaving no GPU idle. Our evaluation shows that Lazarus outperforms existing MoE training systems by up to 5.7x under frequent node failures and 3.4x on a real spot instance trace.

cs.DC

Empowering Distributed Training with Sparsity-driven Data Synchronization

Distributed training is the de facto standard to scale up the training of deep learning models with multiple GPUs. Its performance bottleneck lies in communications for gradient synchronization. Although high tensor sparsity is widely observed, the optimal communication scheme to fully leverage sparsity is still missing. This paper aims to bridge this gap. We first analyze the characteristics of sparse tensors in popular models to understand the fundamentals of sparsity. We then systematically explore the design space of communication schemes for sparse tensors and find the optimal ones. These findings give a new understanding and inspire us to develop a holistic gradient synchronization system called Zen for sparse tensors. We demonstrate that Zen can achieve up to 5.09x speedup in communication time and up to $2.48\times$ speedup in training throughput compared to the state-of-the-art methods.

cs.LG

Locally biH\"{o}lder continuous mappings and their induced embeddings between Besov spaces

In this paper, we introduce a class of homeomorphisms between metric spaces, which are locally biH\"{o}lder continuous mappings. Then an embedding result between Besov spaces induced by locally biH\"{o}lder continuous mappings between Ahlfors regular spaces is established, which extends the corresponding result of Bj\"{o}rn-Bj\"{o}rn-Gill-Shanmugalingam (J. Reine Angew. Math. 725: 63-114, 2017). Furthermore, an example is constructed to show that our embedding result is more general. We also introduce a geometric condition, named as uniform boundedness, to characterize when a quasisymmetric mapping between uniformly perfect spaces is locally biH\"{o}lder continuous.

math.FA

ByteComp: Revisiting Gradient Compression in Distributed Training

Gradient compression (GC) is a promising approach to addressing the communication bottleneck in distributed deep learning (DDL). However, it is challenging to find the optimal compression strategy for applying GC to DDL because of the intricate interactions among tensors. To fully unleash the benefits of GC, two questions must be addressed: 1) How to express all compression strategies and the corresponding interactions among tensors of any DDL training job? 2) How to quickly select a near-optimal compression strategy? In this paper, we propose ByteComp to answer these questions. It first designs a decision tree abstraction to express all the compression strategies and develops empirical models to timeline tensor computation, communication, and compression to enable ByteComp to derive the intricate interactions among tensors. It then designs a compression decision algorithm that analyzes tensor interactions to eliminate and prioritize strategies and optimally offloads compression to CPUs. Experimental evaluations show that ByteComp can improve the training throughput over the start-of-the-art compression-enabled system by up to 77% for representative DDL training jobs. Moreover, the computational time needed to select the compression strategy is measured in milliseconds, and the selected strategy is only a few percent from optimal.

cs.LG

Borderline case of traces and extensions for weighted Sobolev spaces

In this paper, we study the traces and the extensions for weighted Sobolev spaces on upper half spaces when the weights reach to the borderline cases. We first give a full characterization of the existence of trace spaces for these weighted Sobolev spaces, and then study the trace parts and the extension parts between the weighted Sobolev spaces and a new kind of Besov-type spaces (on hyperplanes) which are defined by using integral averages over selected layers of dyadic cubes.

math.FA

$p$-harmonic mappings between metric spaces

In this paper, we solve the Dirichlet problem for Sobolev maps between singular metric spaces that extends the corresponding result of Guo and Wenger [Comm. Anal. Geom. 2020]. The main new ingredient in our proofs is a suitable extension of the theory of trace for metric valued Sobolev maps developed by Korevaar and Schoen [Comm. Anal. Geom. 1993]. We also develop a theory of trace in the borderline case, which investigates a sharp condition to characterize the existence of traces.

math.AP