SearcharxivSearch

arXiv subjects

Huang Cheng

Publications and source records attributed to Huang Cheng.

3 recordsLinked to original sources

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load. Exhaustive profiling resolves this uncertainty, but measures many configurations that do not affect the final resource allocation. We present FleetSieve, which selects measurements according to their expected effect on a resource-coupled, SLO-aware fleet decision. FleetSieve models capacity and tail latency jointly, compares conservative and optimistic allocations, and stops when their remaining decision gap is below a specified tolerance. On a fixed H100 measurement grid for a 31B-parameter open-weight model, FleetSieve reaches the oracle aggregate decision using 22,200 GPU-seconds, 6.9% less than uniform random profiling in the fixed comparison. Across 200 random reveal orders, its mean saving over random profiling is 5.4% (95% bootstrap CI: 3.5-7.2%). The fixed-comparison saving is 21.5% for Chat, while FleetSieve does not use the fewest GPU-seconds for Code. Joint capacity and tail modeling also avoids selecting a configuration whose 46.4-second completion p99 violates a 30-second SLO. In a 16-GPU allocation, an incorrect sparse-profile decision loses up to 1.93 requests/s and 12.4 percentage points of max-min fulfillment. Boundary repeats and BurstGPT measurements support the observed load-dependent tail-latency mechanism.

cs.LG

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

Prefix caching avoids prefill only when a repeated request returns to a server that still holds the prefix KV. Cache-blind balancing disperses that reuse; fixed affinity preserves it but can overload a server. CacheRoute resolves this tradeoff with a periodic routing plan. It admits high-rate keys to a stable warm set and places their assignments by expected load. Hot keys may use more than one destination, although every key in our primary semi-synthetic aggregate uses exactly one. On Llama-3.3-70B in fp8 across 60 H100 GPUs, CacheRoute sustains 176+/-11 QPS at a 3.5-s p99 SLO, 2.3x the strongest of five baselines. Served KV-cache hit rate rises from 64.1+/-1.3% under cache-blind balancing to 93.2+/-0.5%. A second semi-synthetic aggregate and controlled 8B and burst experiments separate the effects of affinity and placement. Two 32B workloads provide the counterexamples: when affinity recovers too little KV work, its residual load skew reduces or erases the improvement. We therefore recommend gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone.

cs.DC

Multiphoton Rabi Oscillations of Correlated Electrons in Strong Field Nonsequential Double Ionization

With quantum calculations, we have investigated the multiphoton nonsequential double ionization of helium atoms in intense laser fields at ultraviolet wavelengths. Very surprisingly, we find a so-far unobserved double-circle structure in the correlated electron momentum spectra. The double-circle structure essentially reveals multiphoton Rabi oscillations of two electrons, which are strongly supported by the oscillating population of a certain doubly excited state and by the oscillating double ionization signals. This two-electron multiphoton Rabi effect provides profound understandings of electronic correlations and complicated multiphoton phenomena and is expected to be a new tool for broad applications, such as quantum coherent control.

physics.atom-ph