Searcharxiv⌕ Search

arXiv · 2610.03228

Execution-Path Qualification and Realized Costs in Speculative Decoding on Consumer Systems

Abstract

Numerical execution paths can change greedy output; a mismatch alone does not predict the remaining trajectory. On the tested B570 stack, we select a draft-free target-only micro-batch-one predictor using exposed discovery data, then freeze complete generated-ID predictions before reserved live speculation. It matches 72/72 capped sequences on 24 repeat-stable prompts, including all 12 runs of four reference-divergent prompts; the same-set reference matches 60/72. The predictor is sufficient for this path, but it neither restores qualification nor identifies a unique numerical cause. Separate B570 same-state forks test intervention effects; RTX proposal/replay programs test reproduction while sharing numerical routines. We assess cost separately. Newly direct-ID-qualified RTX serialization is slower within its instrumented post-prefill boundary ($0.646\times$ target/serialized latency). At fixed measured progress and target work, 358H remains below parity even with draft and residual costs removed. Qualified M4 Guard beats target-only but loses to fixed speculation. These results give testable diagnoses and engineering decisions about lossless compatibility and complete cost on the measured paths; they establish no portable decoding or control gain.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chengzhan Li. 2026-10-02. Execution-Path Qualification and Realized Costs in Speculative Decoding on Consumer Systems. https://arxiv.org/abs/2610.03228

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration

Evaluating distributed LLM serving designs on physical clusters is costly. Yet existing simulators provide only subsets of the capabilities needed for realistic design exploration: stateful closed-loop execution, timing prediction without profiling target hardware, and direct execution of unmodified serving benchmarks. We present ServeTwin, a closed-loop simulator that couples specification-driven analytical timing with a stateful serving loop. This coupling captures feedback among request completion, scheduling, queue state, and KV-cache evolution, including prefill-decode disaggregation and multi-turn agent workloads. ServeTwin avoids target-hardware operator profiling through iSTAGE, an analytical trace generator that derives per-iteration traces from model and engine specifications while representing ragged batches and discrete engine mechanisms such as CUDA-graph batch padding. It further decomposes execution time into component-owned throughput, scheduler, and runtime costs. This ownership lets unaffected parameters transfer across platforms and confines recalibration to changed hardware or software components. ServeTwin implements a vLLM-compatible interface and runs unmodified serving benchmarks. Against real deployments, it reproduces InferenceX's steady-state throughput-interactivity frontier with a 3.6% mean error and predicts LMBenchmark's multi-turn performance with a 9.9% error while tracking KV-cache evolution. Pre-silicon sweeps reveal that the preferred HBM bandwidth-capacity tradeoff reverses across workload states and that increasing concurrency shifts the bottleneck from memory bandwidth to scheduler and runtime overheads. Together, these capabilities enable practical exploration of distributed LLM serving systems before target hardware is available.

cs.PF↗

A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture

State-vector quantum circuit simulation is memory-bandwidth bound, yet the interaction between memory hierarchy, access pattern, and hardware parallelism remains incompletely characterized. We address this using the Apple M4 Pro Unified Memory Architecture (UMA), where CPU and GPU share identical physical LPDDR5X DRAM ($\sim$224 GB/s STREAM bandwidth for both), eliminating memory-technology and interconnect confounds. Using a thermally isolated, multi-trial methodology across 11 simulation backends on GHZ and QFT circuits from 3 to 30 qubits, we make three central contributions. First, a Roofline analysis confirms all gate implementations have arithmetic intensity $\leq$0.38 FLOP/byte, well below the ridge point for any plausible peak compute on modern hardware, establishing structural memory-boundedness. Second, we identify a reproducible 4.46$\times$ timing discontinuity at the 28$\rightarrow$29 qubit transition, confirmed under thermally isolated conditions and cross-validated across GHZ and QFT circuits; tensordot backends exhibit the full discontinuity while direct-index backends maintain $\sim$2$\times$ per-qubit scaling throughout. Third, despite STREAM predicting only 1.85$\times$ GPU speedup (MLX CPU 119.9 GB/s vs. MLX GPU 221.9 GB/s), all three algorithm classes exceed this prediction: tensordot 3.1--4.1$\times$, flat-index 3.5--5.9$\times$, and direct-index 6--10$\times$, demonstrating that peak streaming bandwidth does not predict simulation speedup for non-contiguous memory access patterns, with the gap widening as access irregularity increases. These findings provide a hardware-characterization framework for quantum simulation workloads on UMA.

cs.PF↗

LLTA: A Simplicity-Oriented Open-Source WCET Analyser

Deriving a safe upper bound on the worst-case execution time (WCET) of a real-time task is essential for hard real-time systems. Many WCET analysers exist, but they are either (i) closed source or (ii) do not provide a WCET for an existing microcontroller. Hence, it is difficult to obtain a trustworthy WCET bound for the binary that is flashed to an available embedded hardware platform and to observe whether the bound is safe. To solve this, we present LLTA, an open-source WCET analyser for Commercial-Off-The-Shelf (COTS) microcontrollers. LLTA is designed based on the following goals: it is simple (G1), runs on real COTS hardware (G2), reports results that are unsound due to programming or device constraints (G3), and is available open-source without licensing complications (G4). LLTA generates WCETs for the ESP32-C6 and the MSP430(FR5994), which are broadly available. This paper systematically lays out our design decisions for LLTA with those goals in mind. To demonstrate its ease of use, we propose scenarios where LLTA can be used for lab exercises of real-time systems.

cs.PF↗