Searcharxiv⌕ Search

arXiv · 2610.08373

ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving

Abstract

Energy-efficient LLM serving requires minimizing serving GPU energy while meeting latency and throughput service-level objectives (SLOs). Attention--FFN disaggregation (AFD) enables separate resource allocation and operating controls for attention and expert computation, but their energy effects remain coupled through the execution pipeline. Realizing its energy-saving potential therefore requires navigating a hierarchical configuration space in which deployment structures constrain admissible controls and shape their end-to-end effects. Finding low-energy configurations that meet SLOs is challenging because physical evaluations are costly and only a small fraction of candidates can be measured. We present Energy-Oriented Configuration Optimization (ECO), which jointly searches deployment structures and their admissible operating controls under a limited measurement budget. ECO constructs a structure-aware energy prior from calibrated stage behavior and pipeline dependencies, then learns residual prediction errors with a Gaussian process. Its cost-aware constrained Bayesian optimization prioritizes measurements according to expected energy improvement while accounting for SLO feasibility, execution success, and evaluation cost, and returns the lowest-energy measured feasible configuration. Across all 16 scenarios on A6000 and A100 with Qwen and DeepSeek, ECO's frozen configurations, evaluated on disjoint requests, reduce serving energy by 40.5\% and increase output token rate by 20.7\% on average relative to baselines while meeting target SLOs. Across the 8 A6000 scenarios, its selected feasible energy averages 33.1\% below generic constrained Bayesian optimization and 25.8\% below genetic search.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zou Qingyun, Bin Gao, Zhuobin Huang, Weng-Fai Wong, Tulika Mitra, Bingsheng He. 2026-10-06. ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving. https://arxiv.org/abs/2610.08373

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Predictive Software Scheduling as an Early-Warning Hint Layer for Optical Engine Thermal Drift in Heterogeneous SoIC Packaging

As semiconductor scaling approaches the A16 / 2 nm node, the integration of co-packaged optics (CPO) through TSMC's Compact Universal Photonic Engine (COUPE) architecture introduces critical thermal-optical coupling challenges. Micro-ring resonators embedded in the Photonic Integrated Circuit (PIC) layer are highly sensitive to temperature, with a wafer-level center wavelength deviation of merely +/-1.7 nm across a 300 mm wafer representing the platform's manufacturing control limit. To address this, we propose XRM-SSD V24, a physics-aware scheduling layer that models inference-load density 20-50 ms before execution and issues early-warning hints to the COUPE bias-control firmware, enabling pre-emptive thermal compensation. Simulation-based validation on a software emulation platform (physical characterization pending TSMC tape-out) over 90,000 inference steps yields a simulator-internal thermal-load correlation of R^2 = 0.9911 across a workload density range of pv24 in [0.9, 2.7] (a 3x span), with wavelength drift below 0.354 nm - equivalent to 21% of the +/-1.7 nm wafer-level wavelength control budget and 71% of the tighter +/-0.5 nm per-channel spectral specification. A full Thermal Resistance Fingerprint characterization further confirms Rth = 0.45 deg C/W, a thermal time constant tau = 80 ms, and a thermo-optic coefficient of 0.0852 nm/deg C across five discrete load states (Idle to Peak). Memory stability is reported as zero leakage in the current simulation run; long-duration soak testing to confirm sustained stability remains future work. We establish a formal domain separation between deterministic software scheduling and continuous physical thermal dynamics, ensuring physics-consistent claims suitable for peer review.

cs.AR↗

An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC

While dependent quantization in H.266/VVC delivers a high compression ratio, its strong serial nature and high complexity result in poor real-time performance, making it difficult to deploy in practical scenarios. To improve the real-time performance of dependent quantization with minimal degradation to its compression performance, we propose a interleaved parallel dependent quantization hardware architecture with low BDBR loss, which achieves four-channel parallel dependent quantization by time-division multiplexing most combinational logic. This architecture adopts the proposed intra-CG context simplification scheme and the encoding scheme that skips decAbslevel during rate estimation. The proposed design incurs a BDBR loss of only 0.42% under the All Intra configuration and 0.38% under the Random Access configuration, respectively. Implemented in Verilog HDL, the dependent quantization hardware architecture achieves quantization speeds of 4K@33.2, 91.4, 331.2, and 454.1 fps at QP = 22, 27, 32, and 37, respectively, when implemented on the Xilinx XCZU19 FPGA. When implemented on ASIC using the TSMC 28nm process standard cell library, the corresponding quantization speeds reach 4K@84.0, 231.0, 837.0, and 1147.7 fps.

cs.AR↗

BenchmarkAnything: Agent-Driven Construction of Simulator-Ready Microarchitecture Benchmarks

The selection of benchmark workloads is of paramount importance in computer architecture, as it establishes the yardstick against which architectural innovations are measured and guided. Yet for decades, the SPEC benchmark suites, comprising merely tens of workloads, have been the de facto standard in academic architectural research, where they are frequently treated as a principal evaluation and optimization target. When a suite this small is relied upon so heavily, it risks architectural overfitting; as our research and prior studies demonstrate, an overly narrow focus can mislead design decisions by overvaluing certain innovations, producing cores that excel on SPEC benchmarks yet underperform on broader, realistic workloads. To mitigate this overfitting, adopting a large, comprehensive benchmark suite is the natural solution. However, the immense engineering effort required to strip software into the clean, interference-free binary executables demanded by simulators often makes this highly impractical. In this work, we demonstrate that AI agents provide an elegant solution to this challenge. Rather than manually curating yet another static benchmark suite, we introduce an agent-driven workflow capable of autonomously transforming arbitrary open-source repositories into simulator-ready executables. This automated approach makes workload collection highly scalable, allowing us to rapidly harvest hundreds of diverse applications from public repositories into our benchmark suite. Through a comparative analysis of our agent-generated suite against SPEC, we show that it not only achieves higher-fidelity performance assessments but also uncovers novel architectural insights that traditional, static suites fail to expose.

cs.AR↗