Search arXiv⌕ Search

arXiv · 2610.08373

ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving

Abstract

Energy-efficient LLM serving requires minimizing serving GPU energy while meeting latency and throughput service-level objectives (SLOs). Attention--FFN disaggregation (AFD) enables separate resource allocation and operating controls for attention and expert computation, but their energy effects remain coupled through the execution pipeline. Realizing its energy-saving potential therefore requires navigating a hierarchical configuration space in which deployment structures constrain admissible controls and shape their end-to-end effects. Finding low-energy configurations that meet SLOs is challenging because physical evaluations are costly and only a small fraction of candidates can be measured. We present Energy-Oriented Configuration Optimization (ECO), which jointly searches deployment structures and their admissible operating controls under a limited measurement budget. ECO constructs a structure-aware energy prior from calibrated stage behavior and pipeline dependencies, then learns residual prediction errors with a Gaussian process. Its cost-aware constrained Bayesian optimization prioritizes measurements according to expected energy improvement while accounting for SLO feasibility, execution success, and evaluation cost, and returns the lowest-energy measured feasible configuration. Across all 16 scenarios on A6000 and A100 with Qwen and DeepSeek, ECO's frozen configurations, evaluated on disjoint requests, reduce serving energy by 40.5\% and increase output token rate by 20.7\% on average relative to baselines while meeting target SLOs. Across the 8 A6000 scenarios, its selected feasible energy averages 33.1\% below generic constrained Bayesian optimization and 25.8\% below genetic search.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zou Qingyun, Bin Gao, Zhuobin Huang, Weng-Fai Wong, Tulika Mitra, Bingsheng He. 2026-10-06. ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving. https://arxiv.org/abs/2610.08373

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Adaptive Heterogeneous Architecture for High-Ratio, High-Throughput Lossless Compression

Modern lossless compression enforces an acute dichotomy: industrial streaming codecs (e.g., Zstandard) prioritize throughput (10 to 1000 MiB/s) at the expense of ratio, while context-mixing algorithms achieve superior density at serial speeds (0.1 to 1 MiB/s). We present GPX, an adaptive multi-tier compression architecture for structured data streams. GPX operates transparently across arbitrary stream lengths (N >= 1,024 B, identity fallback below 1 KiB), combining an autonomous structural decision probe (6.7--486.2 us across 49 benchmark instances, median 31.6 us, <0.05% overhead), cache-tiled AVX2 SIMD reversible domain transforms, multi-hypothesis sequence optimization with in-flight frame MDL gating, and adaptive entropy boundaries. On canonical Silesia (202.12 MiB) under a 7-repeat protocol (W = 6 pinned P-cores), GPX Track A produces standard RFC 8878-compliant .zst streams compressing to 56.20 MiB at 191.44 MiB/s with native line-speed decompression (5,182.79 MiB/s at W=6, 1,706.52 MiB/s at W=1 on unmodified libzstd, representing 124.1% single-core speed of Stock Level 9). GPX Track B achieves 54.49 MiB at 299.23 MiB/s (CI [295.3, 305.7]) with 4,919.14 MiB/s decode, strictly dominating Stock Levels 9--14 across size and speed (1.04x faster and 1.95 MB smaller than Level 9; 151.78 KB smaller and 8.24x faster than Level 14). GPX Track A+B achieves 54.37 MiB at 204.87 MiB/s (CI [199.6, 210.3]) with 4,599.00 MiB/s decode, strictly dominating Levels 10--14 under 95% Bootstrap CIs (276.07 KB smaller than Level 14, 1.19x to 5.64x faster). Generalization is confirmed across 6 suites (49 instances, 48 streams): Canterbury (-0.65%), Calgary (-1.66%), PE binaries (-14.10%), and Transformer BFloat16 tensors (-10.61%, saving 3.28 MB vs Stock L9). Downstream SIMT nibble unpacking sustains 103.62 GiB/s (1.05x PCIe speedup). All outputs are bit-exact verified via SHA-256.

cs.AR↗

Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller

DRAM continues to limit the performance, energy efficiency, and robustness of modern systems. Many prior works propose DRAM techniques that support in-DRAM computation, improve memory access latency and parallelism, and enhance DRAM maintenance and reliability. However, adopting each new technique requires repeated modifications to the rigid DRAM interface and memory controller, hindering its deployment. Our goal is to reduce these repeated modifications. We observe that the DRAM commands and memory controller structures of many DRAM techniques are similar. Our key idea is to use these similarities to compose a set of primitives for implementing diverse DRAM techniques. We propose Terracotta, a new framework with two flexible components: (i) custom command extensions that let DRAM vendors define new commands within a single, standardized interface, and (ii) a programmable memory controller that system designers can program to support new DRAM techniques post-silicon. Together, these enable deployment by configuring the memory controller instead of modifying the interface and controller. We design Terracotta for a DDR5-based system and evaluate its performance, energy, and hardware complexity. For four DRAM techniques from four distinct domains (processing-using-DRAM, low-cost DRAM maintenance, subarray-level parallelism, and latency reduction), Terracotta retains almost all of the performance benefits (>96%) of custom implementations. A Terracotta-based composition of two techniques outperforms the Terracotta-based implementation of each technique alone, demonstrating the benefits of adding techniques without repeated interface and controller modifications. Terracotta incurs low DRAM energy (0.6-3.2%), area (0.03%), and power (0.56%) overheads in a high-end server-grade processor. Terracotta's source code is freely available at https://github.com/CMU-SAFARI/Terracotta.

cs.AR↗

Beyond No-Good Benders Cuts: Exact Realizability for Discretely Tunable Clock Trees

Useful-skew schedules can satisfy timing constraints yet remain unrealizable by a fixed clock tree with discrete tuning choices. We formulate this mismatch as exact membership in a finite relative-latency set and develop a checkable feedback interface between the scheduler and the tree model. Arithmetic certificates explain unrealizable targets, while tree-specific contraction reduces the structural size of the exact relations projected onto selected sinks. These relations become realizability cuts that can exclude more candidates than a no-good on the same certificate support, including discrete holes that linear inequalities cannot separate. Controlled experiments confirm fewer oracle calls and faster projection construction. They also expose important limits: globally minimum certificates can cost more than they save, compact arithmetic feedback can fail on non-parity obstructions, and direct monolithic optimization remains faster on the tested additive models. The contribution is an exact, independently verifiable scheduler--tree interface, rather than a claim of universal solver acceleration or physical signoff.

cs.AR↗