Search arXiv⌕ Search

arXiv · 2508.00234

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts

Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities, leading to a significant increase in user demand for LLM services. However, cloud-based LLM services often suffer from high latency, unstable responsiveness, and privacy concerns. Therefore, multiple LLMs are usually deployed at the network edge to boost real-time responsiveness and protect data privacy, particularly for many emerging smart mobile and IoT applications. Given the varying response quality and latency of LLM services, a critical issue is how to route user requests from mobile and IoT devices to an appropriate LLM service (i.e., edge LLM expert) to ensure acceptable quality-of-service (QoS). Existing routing algorithms fail to simultaneously address the heterogeneity of LLM services, the interference among requests, and the dynamic workloads necessary for maintaining long-term stable QoS. To meet these challenges, in this paper we propose a novel deep reinforcement learning (DRL)-based QoS-aware LLM routing framework for sustained high-quality LLM services. Due to the dynamic nature of the global state, we propose a dynamic state abstraction technique to compactly represent global state features with a heterogeneous graph attention network (HAN). Additionally, we introduce an action impact estimator and a tailored reward function to guide the DRL agent in maximizing QoS and preventing latency violations. Extensive experiments on both Poisson and real-world workloads demonstrate that our proposed algorithm significantly improves average QoS and computing resource efficiency compared to existing baselines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jin Yang, Qiong Wu, Zhiying Feng, Zhi Zhou, Deke Guo, Xu Chen. 2025-08-01. Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts. https://arxiv.org/abs/2508.00234

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Secure Polarization-Shift Backscatter Identification Applied to Battery-Free BLE Sensors Powered by Wireless Power Transfer

This paper presents a lightweight and protocolindependent security mechanism for battery-free Bluetooth Low Energy (BLE) sensor nodes operating in Simultaneous Wireless Information and Power Transfer (SWIPT) architecture. The proposed approach exploits polarization-shift backscattering of the wireless power wave to transmit an encrypted device identification prior to data communication. A fail-safe RF switch and orthogonally polarized antennas are integrated as an external add-on module, enabling controlled backscatter without modifying the original energy-harvesting rectifier. The identification payload is encrypted using AES-128 and transmitted with minimal energy overhead. Experimental validation on a battery-free BLE sensor node demonstrates reliable extraction of the backscattered identification signal, seamless coexistence with BLE advertising, and improved RF-to-DC harvesting efficiency compared to rectifier-based backscatter solutions. The results confirm that polarization-shift backscatter identification provides an effective and practical security for battery-free BLE sensing systems.

cs.NI↗

From WPT to Encrypted Telemetry: A Battery-Free Backscattering-based Polarimetric Wireless Sensor

This work introduces an indoor Battery-Free Wireless Sensing Node powered through radiative Wireless Power Transfer (WPT). The proposed platform targets secure, energyefficient active sensing and overcomes key limitations of many prior battery-free approaches, which commonly provide neither on-node computation nor cryptographic protection. The node combines temperature, humidity, pressure and Volatile Organic Compound (VOC) measurements with a low-power microcontroller that executes sensor calibration, derives a VOC index, formats the payload, and applies AES-128 encryption before wireless transmission. Energy harvesting and communication are enabled by a 1-bit controlled Backscatter Rectenna (BR), which both scavenges incident RF power and produces an orthogonally polarized backscattered signal for robust polarimetric operation. Experimental results validate reliable multi-sensor readout and encrypted data transfer, while maintaining a very low energy budget for the complete sense-compute-encrypt-transmit cycle.

cs.NI↗

NebulaSD: Many-for-Many Speculative Decoding

Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Drafting and verification, however, exhibit different service characteristics and favor different batch configurations, making fixed draft-target coupling inefficient under concurrent workloads. Existing distributed designs can physically separate the two stages, but often retain request or batch affinities that prevent their capacities from being shared globally. We present NebulaSD, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools. Such dynamic reassignment removes fixed worker locality, requiring request states to be made available at newly selected workers without introducing migration stalls. NebulaSD addresses this challenge through worker-triggered batch reconstruction and asynchronous KV-state preparation overlapped with model execution. We evaluate NebulaSD from both system and scaling perspectives, showing that dynamic pooling improves request-round processing rate by 50.4% over a physically disaggregated baseline and 72.6% over co-located execution on a four-GPU deployment while substantially increasing effective GPU utilization. Profile-driven simulations further show approximately proportional compute-side capacity scaling under idealized state movement.

cs.NI↗