Search arXivSearch

arXiv · 2002.11101

Deep Reinforcement Learning for Intelligent Reflecting Surfaces: Towards Standalone Operation

Abstract

The promising coverage and spectral efficiency gains of intelligent reflecting surfaces (IRSs) are attracting increasing interest. In order to realize these surfaces in practice, however, several challenges need to be addressed. One of these main challenges is how to configure the reflecting coefficients on these passive surfaces without requiring massive channel estimation or beam training overhead. Earlier work suggested leveraging supervised learning tools to design the IRS reflection matrices. While this approach has the potential of reducing the beam training overhead, it requires collecting large datasets for training the neural network models. In this paper, we propose a novel deep reinforcement learning framework for predicting the IRS reflection matrices with minimal training overhead. Simulation results show that the proposed online learning framework can converge to the optimal rate that assumes perfect channel knowledge. This represents an important step towards realizing a standalone IRS operation, where the surface configures itself without any control from the infrastructure.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abdelrahman Taha, Yu Zhang, Faris B. Mismar, Ahmed Alkhateeb. 2020-02-25. Deep Reinforcement Learning for Intelligent Reflecting Surfaces: Towards Standalone Operation. https://arxiv.org/abs/2002.11101

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the Distribution of Age of Information in Time-varying Updating Systems

Age of Information (AoI) is a crucial metric for quantifying information freshness in real-time systems where the sampling rate of data packets is time-varying. Evaluating AoI under such conditions is challenging, as system states become temporally correlated and traditional stationary analysis is inapplicable. We investigate an $M_{t}/G/1/1$ queueing system with a time-varying sampling rate and probabilistic preemption, proposing a novel analytical framework based on multi-dimensional partial differential equations (PDEs) to capture the time evolution of the system's status distribution. To solve the PDEs, we develop a decomposition technique that breaks the high-dimensional PDE into lower-dimensional subsystems. Solving these subsystems allows us to derive the Aol distribution at arbitrary time instances. We show AoI does not exhibit a memoryless property, even with negligible processing times, due to its dependence on the historical sampling process. Our framework extends to the stationary setting, where we derive a closed-form expression for the Laplace-Stieltjes Transform (LST) of the steady-state AoI. Numerical experiments reveal AoI exhibits a non-trivial lag in response to sampling rate changes. Our results also show that no single preemption probability or processing time distribution can minimize Aol violation probability across all thresholds in either time-varying or stationary scenarios. Finally, we formulate an optimization problem and propose a heuristic method to find sampling rates that reduce costs while satisfying AoI constraints.

cs.IT

Properties of Random Code Ensembles over Classical-Quantum Channels

We show two properties of i.i.d. code ensembles over classical-quantum (CQ) channels with arbitrary output states. The first property is that the probability distribution of the error exponent across the ensemble accumulates above a threshold that is strictly larger than the CQ random coding exponent (RCE) at low rates, while coinciding with it at rates close to the mutual information of the channel. This result, combined with the works by Dalai, Renes, Li and Yang and Cheng and Liu, implies that the ensemble distribution of error exponents concentrates around the CQ RCE in the high rate regime. Moreover, in the same rate region the threshold we derive coincides with the ensemble-average of the exponent, that is, the CQ typical random coding (TRC) exponent. The second property we derive is that the probability that a randomly selected code from an ensemble contains a nested code of the same rate such that each of its codewords achieves the CQ expurgated exponent goes to one asymptotically in the codeword length. Such nested code can be obtained by expurgating a vanishingly small fraction of codewords as the codeword length increases. This result refines Holevo's work [5], which proved that the expurgated exponent can be achieved by discarding the worst half of the codewords from at least one code in the ensemble.

cs.IT

Taming Subpacketization without Sacrificing Communication: A Packet Type-based Framework for D2D Coded Caching

Finite-length design is essential for making coded caching practical, as the optimal communication gains of existing schemes often require prohibitively large subpacketization. This paper studies rate-optimal device-to-device (D2D) coded caching with reduced subpacketization. We propose a packet type-based (PT) framework that exploits the geometric structure induced by user grouping. Under this structure, subfiles, packets, and multicast groups are classified into types, allowing the originally symmetric Ji-Caire-Molisch (JCM) design~\cite{ji2016fundamental} to be systematically relaxed without sacrificing the optimal D2D communication rate. The key feature of the PT framework is that subpacketization reduction is achieved through two complementary mechanisms: \emph{subfile saving}, by excluding redundant subfile types, and \emph{further-splitting saving}, by assigning type-dependent further-splitting factors to subfiles through transmitter selection. The type-dependent splitting factors are then coordinated across multicast group types to produce a globally consistent file-splitting structure. Based on this framework, we construct several classes of rate-optimal D2D coded caching schemes that strictly improve upon the JCM subpacketization. The proposed schemes achieve either order-wise reductions in the number of users or constant-factor reductions over broad memory regimes, while preserving the optimal rate. These results reveal a structural distinction between D2D and shared-link coded caching: unlike in the shared-link setting, full symmetric subpacketization is not necessary for rate-optimal D2D caching.

cs.IT