Search arXivSearch

arXiv · 2010.15597

Enhancing reinforcement learning by a finite reward response filter with a case study in intelligent structural control

Abstract

In many reinforcement learning (RL) problems, it takes some time until a taken action by the agent reaches its maximum effect on the environment and consequently the agent receives the reward corresponding to that action by a delay called action-effect delay. Such delays reduce the performance of the learning algorithm and increase the computational costs, as the reinforcement learning agent values the immediate rewards more than the future reward that is more related to the taken action. This paper addresses this issue by introducing an applicable enhanced Q-learning method in which at the beginning of the learning phase, the agent takes a single action and builds a function that reflects the environments response to that action, called the reflexive $γ$ - function. During the training phase, the agent utilizes the created reflexive $γ$- function to update the Q-values. We have applied the developed method to a structural control problem in which the goal of the agent is to reduce the vibrations of a building subjected to earthquake excitations with a specified delay. Seismic control problems are considered as a complex task in structural engineering because of the stochastic and unpredictable nature of earthquakes and the complex behavior of the structure. Three scenarios are presented to study the effects of zero, medium, and long action-effect delays and the performance of the Enhanced method is compared to the standard Q-learning method. Both RL methods use neural network to learn to estimate the state-action value function that is used to control the structure. The results show that the enhanced method significantly outperforms the performance of the original method in all cases, and also improves the stability of the algorithm in dealing with action-effect delays.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hamid Radmard Rahmani, Carsten Koenke, Marco A. Wiering. 2020-10-25. Enhancing reinforcement learning by a finite reward response filter with a case study in intelligent structural control. https://arxiv.org/abs/2010.15597

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Analysis of Regularized Learning in Banach Spaces for Linear-functional Data

This article delves into the study of the theory of regularized learning in Banach spaces for linear-functional data. It encompasses discussions on representer theorems, pseudo-approximation theorems, and convergence theorems. Regularized learning is designed to minimize regularized empirical risks over a Banach space. The empirical risks are calculated by utilizing training data and multi-loss functions. The input training data are composed of linear functionals in a predual space of the Banach space to capture discrete local information from multimodal data and multiscale models. Through the regularized learning, approximations of the exact solution to an unidentified or uncertain original problem are globally achieved. In the convergence theorems, the convergence of the approximate solutions to the exact solution is established through the utilization of the weak* topology of the Banach space. The theorems of regularized learning are utilized in the interpretation of classical machine learning, such as support vector machines and artificial neural networks.

cs.LG

On Minimal Depth in Neural Networks

Understanding the relationship between the depth of a neural network and its representational capacity is a central problem in deep learning theory. In this work, we develop a geometric framework to analyze the expressivity of ReLU networks with the notion of depth complexity for convex polytopes. The depth of a polytope recursively quantifies the number of alternating convex hull and Minkowski sum operations required to construct it. This geometric perspective serves as a rigorous tool for deriving depth lower bounds and understanding the structural limits of deep neural architectures. We establish lower and upper bounds on the depth of polytopes, as well as tight bounds for classical families. These results yield two main consequences. First, we provide a purely geometric proof of the expressivity bound by Arora et al. (2018), confirming that $\lceil \log_2(n+1)\rceil$ hidden layers suffice to represent any continuous piecewise linear (CPWL) function. Second, we prove that, unlike general ReLU networks, convex polytopes do not admit a universal depth bound. Specifically, the depth of cyclic polytopes in dimensions $n \geq 4$ grows unboundedly with the number of vertices. This result implies that Input Convex Neural Networks (ICNNs) cannot represent all convex CPWL functions with a fixed depth, revealing a sharp separation in expressivity between ICNNs and standard ReLU networks.

cs.LG

DeepSPoC: A Deep Learning Based Sequential Propagation of Chaos

Classical particle methods based on propagation of chaos (PoC) have been developed for solving mean-field stochastic differential equations and their associated nonlinear Fokker--Planck equations. However, direct PoC implementations are difficult to apply to high-dimensional problems because they require simulating and storing large numbers of interacting particles, often with high particle-particle interaction costs. Motivated by these limitations, we build on the recently proposed sequential propagation of chaos (SPoC) framework, which replaces the fully interacting particle system in PoC with a sequential interaction mechanism. Based on this structure, we present DeepSPoC, a neural particle method that embeds a neural density representation into the sequential particle dynamics. DeepSPoC simulates particles batch by batch, while the neural network represents the evolving empirical law and is substituted into the coefficients of the mean-field SDE, thereby replacing direct particle-particle interactions with particle-network interactions. In DeepSPoC, a recently developed normalizing flow model called KRnet is used to approximate the empirical measure of particles. Compared with direct particle implementations, DeepSPoC substantially reduces memory consumption and evaluates interaction terms more efficiently, thereby improving scalability for high-dimensional problems. We apply DeepSPoC to a wide range of mean-field equations and verify its effectiveness and computational advantages.

cs.LG