Search arXiv⌕ Search

arXiv · 2609.36421

Sample Complexity of Equivariant Reinforcement Learning

Abstract

Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake. 2026-09-29. Sample Complexity of Equivariant Reinforcement Learning. https://arxiv.org/abs/2609.36421

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Deterministic Depth-4 PIT and Normalization

In this paper, we initiate the study of deterministic PIT for $Σ^{[k]}ΠΣΠ^{[δ]}$ circuits over fields of any characteristic, where $k$ and $δ$ are bounded. Our main result is a deterministic polynomial-time black-box PIT algorithm for $Σ^{[3]}ΠΣΠ^{[δ]}$ circuits, under the additional condition that one of the summands at the top $Σ$ gate is squarefree. Our techniques are purely algebro-geometric: they do not rely on Sylvester--Gallai-type theorems, and our PIT result holds over arbitrary fields. The core of our proof is based on the normalization of algebraic varieties. Specifically, we carry out the analysis in the integral closure of a coordinate ring, which enjoys better algebraic properties than the original ring.

cs.CC↗

Local Search with Correlated Randomness

How much does an algorithm's running-time distribution under independent randomness reveal about its behavior when independence is no longer guaranteed? We study sources satisfying $ν[w]\le DP[w]^s$ for every finite prefix $w$, where $P$ is an independent reference law, $0<s\le1$, and $D\ge1$. The constraint controls complete-prefix probabilities while allowing individual choices to be predictable, even fully determined by the past. For retry tasks, all deterministic history-dependent selectors have the same independent-source running-time law. Yet two orders have worst-case failure probabilities $1$ and $\exp[-Θ(n)]$ at the same linear deadline under the same source constraint. We identify a static priority rule that is optimal at every deadline and every $D$. For the standard local walk on a $k$-CNF with at least $r$ true literals per clause under some assignment, $k/2<r<k$, we determine the sharp source threshold $s_*$. At and above it, the expected flip count is $O_{k,r}(\min\{L^3,L/(s-s_*)\})$, where $L=h+\log D+1$, $h$ is the initial Hamming distance to that assignment, and $L/0=\infty$. The bound allows arbitrary clause overlap and history-dependent clause selection. Matching instances admit one source forcing this delay with probability one for every selector. At criticality and fixed $D$, the delay is cubic despite a linear independent-source expectation. Variable-depth prefix covers, together with classical tree max-flow/min-cut, yield an exact criterion for restoring exponential tails by restarting on the same tape. We synthesize updates and restarts for explicit finite-state processes. Under a sufficient prefix guarantee, we also obtain noisy predecessor search with error at most $η$ and expected query count polynomial in the correct leaf's depth and $\log(D/η)$, without knowing the depth or tree height.

cs.CC↗

Three-Color Free-Flood-It on Fixed-Height Grids Is Polynomial-Time Solvable

We give a deterministic algorithm for \textsc{Free-Flood-It} on rectangular grids $P_k\square P_n$ with at most three colors. For every fixed height $k$, it computes the minimum number of moves and an optimal sequence in $N^{O(k^2)}$ time, where $N=kn$. This resolves the previously open three-color case on complete $3\times n$ boards. The proof uses a representation of flooding strategies as paintings by connected regions. We show that an optimal painting can be chosen so that its regions form a rooted tree with strong restrictions on the colors along ancestral paths. These restrictions make the part of the tree visible in any fixed-size connected window admit only polynomially many descriptions. The remaining common ancestors may form an arbitrarily long chain. We retain that chain on a stack and use a finite context-free recurrence to minimize the total painting cost without enumerating all possible stack contents. Connectivity information at the boundary ensures that the resulting local descriptions assemble into one valid global painting. The argument also applies to graphs supplied with an ordering into connected layers of bounded size, provided every edge lies within one layer or joins consecutive layers. A family of three-row boards shows that the unbounded ancestor chains handled by the algorithm are necessary even for optimal paintings. For any fixed height and palette size, we also give a deterministic EPTAS and a randomized sampling variant with an explicit failure-probability bound. These approximation results use a separate algorithm with single-exponential dependence on the move budget.

cs.CC↗