arXiv · 2609.07461
Temporal-Causal Inference for Reinforcement Learning via Automata Learning
Abstract
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.
Explore related subjects
Keep this discovery
Jan Corazza, Daniil Kaminskyi, Simon Lutz, Patrick Nossol, Hadi Partovi Aria, Zhe Xu, Daniel Neider. 2026-09-07. Temporal-Causal Inference for Reinforcement Learning via Automata Learning. https://arxiv.org/abs/2609.07461
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.