Search arXiv⌕ Search

arXiv · 2609.31076

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Abstract

Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bartłomiej Cupiał, Jens Tuyls, Maciej Wołczyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Miłoś, Karthik R. Narasimhan. 2026-09-25. Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents. https://arxiv.org/abs/2609.31076

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution

This note is a substantially revised version of the author's earlier preprint (version~1, 2023), which proposed a ``Categorical Graph Model for Conflict Resolution'' (C-GMCR). Several of the central definitions and claims of version~1 are incorrect, and this version replaces them. We first observe that the states and one-step moves of a graph model do not form a category; the correct construction is the free category on a quiver whose arrows are moves labelled by the decision maker (DM) who controls them. Controller labels are functorial: they define a functor into the free monoid on the set of DMs, and the legal move sequences underlying coalition reachability are exactly the paths whose label words contain no immediate repetition. Preferences, in contrast, are not functorial. We show that a map representing a DM's preference in a preordered set exists only if the preference is transitive, and that even then it extends to a functor on the path category if and only if every move, regardless of its controller, is weakly improving for the focal DM; in that case general metarationality, symmetric metarationality and sequential stability all collapse to Nash stability for the focal DM. Preferences therefore enter the categorical picture as a selection of a subquiver of improvements, not as a functor. We restate the four standard stability concepts in path language, work them out on the Prisoner's Dilemma, and take a first step towards comparing graph models via induced embeddings: Nash stability is reflected by induced embeddings, whereas general metarationality, symmetric metarationality and sequential stability are neither preserved nor reflected. A section lists each correction to version~1 explicitly.

cs.AI↗

From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark

Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs. Our code is available at https://github.com/you68681/kaar.

cs.AI↗

GLOVE: Global Verifier for LLM Memory-Environment Realignment

Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through internal model cognition, such as reflection, for editing memory entries. However, these assumptions often break down in practical environments with dynamic drifts. We propose the Global Verifier (GLOVE), a framework that introduces a new design dimension for LLM memory systems by establishing a relative notion of truth. Through active probing to detect inconsistencies between retrieved memories and fresh observations, GLOVE enables memory-environment realignment by verifying and updating memory without access to ground-truth supervision or strong reliance on model introspection. We evaluate GLOVE on diverse benchmarks spanning web navigation, planning, and control, augmented with controlled environmental drifts that introduce non-stationarity beyond the original benchmark settings. Our results show that GLOVE substantially improves agent success rates, suggesting a robust pathway to cognitive agents capable of self-evolving.

cs.AI↗