Search arXiv⌕ Search

arXiv · 2610.03140

Tracking State Footprints: How Agents Can Transact

Abstract

Can AI agents transact? We argue that they must: as multi-agent systems (MASs) increasingly write code, deploy infrastructure, modify databases, and call web services, lost updates or stale reads can be catastrophic. Although MASs increasingly execute plans in parallel, current orchestrators do not track the state that agents read and write. As a result, concurrency anomalies manifest even in simple coding tasks. We frame MAS coordination as a data management problem and propose to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems. We posit that MASs require guarantees similar to those of databases, but providing them raises new challenges and opportunities: unlike database transactions, agents do not read from a fixed schema or an isolated snapshot, and cannot be replayed deterministically upon failure. They can, however, resolve conflicts semantically instead of aborting, enabling new forms of concurrency control and conflict resolution. Towards agents that can transact, we outline a vision for next-generation agent orchestrators and transactional interfaces for external systems to participate in agentic transactions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Oto Mraz, Rares Şerban, Kyriakos Psarakis, Burcu Kulahcioglu Ozkan, Asterios Katsifodimos. 2026-10-02. Tracking State Footprints: How Agents Can Transact. https://arxiv.org/abs/2610.03140

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grain Theory: Type-Level Granularity Correctness in Data Pipelines

Data pipelines fail silently when a transformation changes the grain of the data, the level of detail each element represents, without the designer intending it; fan traps and chasm traps are the familiar symptoms. Dimensional modeling has used grain informally since the 1990s. We define grain by irreducibility and isomorphism alone, with no appeal to functional dependencies, and compute it compositionally over products, sums and the inductive and coinductive fixed points, so it is defined on arrays, event streams and the other nested and recursive types where dependency theory has no referent, and the grain of a type's entity is its entity key, where collections of different grain integrate. Grain is more than a key: one structure under different declared grains denotes different things. The grain ordering is a partial order on types, axiomatized by Armstrong-style axioms whose completeness transfers on product types. For the relational algebra we prove inference rules, chief among them an equi-join theorem computing the exact, minimal grain. Grain projection commutes with transformation and grain lifts compose, and the algorithm CalcG decides pipeline grain-correctness at design time, from the schema alone, in time linear in the pipeline. The errors then follow from a single comparison: a grain the calculus computes, set against a grain the design intends, stands in exactly one relation to it in the grain ordering, and that position names the failure. Fan traps, dropped grain fields, ungrounded multi-version reads, wrong-grain aggregations and behavioral-class violations are the instances of this classification; chasm traps, a data-instance failure the theory does not model, are localized but not decided. The theory is mechanized in Lean 4.

cs.DB↗

Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents

However, building reliable data agents remains difficult because real enterprise data is fragmented across many heterogeneous database systems, with duplicated and inconsistent data, and key information often buried in unstructured text, requiring agents to go beyond just writing SQL or data science scripts to answer questions. Existing benchmarks tackle only individual pieces of this end-to-end workflow (e.g., text-to-SQL over a single database) and are increasingly saturated and contaminated. We present a new benchmark for LLM agents, the Data Agent Benchmark (DAB), grounded in a study of enterprises building production data agents across six industries. DAB comprises 104 queries across 17 datasets and 4 database management systems. On DAB, the best state-of-the-art agent achieves only 57% pass@1. We analyze agent failure modes and distill takeaways for future data-agent development. Our benchmark and experiment code are published at github.com/ucbepic/DataAgentBench.

cs.DB↗

Query Performance Tuning with Optimal Exploration of Optimizer Cost Model Parameter Space

Modern query optimizers use analytical cost models to estimate the cost of a given query plan. Such cost models are typically functions of a set of "cost units" that specify unit CPU cost when processing a row or unit IO cost when accessing a disk page. These cost units are traditionally viewed as platform-dependent constants, that is, they require a one-shot calibration when a database is deployed on a hardware/software platform, but are fixed afterward regardless of the query being optimized for. Some very recent work has taken a different perspective by viewing these cost units as tunable parameters that we call "cost model parameters (CMPs)" in this paper. However, so far there is no approach that offers any optimality guarantee for the tuning results. We present a new approach to systematically explore the query plan space spanned by the CMPs and find the best plan in terms of execution time. Compared to alternative exploration approaches that use random search (RS) or Bayesian optimization (BO), our new approach is deterministic and, therefore, avoids the undesirable instability that is inevitable when applying RS or BO. Moreover, it is guaranteed to find all candidate plans in the query plan space without suffering from the overhead of an exhaustive enumeration. We also present a set of optimization techniques to reduce the overall evaluation time spent on executing the candidate plans found, a factor that is often overlooked by previous work but is critical from a practical point of view. Experimental evaluation on top of PostgreSQL and Microsoft SQL Server demonstrates the efficacy of tuning the CMPs, which can find query plans that are orders of magnitude faster in execution time than the ones found by RS or BO.

cs.DB↗