Search arXiv⌕ Search

arXiv · 2609.29647

AgentKernel: The Trust-Native Agentic Operating System

Abstract

Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control. We introduce AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint. AgentKernel wraps the agent lifecycle in a mandatory enforcement boundary organized into four pillars: Identity, Perception, Cognition, and Execution. Each pillar adapts classical OS security principles to failures at the semantic plane, including delegation abuse, prompt injection, memory poisoning, and tool misuse. AgentKernel treats structural security as a capability multiplier. Kernel-managed identity supports trustworthy cross-organization collaboration; graduated perception replaces brittle single-point filters; information-flow-controlled memory improves retrieval fidelity while limiting poisoning; and semantic-to-kernel enforcement permits broader tool privileges behind a non-bypassable boundary. We position AgentKernel as the missing OS layer beneath orchestration frameworks, agent runtimes, governance platforms, and execution sandboxes, and use systematic comparison and security analysis to show how a single integrated architecture can enforce security across the full agent lifecycle.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao, Shuo Li, Zhuotao Liu. 2026-08-29. AgentKernel: The Trust-Native Agentic Operating System. https://arxiv.org/abs/2609.29647

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Server-Enforced Watermarking in U-Shaped Split Federated Learning

U-shaped split federated learning (U-SFL) enables resource-constrained Internet of Things devices to collaboratively train models with an edge server while retaining raw data and labels locally. We consider model-service deployments in which a provider supplies proprietary models whose client-side segments reside on participating devices, creating risks of unauthorized copying and redistribution. Protecting these segments is challenging because the server cannot directly access client data, labels, or model parameters, and potentially malicious clients may refuse to perform watermark embedding. To address this challenge, we propose Sigil, a server-enforced watermarking framework for U-SFL. Sigil defines a secret watermark constraint in the server-visible activation space and embeds the watermark into client-side models by injecting a watermark gradient into the gradients returned during training. This mechanism requires neither access to clients' raw data and labels nor client-side watermarking operations. To limit interference with the main task and reduce detectability by gradient anomaly detectors, Sigil adaptively clips the watermark gradient relative to the main-task gradient. Experiments on two datasets and four model architectures demonstrate high watermark detection rates with limited impact on task accuracy, robustness against the evaluated removal attacks, and stealthiness against the evaluated gradient anomaly detector.

cs.CR↗

Digital Agriculture Sandbox for Collaborative Research

Digital agriculture is transforming the way we grow food by utilizing technology to make farming more efficient, sustainable, and productive. This modern approach to agriculture generates a wealth of valuable data that could help address global food challenges, but farmers are hesitant to share it due to privacy concerns. This limits the extent to which researchers can learn from this data to inform improvements in farming. This paper presents the Digital Agriculture Sandbox, a secure online platform that solves this problem. The platform enables farmers (with limited technical resources) and researchers to collaborate on analyzing farm data without exposing private information. We employ specialized techniques such as federated learning, differential privacy, and data analysis methods to safeguard the data while maintaining its utility for research purposes. The system enables farmers to identify similar farmers in a simplified manner without needing extensive technical knowledge or access to computational resources. Similarly, it enables researchers to learn from the data and build helpful tools without the sensitive information ever leaving the farmer's system. This creates a safe space where farmers feel comfortable sharing data, allowing researchers to make important discoveries. Our platform helps bridge the gap between maintaining farm data privacy and utilizing that data to address critical food and farming challenges worldwide.

cs.CR↗

Stateful Agent Backdoors: Constructing Cross-Session Attack Programs

Multi-step attacks on large language model agents may depend on opportunities distributed across sessions, such as access to target information or the availability of required tools. To combine these opportunities, an attack needs to retain its state and intermediate results, and choose actions based on current conditions. In this work, we study such attacks as cross-session attack programs. We focus on their shared control structures, including sequential execution with conditional waiting, branching and merging, looping, and condition accumulation. We represent these programs as Mealy machines and construct a sub-backdoor for each transition. We build single-session training trajectories for each sub-backdoor, combine them into a training dataset, and fine-tune the model to learn the local behaviors jointly. After a single injection of the initial trigger, the agent uses persistent memory to connect local behaviors into a complete cross-session attack program. We evaluate four instantiations of these control structures across four models in a LangChain-based agent environment. The primary instantiation achieves mean complete-program success rates of 71.7\%--94.7\% in LangChain. In the official OpenClaw runtime under a controlled configuration, it achieves a complete-program success rate of 85\% for each of two evaluated models. These results demonstrate the feasibility of end-to-end execution of cross-session attack programs with different control structures.

cs.CR↗