arXiv · 2609.33141
On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation
Abstract
Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device -- reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Moghis Fereidouni, Anthony Arnold, Sumit Gulwani, Mark Marron, A. B. Siddique. 2026-09-27. On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation. https://arxiv.org/abs/2609.33141
Cite the original work for its findings. Save a collection to share your selection of sources.