Search arXiv⌕ Search

arXiv · 2610.06848

TranScope: What the Software Hides About LLM Training Data, the Hardware Reveals at Scale, and Accelerators Magnify

Abstract

Membership is the root privacy primitive in machine learning: to date, no hardware-based out-of-distribution detection on black-box models has been demonstrated against constant-time, static neural networks with masked confidence. This paper performs the first cycle-level examination of how large language models and vision transformers interact with various modern microarchitecture components, including integrated accelerators, as LLMs scale in size and answers the question of whether the data that a model was trained on affects its execution footprint even without any input-dependent branch, dynamic optimization, or early exit and in constant-time models. The results confirm that the answer is yes and identify which modern hardware components, such as TLBs or on-core accelerators, reveal or amplify that effect. The results also answer whether the signal is informative enough to reliably classify the in-/vs/out-of-distribution property of membership. To understand why, we perform a systematic root cause analysis and find that the transformer's tokenization steps, which happen during training, alter the locality of the accesses the model makes to fetch the vocabulary token later during inference and, as a result, change the page table access patterns and TLB in a previously unknown data-dependent way, causing microarchitectural state to vary significantly based on whether or not the input was in the distribution of the transformer training data. Building on the above observation, we introduce TranScope: the first microarchitecture tool for detecting membership information with low cost, no need for a surrogate model, and significantly higher robustness, e.g., 0.6 AUC for PETAL (best previously reported) vs 0.9 AUC (ours). This reintroduces hardware as both an opportunity, e.g., a tool for checking copyright violation for the first time, and a new channel for inferring membership (MIA).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joshua Kalyanapu, Darsh Asher, Kaushal Mhapsekar, Bita Aslrousta, Rushiraj Chaitanyakumar Sheth, Maharshi Mukeshkumar Oza, Achyuta Kannan, Samira Mirbagher Ajorpaz. 2026-10-05. TranScope: What the Software Hides About LLM Training Data, the Hardware Reveals at Scale, and Accelerators Magnify. https://arxiv.org/abs/2610.06848

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Speed Kills: Exploring Confused Deputy Attacks Through Edge AI Accelerators

AI Accelerators (AIAs) are specialized hardware, such as Tensor Processing Units (TPUs), that enable optimal and efficient execution of AI applications and on-device inference. The growing demand for AI applications has led to the widespread adoption of AIAs on edge and embedded devices. Unlike applications, AIAs are not bound by Operating System (OS) restrictions and have limited visibility into Application Processor (AP) security mechanisms (e.g., kernel versus application memory and process isolation). This semantic gap can lead to confused deputy vulnerabilities, where an AIA can be tricked by a malicious application into performing privileged operations on its behalf. In this paper, we conduct the first in-depth study of Confused Deputy Attacks (CDAs) using AIAs. We design DeputyHunt, a Large Language Model (LLM)-assisted framework to extract CDA-relevant information for a given AIA through a combination of dynamic and static analysis. We use this information to explore the feasibility of CDAs on seven different AIAs from popular vendors, including Google, NVIDIA, Hailo, Texas Instruments, NXP, AWS, and Rockchip. Our analysis reveals that CDAs are feasible on six out of the seven AIAs, impacting over 128 System-on-Chips (SoCs) and over 100 million devices. Our findings highlight critical security risks posed by AIAs to system security. Our work has been acknowledged by the corresponding vendors and assigned CVE-2025-66425. We propose an on-demand validation defense against CDAs, and our evaluation on the Gem5-salam simulator shows that the average runtime overhead of our mechanism is significantly lower (~15% versus ~56%) than that of traditional defenses.

cs.CR↗

Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configurations, query budgets, and security objectives, limiting comparability across methods. To address this gap, we introduce a unified benchmark covering six extraction attacks, ten defenses, and two adaptive attacks that paraphrase or back-translate protected responses before surrogate training. The benchmark controls model configurations, query data, budgets, and held-out evaluation conditions within each comparison while preserving attack-specific querying and training procedures. We measure surrogate capability, fidelity to the victim, output quality using Rep-4, and query-budget sensitivity; defenses use their own security metrics paired with surrogate performance. For the adaptive attacks, we jointly measure provenance-detector scores and the capability and fidelity of surrogates trained on rewritten responses. The benchmark thus provides a reproducible basis for comparing extraction methods and their interactions with defenses under text-only access. Code and artifacts are available at https://github.com/sliu11-byte/MEA-Bench.

cs.CR↗

SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models

Large language models (LLMs) offer a flexible approach to network intrusion detection, but direct classification of numerical flow records can be unreliable without traffic- specific decision boundaries. Retrieval-augmented generation (RAG) enables LLMs to leverage external experience; however, free-form textual experience requires the model to infer the numerical conditions that distinguish traffic classes. We propose SERA-IDS, a Structured Experience Retrieval-Augmented Intru- sion Detection framework that learns decision knowledge from classification errors. A Small Language Model (SLM) analyzes misclassified flows and generates structured rules containing behavioral descriptions, feature conditions, class-confusion infor- mation, tool-derived confidence, and provenance. A confidence gate admits supported rules into the Experience Library. During testing, the library is frozen, and retrieved class-diverse rules are condition-matched to the query flow and provided as evidence to the decision SLM. The same library is evaluated with three small, locally deployable SLMs, namely Llama 3.1, Phi-4:14B, and Qwen2.5:7B, without fine-tuning or paid hosted inference. On the test sets, SERA-IDS improves macro F1 on NF-BoT- IoT from 11.66%, 9.26%, and 7.83% to 86.46%, 69.50%, and 83.51%, respectively, and on NF-ToN-IoT from 3.82%, 6.17%, and 1.94% to 84.16%, 87.66%, and 83.37%. These results show that structured experience retrieval enables small, free, and locally deployable SLMs to achieve strong intrusion-detection performance, including performance exceeding that of the re- ported existing work used in our comparison.

cs.CR↗