Search arXiv⌕ Search

arXiv · 2610.04243

Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD

Abstract

Active Directory (AD) remains the predominant identity and access management infrastructure in enterprise environments, and its compromise represents the highest-impact outcome in internal penetration tests. Recent work has shown that large language models (LLMs) can autonomously conduct assumed-breach penetration testing against AD, but these studies employ standalone agents lacking structured guardrails, deterministic validation, and multi-stage chain orchestration. We present a benchmark evaluation of NeuroSploit v4.2.0, an open-source Rust-based autonomous pentest harness, against the Game of Active Directory (GOAD), a deliberately vulnerable multi-forest AD lab maintained by Orange Cyberdefense comprising five virtual machines, two forests, and three domains. The harness orchestrates 22 AD-specific agents and 7 multi-stage attack-chain playbooks covering the full AD kill chain: enumeration, Kerberoasting, AS-REP roasting, NTLM relay and coercion, Kerberos delegation abuse, AD CS exploitation (ESC1-ESC8), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse, and persistence detection. We benchmark nine frontier LLMs (Claude Opus 4.6/4.7/4.8, GPT-6 Astra, GPT-5.6 Sol, Grok 4.6, Qwen 3.8, GLM 5.3, Kimi k3) within the harness, comparing against direct invocation across 14 technique categories and 7 chains. The harness achieves 96-100% technique coverage with 90-97% precision, while direct invocation covers only 21-54% and produces 3.2x more false positives. Time to full three-domain compromise with Opus 4.8 was 134 minutes with guardrail activations preventing lockout-triggering sprays, unauthorized DCSync dumps, and out-of-scope reconnaissance. Results demonstrate that structured harness orchestration with domain-specialized agents, POMDP belief tracking, and cross-model voting substantially outperforms unstructured LLM usage for AD penetration testing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joas Antonio dos Santos Barbosa. 2026-10-03. Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD. https://arxiv.org/abs/2610.04243

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16 times amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques. Our code is available at https://github.com/multimodal-ai-lab/EEB.

cs.CR↗

Condense to Conduct and Conduct to Condense

In this paper, we present the first explicit examples of low-conductance permutations. The notion of conductance of permutations was introduced by Dodis et al. in "Indifferentiability of Confusion-Diffusion Networks", where the search for low-conductance permutations was first initiated and motivated. As part of our contribution, we not only provide these examples, but also offer a general characterization of the problem: we show that low-conductance permutations are equivalent to permutations possessing the information-theoretic properties of Multi-Source-Somewhere-Condensers, a specific variant of somewhere condensers.

cs.CR↗

Potential and Challenges of Large Language Models for Reverse Engineering

Reverse engineering (RE) is central to cybersecurity, supporting tasks such as decompilation, deobfuscation, and security analysis. However, RE remains labor-intensive and expertise-demanding, as analysts often manually recover high-level semantics from low-level program representations. Recent advances in large language models (LLMs) provide a promising way to address these challenges through program artifact understanding, semantic reasoning, and tool-augmented problem solving. This potential has stimulated interest in LLM-assisted RE, but existing studies remain scattered across different tasks, targets, methodologies, and evaluation practices. Despite these advances, the literature still lacks a comprehensive survey that consolidates progress, systematizes technical choices, and clarifies open challenges and future opportunities. To fill this gap, we present a systematic survey of LLM-assisted RE, covering 48 peer-reviewed and published research articles identified through our search and selection protocol as of July 1, 2026. We develop a faceted taxonomy that organizes prior studies along six dimensions: task objective, analysis target, methodological approach, evaluation protocol, training scale, and data quality. We extract task formulations and experimental settings from existing studies to support comparison, reproducibility, and future research. From this review, we synthesize research gaps, characterize key challenges, and outline future directions toward more reliable, reproducible, and security-relevant applications of LLMs in RE.

cs.CR↗