arXiv · 2609.23953
Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
Abstract
AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how reliably an off-the-shelf coding agent, driving one of seven open-weight models with a shell and the stock Python PDF stack, alters one dollar amount, date or address in a real filed financial document from a single sentence of intent, graded by rules rather than by a model. Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide. A deterministic script with no model in it solves 98 of the 125 documents; the agents solve 124, and none the script solves alone. Agents misreport 41% of their wrong edits as done, no model refused, and the cheapest verified forgery costs 2.4 cents. The raw rate overstates the threat by about a factor of two; the strict rate is still large.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Simiao Ren, Ankit Raj, Tommy Duong, Yuxin Zhang, Dennis Ng, Xingyu Shen, Kidus Zewde, Yuchen Zhou, Neo Tiangratanakul. 2026-09-20. Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control. https://arxiv.org/abs/2609.23953
Cite the original work for its findings. Save a collection to share your selection of sources.