How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings
Document parsing models are improving rapidly, but existing benchmarks are still insufficient for a comprehensive evaluation of their capabilities. We introduce PureDocBench, a benchmark with 10 domains, 66 subcategories, and 1,475 pages. Each page has matched Clean, Digital, and Real versions, covering clean renders, simulated degradation, and real capture or sharing conditions. The resulting 4,425 images have source-derived annotations for text, formulas, tables, and reading order, checked through automated validation and human review. Keeping page content and annotations fixed allows direct comparisons across image conditions. We evaluate 58 models, including pipeline-based parsers, end-to-end parsers, and general vision-language models. The best three-track average score is 79.69 out of 100, and the mean across models is 63.84. Scores drop by 3.34 points on Digital and 11.61 points on Real on average, relative to Clean. TeleOCR leads on Clean, GLM-5.3-Flash on Digital, and Gemini 3.6 Flash on Real. The benchmark dataset, source documents, generation tools, and evaluation code are publicly available at https://github.com/zhihengli-casia/PureDocBench.