arXiv · 2609.33447
From PDF to Evidence: Structure-Aware Retrieval for Clinical Practice Guidelines
Abstract
Guideline documents are published as unstructured PDFs whose evidence is locked in visual structures---tables, flowcharts, and graded recommendations---that standard retrieval pipelines flatten into fixed-size text chunks. We cast evidence access as a document image analysis problem: parse each page image into typed structural elements, then retrieve structure-aware evidence units that follow the document's own layout (sections, table rows, flowchart paths, graded recommendations), each keeping its structural context so a result points to a specific element rather than a page. On 26 clinical practice guidelines from 9 sources (3,619 pages, Chinese and English) with 199 evidence queries, structure-aware units rank the gold element first under BM25, dense, and hybrid retrieval (hybrid Element Hit@1 of 0.382), with a significant element-level ranking gain over per-element OCR text (MRR_e +0.107, p=0.002; the Hit@5 gain is directional, p=0.17), while matching page-level recall (Page Hit@5 0.879 vs. 0.889, p=0.75) at 3.8x less context and clearly outperforming a ColPali visual-RAG baseline (PH@5 0.497).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xingyu Lin, Dehui Du. 2026-09-27. From PDF to Evidence: Structure-Aware Retrieval for Clinical Practice Guidelines. https://arxiv.org/abs/2609.33447
Cite the original work for its findings. Save a collection to share your selection of sources.