arXiv · 2609.25651
SurgGraph: Quantitative Laparoscopic Video Understanding via Geometry-Grounded Scene Graphs
Abstract
Surgical videos are a primary resource for teaching trainees anatomy, tool usage, and procedural skills. Yet learning from them at scale requires systems that understand surgical scenes. Existing approaches fall short: vision-language models lack fine-grained domain reasoning, task-specific models do not generalize, and prior scene graphs omit clinically meaningful detail. We present SurgGraph, a training-free pipeline that generates quantitative scene graphs from surgical videos. Operating on segmentation masks and depth maps, SurgGraph encodes each clinically meaningful relation (attachment, occlusion, separation, tool actions) as a tuple whose numeric value quantifies the relation's extent over time. Technical evaluations show more precise scene understanding than state-of-the-art surgical VLM baselines. We then build SurgGraphQA, a proof-of-concept learning application that retrieves meaningful and boundary-case exemplars and generates visual explanations and feedback. A study with 17 medical students and 2 resident surgeons shows significant learning gains, demonstrating its educational value.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jingying Wang, Rosiana Natalie, Marquise D Singleterry, Filippos Bellos, Brian George, Gurjit Sandhu, Jason J Corso, Anhong Guo, Vitaliy Popov, Xu Wang. 2026-09-22. SurgGraph: Quantitative Laparoscopic Video Understanding via Geometry-Grounded Scene Graphs. https://arxiv.org/abs/2609.25651
Cite the original work for its findings. Save a collection to share your selection of sources.