A Surgical Foundation Model Reveals Task-Dependent Label Efficiency
Developing label-efficient models is a central challenge in surgical AI due to the high cost and scarcity of expert annotation. While self-supervised foundation models adapt well to new tasks with minimal data, how label efficiency varies across different surgical tasks remains largely unexplored. Here, we introduce SURGE, a surgical foundation model trained on SurgSpectrum-30M+, the largest pretraining dataset comprising over 30 million frames, with checkpoints released to enable further research. We systematically evaluate label efficiency across 5 task categories and 15 benchmarks. These range from temporal and spatial scene understanding to fine-grained reasoning tied to instrument-anatomy interactions and safety-critical maneuvers. SURGE outperforms prior state-of-the-art on all benchmarks, even surpassing task-specific models on complex reasoning tasks. Crucially, we reveal a task-dependent scaling behavior: while scene understanding tasks saturate with minimal supervision, fine-grained reasoning tasks continue improving with substantially larger annotation budgets, providing a blueprint for allocating expert effort in complex domains. Code: https://github.com/CAMMA-public/SURGE