Benchmarking deep learning pipelines for surgical instrument segmentation in endoscopic spine surgery: cross-dataset evaluation under deployment-realistic conditions

  • Mun, Bong-Su; 
  • Jang, Seok-In; 
  • Park, Sang-Min; 
  • Park, Jiwon; 
  • Park, Hyun-Jin; 
  • 외 2명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Purpose Automated surgical instrument segmentation is a prerequisite for AI-assisted guidance in endoscopic spine surgery. Deployment-realistic comparisons of foundation-model and conventional deep learning architectures - accounting for the detector dependency of bbox-prompted models - remain lacking.Methods Seven deep learning pipelines were evaluated for binary instrument segmentation: U-Net, Attention U-Net, U-Net++, SegFormer-B0, DeepLabV3+, nnU-Net, and a fully automatic two-stage foundation-model pipeline (MedSAM2-Tiny segmentor + YOLOv8-nano detector). Strict patient-level data splitting was applied across two independent datasets (SNUBH biportal, n = 56 patients; SEA uniportal, n = 60 patients) with three-tier testing (internal, temporal external, cross-dataset external) and two training variants (biportal-only; combined biportal-uniportal).Results With biportal-only training, the fully automatic foundation-model pipeline ranked last on cross-dataset external testing (DSC 0.776; 0.737 under a strict empty-mask convention), about 9% points below DeepLabV3+ (0.865). The same MedSAM2-Tiny segmentor reached DSC 0.946 with ground-truth bounding-box prompts, leaving a 16.9-point oracle-to-deployment gap driven by detector domain mismatch: 61% missed detections and 39% box-localization error. With combined training, all seven pipelines exceeded DSC 0.94; nnU-Net was highest (0.978), and the foundation-model pipeline (0.962) was not significantly different from SegFormer-B0 (0.960) or DeepLabV3+ (0.958) after Bonferroni correction, although the small external sample does not establish equivalence. The improvement was driven primarily by detector rather than segmentor learning, and all seven main pipelines except the nnU-Net 5-fold ensemble met the 30 frames-per-second real-time threshold.Conclusions Foundation-model two-stage pipeline performance depends critically on detector training data coverage; once the segmentor approaches its oracle ceiling, detector domain coverage becomes the limiting factor for fully automatic deployment. When combined training is feasible, foundation-model and conventional architectures converge to comparable accuracy, although these rankings are specific to the datasets studied and require validation on independent cohorts. Reporting deployment DSC alongside the ground-truth bounding-box upper bound is recommended for future surgical-segmentation benchmarks.

키워드

Endoscopic spine surgery; Instrument segmentation; Deep learning; Foundation model; Object detection; Domain generalization; Artificial intelligence; MODELS
제목
Benchmarking deep learning pipelines for surgical instrument segmentation in endoscopic spine surgery: cross-dataset evaluation under deployment-realistic conditions
저자
Mun, Bong-Su; Jang, Seok-In; Park, Sang-Min; Park, Jiwon; Park, Hyun-Jin; Kim, Ho-Joong; Yeom, Jin S.
DOI
10.1007/s00586-026-10162-5
발행일
2026-07
유형
Article; Early Access
저널명
European Spine Journal