What determines performance in automated cervical vertebral maturation staging? A unified analysis of input representation, output formulation, and landmark quality

Citations

WEB OF SCIENCE

0

초록

This study examined how input representation (image-based vs. landmark-based), output formulation (scalar vs. ordinal), and landmark quality (ground-truth vs. predicted) influence automated cervical vertebral maturation staging on lateral cephalograms, including the effect of training-deployment input mismatch. A total of 1750 lateral cephalograms were retrospectively collected. Two orthodontists independently assigned six stages and resolved disagreements by consensus. A test set of 150 images was held out, and 5-fold cross-validation was applied to the remaining 1,600 images. An image-based convolutional neural network used cervical images, whereas a landmark-based multilayer perceptron used landmark coordinates and derived measurements from either ground-truth or predicted landmarks. Scalar and ordinal output formulations were compared. Performance was assessed using 6-stage accuracy, +/- 1-stage accuracy, 3-stage accuracy, and weighted Cohen's kappa. In the image-based model, scalar and ordinal outputs showed similar performance (6-stage accuracy: 0.673 vs. 0.677; weighted Cohen's kappa: 0.912 vs. 0.906). In the landmark-based model trained and evaluated with ground-truth landmarks (GT -> GT), ordinal output yielded numerically higher 6-stage accuracy than scalar output (0.687 vs. 0.627), whereas weighted Cohen's kappa was similar (0.918 vs. 0.920) and no consistent ordinal advantage was observed across the other metrics. Six-stage accuracy decreased when predicted landmarks were evaluated using a ground-truth-trained model (Pred -> GT: 0.420-0.453), with partial numerical recovery when both training and evaluation used predicted landmarks (Pred -> Pred: 0.453-0.460). Most errors occurred between adjacent stages (+/- 1-stage accuracy: 0.820-0.967). Within the evaluated dataset and model architectures, landmark quality was associated with the largest observed performance differences in the landmark-based pipeline. The effects of scalar versus ordinal output were comparatively small and metric-dependent. Automated CVM staging should therefore be considered decision-support information rather than a replacement for expert assessment.

키워드

Cervical vertebral maturation; Lateral cephalogram; Skeletal maturity; Artificial intelligence; Deep learning; HAND-WRIST; CVM METHOD; RELIABILITY
제목
What determines performance in automated cervical vertebral maturation staging? A unified analysis of input representation, output formulation, and landmark quality
저자
Lim, Bang Hyun; LEE,, Soo Young; Zou, Bingshuang; Jung, Seok-Ki
DOI
10.1038/s41598-026-66307-5
발행일
2026-08
유형
Article
저널명
Scientific Reports
권
16
호
1