Abstract / Summary
Purpose: Volumetric PET biomarkers carry substantial prognostic value in head and neck squamous cell carcinoma (HNSCC), yet their clinical implementation is hampered by segmentation heterogeneity. A fundamental and unresolved question is which segmentation method is most suitable for routine clinical use, both for a reader experienced in PET segmentation and for a reader without PET segmentation experience, such as an ENT surgeon or a radiation oncologist. This study addressed this question by assessing the intra- and interobserver reproducibility of four PET segmentation methods in MIM Maestro Software Version 7.4.2, identifying tumour-related determinants of variability, and examining concordance between PET-derived metabolic tumour volume (MTV) and CT-based gross tumour volume (GTV). Methods: Two readers with contrasting PET segmentation experience, a nuclear medicine physician experienced in PET segmentation (R1) and an ENT surgeon without prior PET segmentation training (R2), applied four methods to the primary tumours of 154 patients with HNSCC (oral cavity, oropharynx, hypopharynx or larynx) treated with definitive (chemo)radiotherapy between 2010 and 2021: fixed absolute thresholds at SUV 3.0 and 4.0, a relative threshold at 41% of the maximum standardised uptake value (SUVmax), and a gradient edge-detection algorithm (PET Edge+). R1 segmented twice with a washout of at least 4 weeks; R2 performed one session and delineated CT-GTV in the 85 patients whose PET/CT-derived CT served directly as the radiotherapy planning scan. The second session of R1 served as the reference for all interobserver and MTV–GTV analyses. Agreement was assessed by ICC (two-way mixed model), coefficient of variation, repeatability coefficient and Bland–Altman analysis; determinants of variability by a multivariable linear mixed-effects model with a random intercept per patient. Results: All four methods demonstrated excellent intraobserver (ICC 0.993 to >0.999) and interobserver (ICC 0.989–0.996) reproducibility. SUV3.0 generated significantly more relative error than SUV41% (β = +3.1%; 95% CI 0.89, 5.3). PET Edge+ was the only method for which a significant method × reader interaction was observed in this reader pair: the reader without PET segmentation experience generated more error than the reader experienced in PET segmentation (β = +7.1%; 95% CI 3.9, 10), an effect concentrated in smaller tumours (below approximately 15 mL); in the 85-patient GTV subset, where tumours were larger on average, the interobserver ICC for PET Edge+ was 0.997. GTV concordance was good for three of four methods (ICC 0.808–0.895); SUV41% showed only moderate concordance (ICC 0.572). Conclusions: All four methods provided excellent volumetric reproducibility within a single clinical software platform, but they differed in reader dependence and in agreement with CT-defined GTV. Rather than supporting a universal segmentation standard, these findings suggest that the method should be selected according to its intended application, pending multicentre confirmation.