Abstract / Summary
Background: Comprehensive evaluations of lung nodule detection methods remain limited, particularly regarding data from clinical routine, novel photon-counting CT (PCCT) technology, and the comparison of academic and commercial models. Materials and Methods: This retrospective study collected 1860 routine PCCT scans acquired at Hannover Medical School from 2021 to 2024. We compared four deep-learning-based lung nodule detection models, which were developed for conventional CT, namely the publicly available TotalSegmentator, nnDetection, and RadYOLO methods, as well as a commercial lung computer-aided detection (CAD) system. We evaluated the inter-model agreement in an unannotated cohort (n=1699), sensitivity in a smaller cohort of 25 manually annotated scans, and false positives as detections in 136 report-negative scans. Results: In the unannotated cohort, the commercial CAD system detected 7071, TotalSegmentator 4264, nnDetection 5254, and RadYOLO 3122 nodules. Overall, the inter-model agreement was low, and the models produced substantially different candidate nodules. Approximately 1300 nodules were detected by all models, representing 18.7% - 42.8% of each model's detections. 15.4% - 37.4% of a model's detections did not match any other model. The models showed moderate sensitivity (between 0.71, 95%-CI [0.64, 0.77], and 0.83, 95%-CI [0.77, 0.88]), and TotalSegmentator and RadYOLO produced few false positives per scan (both 0.85, 95%-CI [0.70, 1.02]). Conclusion: The lung nodules detected by four deep learning models showed high variation and a lack of consensus on clinical PCCT data. These findings highlight the importance of model benchmarking and local validation when deploying models in real-world clinical workflows.