Abstract / Summary
Background To develop and validate a Vision Transformer-based model for semi-quantitative grading of synovitis on musculoskeletal ultrasound and to assess its potential clinical utility. The study also examined whether transformer-based modeling offers advantages over conventional CNNs in ordinal ultrasound grading.Patients This retrospective diagnostic study included 312 patients, 624 joints, and 1,248 ultrasound images, which were divided at the patient level into training, validation, internal test, and external test cohorts. A Swin Transformer–based ordinal grading model was trained on gray-scale and power Doppler ultrasound data and compared with CNN baselines using internal validation, external testing, stratified subgroup analysis, interpretability assessment, and a reader-assistance evaluation.Results The Vision Transformer achieved the best overall performance, with an internal-test accuracy of 0.809, macro F1 score of 0.806, weighted kappa of 0.842, and AUC of 0.929, while evaluation on an external test set, curated from a non-overlapping period at the same institution, remained acceptable with an accuracy of 0.771 and AUC of 0.904. Most errors occurred between adjacent grades, multimodal GSUS plus PDUS input outperformed single-modality models, and AI assistance improved junior-reader agreement and efficiency.Conclusions This preliminary study confirms that Vision Transformer can achieve accurate and clinically applicable synovitis grading on musculoskeletal ultrasound, showing better robustness than traditional CNNs. This method may help standardize and streamline ultrasound evaluation in clinical practice and multicenter studies.