Abstract / Summary
Abstract Objectives Large language models (LLMs) can automate the conversion of radiologic free text into structured reports, but multidisciplinary reader preferences for the resulting output remain insufficiently studied. This prospective, two-phase, fixed-sequence within-participant study evaluated reader assessments of LLM-structured versus conventional free-text oncologic CT reports. Material & Methods A paired convenience sample of radiologists, radiology technologists, oncologists, and medical students was recruited between May and November 2025. In Phase 1, participants rated 5 free-text oncologic CT reports for readability and comprehension (3 items), oncologic information quality (5 items), and clinical utility (4 items) on a 9-point Likert scale. After a washout of at least 4 weeks, participants rated the LLM-structured version generated by Llama-4-Scout-17B-16E-Instruct. Results Sixty-six paired participants (24 medical students, 22 radiologists, 12 radiology technologists, 8 oncologists) were included (mean [standard deviation] age, 30.5 [7.6] years; 37 [56.1%] female). Structured reports showed significantly higher odds of agreement for 11 of 12 items (odds ratio [OR] range, 1.70–4.81), with strongest effects for clinical utility (OR, 4.81; 95% CI, 3.24–7.15), information findability (OR, 4.31; 95% CI, 3.12–5.96), recommendation clarity (OR, 4.14; 95% CI, 2.95–5.83), and patient communication suitability (OR, 3.45; 95% CI, 2.59–4.60). Radiology technologists showed significantly higher ratings for all 12 items, radiologists for 6, students for 7, and oncologists for none. Conclusion In this single-center paired sample, LLM-structured oncologic CT reports received higher reader ratings than conventional free-text reports across most evaluated dimensions, with the largest differences observed for information findability and clinical utility. These findings support further evaluation of radiologist-supervised LLM-assisted report structuring, including formal source-to-output validation and randomized multicenter assessment, before clinical implementation.