Abstract / Summary
Background Large-scale retrospective studies often require researchers to adjudicate outcomes or cohort eligibility from unstructured clinical records. Manual review can provide reliable labels but is difficult to scale. We developed and evaluated a consensus workflow using multiple large language models (LLMs) to automatically assign high-confidence outcome labels while deferring ambiguous cases for manual review. Methods We included 6,718 brain MRI reports from 5,856 patients. The expert-adjudicated reference set included 680 reports from 675 patients: 372 normal (age-appropriate without significant intracranial pathology), 90 white matter hyperintensity-only (WMH-only; isolated nonspecific WMHs potentially ambiguous for binary classification), and 218 abnormal (other intracranial pathology) reports. We evaluated proprietary-inclusive (Qwen3/Kimi-Linear/GPT-5) and open-source (Qwen3/Kimi-Linear/DeepSeek-R1) ensembles against individual LLMs. Ensembles assigned a normal or abnormal label only when all models agreed; all other reports were classified as uncertain and deferred for manual review. Individual LLMs assigned normal or abnormal predictions and classified WMH-only predictions as uncertain. In the full cohort, we compared the prevalence of active central nervous system (CNS)-related International Classification of Diseases, Tenth Revision (ICD-10) diagnoses across the normal, uncertain, and abnormal groups generated by each ensemble. Results In the expert-adjudicated reference set, both ensembles achieved 100% precision for automatically assigned normal and abnormal reports. GPT-5 was the strongest individual LLM but had lower precision than either ensemble, with normal and abnormal precision of 98.9% and 94.1%, respectively, and 17 expert-label mismatches versus none for the consensus ensembles. In the full cohort, 1.6% of reports classified as normal were associated with CNS-related ICD-10 diagnoses, compared with 29.7% of uncertain and 51.3% of abnormal reports for the proprietary-inclusive ensemble; corresponding rates were 1.6%, 31.1%, and 51.6% for the open-source ensemble. Conclusions Requiring unanimous agreement across multiple LLMs enabled high-precision automatic adjudication while directing ambiguous cases to manual review. This strategy may reduce the burden of manual adjudication and support scalable construction of high-purity research cohorts from unstructured clinical text.