Abstract / Summary
Chest X-ray (CXR) imaging remains a widely used diagnostic modality for the detection of thoracic diseases, including pneumonia, tuberculosis, pleural effusion, lung cancer and cardiomegaly. Despite the extensive use of CXR examinations in clinical practice, the development of robust multi-label artificial intelligence (AI) systems remains constrained by limitations in dataset availability, annotation consistency and reproducible labeling methodologies. Existing public datasets frequently rely on heterogeneous labeling protocols, limited pathology coverage or annotations that inadequately account for linguistic negation and uncertainty within radiology reports. In this paper, we introduce MultiCaRe-Thorax, a curated thoracic benchmark dataset derived from the publicly available MultiCaRe repository, with a reproducible negation-aware natural language processing (NLP) framework for automated generation of multi-label thoracic disease annotations. Unlike existing CXR benchmarks that are distributed in their pre-curated form, MultiCaRe-Thorax provides a fully reproducible pipeline from raw case reports to curated dataset, establishing a new resource that is both scalable and transparent. Our pipeline performs thoracic case identification, image–report matching, duplicate removal, data cleaning, disease-specific label extraction and quality-control verification to produce a curated cohort comprising 5252 chest radiographs from 3415 unique patients. Sixteen clinically relevant thoracic pathologies were automatically annotated using a dictionary- and pattern-based NLP framework incorporating explicit negation detection and n-gram mining, achieving 95% accuracy against 500 manually reviewed reports (Cohen’s κ=0.89). Sixteen clinically relevant thoracic pathologies were automatically annotated using a dictionary- and pattern-based NLP framework incorporating explicit negation detection and n-gram mining, achieving 95% agreement against 500 manually reviewed reports. Five ImageNet-pretrained convolutional neural network (CNN) architectures (ResNet50, DenseNet121, EfficientNetB0, EfficientNetB7 and InceptionV3) with two hybrid ensemble configurations were evaluated under a patient-level data partitioning protocol. Among individual architectures, DenseNet121 attains the best AUC on tuberculosis (0.757) and pleural effusion (0.767) with roughly one-third the parameters of ResNet50, while ResNet50 attains the higher mean AUC across all sixteen pathologies (0.688 vs. 0.682) and on interstitial lung disease (0.769) and lung mass/cancer (0.777); we report this as an accuracy–efficiency trade-off rather than a single best architecture. The DenseNet121-InceptionV3 ensemble achieves the highest reported AUC of 0.821 for lung mass/cancer detection. We report 95% confidence intervals for the evaluated ensemble and remaining backbones and note substantial estimation uncertainty for the rarest classes (e.g., congenital heart disease, n=8 test cases). Grad-CAM visualizations were also employed to provide clinically interpretable explanations of model predictions.