Abstract / Summary
Abstract AI-oncology evidence is expanding too quickly for static manual surveillance. We developed OncoTagger, a rule-based abstract-level evidence-surveillance pipeline, and applied it to an English-language, open-access article corpus indexed in the Web of Science Core Collection and analyzed at the title-, abstract-, keyword-, and metadata level from 2019 to 2025. From 59,994 initial records, deduplication, year restriction, automated screening, and manual adjudication yielded 20,766 records. Prediction-stratum-weighted corpus-level estimates showed metric-detection accuracy of 92.3% (95% CI 88.9–95.2%), sensitivity of 89.3% (85.4–93.3%), and specificity of 98.2% (94.5–100.0%). Ordinal metric categories showed exact agreement of 73.6% (69.0–77.8%) for the weighted-category output and 76.8% (72.3–80.8%) for the composite-metric output, with linear weighted Cohen’s kappa of 0.588 (0.513–0.659) and 0.615 (0.537–0.689), respectively. Primary-task assignment showed moderate agreement with manual consensus (68.0% exact agreement, 95% CI 63.3–72.4%; Cohen’s kappa 0.508, 0.442–0.572), and a complete task-unassigned census identified systematic dictionary-coverage gaps. The resulting resource describes abstract-reported metric patterns, pipeline-derived task mix, geography, and an exploratory candidate translational-signal subset. It should be interpreted as reproducible aggregate surveillance infrastructure, not as a full census of the AI-oncology field, a validated article-level classifier, or a comparative evaluation of algorithmic performance.