Abstract / Summary
Abstract Large language models require reproducible outputs for systematic evaluation in tumor boards. We developed a sarcoma tumor-board simulator using Claude Opus 4.5 at temperature 0.0, schema-enforced JSON output, and prompt-specified semantic and conditional-logic rules. The framework was evaluated on 51 de-identified cases across 154 runs. Reproducibility was defined as agreement across all runs of a case per decision code. In the 41 cases not used during development, reproducibility was 94.2% (cluster-bootstrap 95% CI 91.9–96.1); in the 10 development cases, 93.3% (95% CI 88.6–97.6). Pooled, it was 94.0% (95% CI 92.0–95.8); 17 cases agreed on all 21 codes. Estimates are model-specific. A post hoc ablation on 15 cases compared the framework with and without these rules: 94.6% versus 82.5% (difference 12.1 points; 95% CI 7.6–17.1). The improvement was driven mainly by exact wording agreement in three free-text fields (+ 60.0 points), whereas the 18 categorical and numeric codes differed by 4.1 points (95% CI − 0.7 to 9.3). Without any schema, a decision was stated for only 37.6% of instances versus 100% under schema enforcement; seven codes were never addressed. Single-reviewer screening detected no fabricated numbers, references, or measurements. Clinical correctness was not assessed. Consistent recommendations may still be incorrect.