Abstract / Summary
Abstract Multimodal artificial intelligence (AI) could support audiology reporting, but an end-to-end score can obscure whether errors arise from image transcription, specialty-rule execution, or communication. We evaluated these as separate modules in 155 de-identified outpatient records representing 151 patients. In Module 1, schema-only and audiologist-skill prompts were compared in 50 development records; the locked skill was then tested with a Codex GPT5.5 agentic workflow in 101 independent patients. A post hoc robustness analysis tested matched schema-only and skill-assisted direct-API conditions with DeepSeek-V4-Flash-Vision-Exp in the same 101 patients. The skill increased hearing-loss-type agreement from 85.0% to 94.0% and acoustic-reflex agreement from 70.4% to 98.4% in development. Codex validation yielded 1809/1820 (99.4%) exact numeric thresholds, but only 9/15 (60.0%) no-response entries were correct. In the DeepSeek robustness analysis, the skill increased type agreement from 37.6% to 70.3%, while exact numeric-threshold agreement remained 72.1% and no-response agreement remained 0%. In Module 2, AI workflows received adjudicated structured values, not Module 1 outputs. After two audiologists refined reports using 60 records, 91 independent patients were evaluated. All 25 rule-derived fields were correct in 91/91 Codex and 87/91 DeepSeek cases; both audiologists rated all Chinese reports accurate or basically accurate. These were different operational workflows, not a matched foundation-model comparison. The findings support prospective evaluation of modular, auditable clinician assistance, not autonomous diagnosis, end-to-end performance, or patient usability.