Abstract / Summary
International travel can increase infection risk and spread antimicrobial resistance (AMR), yet clinical systems rarely capture travel history in a structured format. Extracting this information from free-text clinical documentation at scale could improve alerts for high-consequence infections and AMR surveillance. We developed an annotation framework covering key aspects of travel history, including destination, exposures, travel duration, and multiple temporal variables. The framework was used to annotate a corpus of 100 clinical case reports by five clinicians, yielding over 5,000 clinician annotations. Several LLMs were evaluated on the same extraction task, with their outputs compared against clinician consensus. Agreement among clinicians was high for straightforward variables such as presence of travel and destination, but dropped substantially for temporal variables such as time since return and symptom onset relative to return. LLMs showed a similar pattern. They performed well on high-level variables but struggled with the same temporal fields that challenged clinicians. Model errors were also concentrated in cases where clinicians disagreed. Across models, outputs were more sensitive to sampling temperature than to prompt wording. A smaller open-weight model performed similarly to substantially larger models on several variables but remained weaker on temporal fields. Overall, LLMs could extract some travel history variables, but reliable extraction of temporal information remained difficult.