Abstract / Summary
Patients with primary immunodeficiency (PI) face prolonged diagnostic delays and may increasingly turn to large language models (LLMs) to interpret their symptoms during this period. We evaluated whether an LLM could recognize PI from symptom descriptions derived from interviews with 21 PI patients. In prior work, we showed that GPT-4o identified PI in 96% of cases when prompted with physician-written patient histories (Rider et al., 2025). Here, when prompted with symptom descriptions in patients' own words, GPT-5.2 identified PI in only 7 cases (33%), although it suggested general immune system concerns in 17 cases (81%). This expert-versus-patient gap may indicate that LLMs are sensitive to the language and framing of symptom descriptions, performing substantially worse when patients describe their own symptoms in everyday language than when clinicians summarize patient histories in structured medical terms. This study underscores the need to carefully evaluate how LLMs are used in patient-facing applications.