Abstract / Summary
Large language models (LLMs) offer potential to alleviate the growing public health burden of eye disease; however, few models have been tailored for ophthalmology or evaluated against clinically relevant benchmarks and real-world workflows. Here, we present LEME, a suite of open-weight LLMs for ophthalmology developed via a two-stage framework: instruction tuning on 211,149 examples curated from authoritative sources, and reinforcement learning with 29,747 preference-labeled examples to enhance clinical reasoning and informativeness. LEME was evaluated under zero-shot settings using five curated benchmarks, covering question answering and patient-physician consultation. LEME outperformed all seven baselines (all p < 0.004), notably exceeding GPT-4o by 3.32% in average ROUGE-L. Furthermore, LEME was assessed on three downstream tasks using deidentified patient data. In answering patient queries, LEME received the highest overall ratings from attending clinicians; its completeness rating even surpassed that of the expert-written answers ( p = 0.015). For visual acuity extraction from clinical notes, LEME achieved the highest F1 scores, outperforming LLaMA-3 70B by 14.1% and Eye-LLaMA by 59.0%. Finally, in assessment and treatment plan generation, LEME achieved ratings approaching those of attending clinicians, demonstrating promise for ophthalmic decision support. All models, datasets, and code are made open to support further development and validation.