Abstract / Summary
Foundation models for colour fundus photography (CFP) are candidate visual encoders for multimodal medical AI, but existing models are large and are assessed under fine-tuning rather than the frozen-encoder conditions such systems impose. We present RFM (Retinal Foundation Model), a 64-million-parameter Vision Transformer initialised from DINOv3-Base and domain-adapted with the LeJEPA joint-embedding predictive self-supervised objective with domain-specific pretext tasks, on 4.3 million CFPs from the North East London Diabetic Eye Screening Programme. A single linear head per task on one frozen backbone measures representation generalisability, not fine-tuning capacity. Across six external datasets spanning diabetic retinopathy (DR) grading, glaucoma detection, ten UK Biobank biomarker phenotypes and 3-year incidence of myocardial infarction, ischaemic stroke and Parkinson's disease, RFM attains the lowest mean absolute error of any frozen model on all seven regression phenotypes and a higher area under the receiver operating characteristic curve (AUROC) than any RETFound variant on all eight matched classification benchmarks. It also attains the highest AUROC of any model evaluated on four of the five ophthalmic benchmarks, with 4.8x fewer parameters than RETFound-MAE and RETFound-DINOv2. Applied to UK Biobank without retraining, RFM exceeds the best published retinal foundation-model AUROC for 3-year myocardial infarction and ischaemic stroke; the Parkinson's endpoint (11 cases) is inconclusive. DINOv3-Base without retinal pre-training, read the same way, stays competitive on DR grading but falls clearly behind on glaucoma, and the feature-extraction convention alone is worth up to 0.154 AUROC. One lightweight backbone yields generalised representations serving ophthalmic grading, biomarker regression and incidence prediction.