Abstract / Summary
Abstract Real-world fundus photographs often contain multiple co-existing ophthalmic signs, yet automated analysis is commonly framed as binary screening or single-disease classification. We introduce FundusSigns-47, an institutional multi-label dataset of 7,111 color fundus images annotated with 47 clinically observable signs. The dataset exhibits fine-grained appearance, severe long-tail imbalance, and diverse label combinations. We also propose Selective Candidate-Verified Prompt Tuning (SCVPT). A prompt-tuned SigLIP2 ViT-g model generates candidate logits, while a learned InceptionV3 residual verifier supplies complementary visual evidence through selective class-wise routing. On the fixed test set, SCVPT achieves 57.87% mean average precision, 56.15% macro-F1, 65.82% micro-F1, 92.40% macro area under the curve, and 62.56% sample-F1. Mechanistic and fusion controls indicate that residual visual verification, rather than candidate gating alone, provides the principal improvement. An independent SigLIP2 and InceptionV3 ensemble yields higher ranking metrics, whereas SCVPT yields higher macro-F1, recall, and sample-F1. These findings support residual verification as a balanced operating strategy for multi-label ophthalmic sign recognition and motivate external validation.