Abstract / Summary
Abstract Objective Retinal thickness is an important structural biomarker for the assessment and monitoring of retinal diseases. In routine practice, thickness measurements are obtained from optical coherence tomography (OCT), which provides depth-resolved retinal imaging but may not be available in all clinical settings. This study investigates whether retinal thickness maps (RTMs) can be synthesized directly from widely available color fundus photographs (CFPs) using paired CFP-RTM data. Methods We propose an anatomically informed latent diffusion framework for cross-modal retinal topography synthesis. The framework builds on latent diffusion generative modeling [1] and uses a perceptually compressed latent representation learned with a variational autoencoder (VAE) [2]. To better preserve modality-relevant retinal structure during conditional generation, we incorporate a spatially weighted cross-attention (SW-CA) module that applies anatomically informed spatial weighting within the attention mechanism, inspired by attention-based conditioning strategies used in recent vision architectures [3, 4]. Model performance was evaluated using image-fidelity and perceptual metrics, including PSNR, SSIM, LPIPS, and FID [5–8], and compared with representative convolutional, generative adversarial, and transformer-based image synthesis approaches. Results Across quantitative evaluations, the proposed AC-LDM achieved a PSNR of 30.88 dB and an SSIM of 0.871. Perceptual similarity and distributional consistency were supported by lower LPIPS (0.105) and FID (29.5) scores relative to the evaluated baseline models. Conclusion These findings suggest that conditional latent diffusion combined with anatomically guided attention may provide a promising framework for synthesizing retinal thickness representations from fundus photography. This approach may support future research on cross-modal retinal imaging and computational ophthalmology, although clinical validation with thickness-specific error metrics remains necessary before translational use.