Abstract / Summary
Although largely preventable, cardiovascular disease (CVD) is the leading cause of death worldwide. Synthetic data generation (SDG) of health data can facilitate data sharing for research and development of personalised cardiovascular disease prevention tools. Current state-of-the-art SDG methods are limited to cross-sectional data or focused on short-term or in-patient EHR applications lacking benchmarks for generalisability. Since prevention is a long-term longitudinal objective, these limitations must be addressed. We evaluated two tabular generative adversarial networks, TGAN and CTGAN, for privacy and fidelity on a structured 6-year longitudinal trial dataset within CVD prevention. We compared performance with a rudimentary perturbation method. In addition, we performed a variable-wise evaluation and an age-stratified analysis of the higher performing model to assess preservation of clinically relevant characteristics. Overall, CTGAN showed better fidelity than TGAN. The models performed similarly within privacy, showing better results than the perturbation method. CTGAN had lower performance in comparison to the perturbation method within numerical fidelity of continuous variables and overall temporal fidelity. Furthermore, CTGAN's performance was lower for younger age groups than older. CTGAN shows better than or equal performance within fidelity and privacy than TGAN and the perturbation method on a long-term longitudinal dataset within CVD prevention. To the best of our knowledge, this is the first study evaluating CTGAN on a longitudinal dataset and our findings suggest that handling of temporal characteristics and imbalanced subgroups need further refinement.