Abstract / Summary
Abstract Background Cell line experiments remain one of the most widely used tools in preclinical cancer research. As prostate cancer cell lines underpin much of this research, evaluating their clinical relevance is essential to improve the translation of laboratory findings into patient care. Here, we assess the clinical relevance and inclusivity of 46 prostate cancer cell lines using an integrated clinical and multiomics approach. Methods Clinical and molecular data from more than 10,000 patient cases across 24 studies in nine countries were harmonised using publicly available databases and peer-reviewed literature. Analyses included mutations, copy number alterations, structural variants, microsatellite instability, gene expression and protein expression. Clinical characteristics, including age, geographical location, ethnicity, race, cancer type, Gleason grade and metastatic stage, were evaluated to determine how well prostate cancer cell lines represent the broader patient population. Statistical comparisons were normalised and cross-validated using bioinformatics approaches. Results Here we show that prostate cancer cell lines do not broadly represent the wider patient population and that the patients from whom these models were derived constitute a demographically and clinically biased subset. Our integrated analyses identify important gaps in the molecular and clinical diversity represented by commonly used cell lines and demonstrate differences in their suitability for modelling specific disease subtypes. Conclusions Our findings highlight limitations in the representativeness of widely used prostate cancer cell lines and emphasise the importance of selecting models according to their clinical and molecular relevance. The proposed data-driven framework provides a practical approach for selecting more representative cell lines and has the potential to improve the translation of preclinical prostate cancer research into clinical practice.