Abstract / Summary
Genotype calling of sequencing data has improved over recent years thanks to technological and statistical breakthroughs. One such advance is the use of imputation-based genotyping methods in human genetics. For haploid and mostly non-recombining organisms, a natural equivalent is to leverage phylogenetic signal for imputation-based genotype calling. SARS-CoV-2 is one such case, and the use of improved genotyping methods in this species is particularly pertinent given the well-documented systematic problems with read alignment and assembly in rapidly evolving regions of its genome, motivating the need for statistically principled methods for genotype correction that complement existing community efforts to curate sequencing data. We present a statistical framework to combine phylogenetic prior probabilities with the information from aligned reads, and demonstrate its utility in correcting consensus calls made on short-read, viral sequencing data. We derive phylogenetically-informed posterior probabilities based on genome placement in a reference phylogeny, and read likelihoods calculated from machine-reported base quality scores. We further modify the prior at each site by estimating a maximum-likelihood scaling parameter jointly with reads at all other sites. We find that the phylogenetically-informed base calling method stays highly accurate across simulated depth and error conditions. In real datasets, the method vastly improves call rate while still maintaining similar accuracy as reference-based assembly methods. In problematic genomic regions, the phylogenetically-informed method confers both accuracy and call rate gains. Accuracy and call rate do not stratify with sequencing center and primer scheme used, suggesting broad applicability across sequencing technologies and potentially diverging sequences. Our results in real and simulated data demonstrate that phylogenetically-informed genotyping correction results in more complete sequences, achieving a higher call rate without compromising accuracy. In cases of low mean depth, sparse coverage, or difficult-to-sequence regions that may still retain a low call rate, the method improves the accuracy of genotype calls.