Abstract / Summary
Background: Next-generation sequencing (NGS) is vital for tracking viral evolution during outbreaks. Oxford Nanopore Technologies (ONT) offers long-read, cost-effective sequencing, but accurate sub-consensus variant detection remains challenging. LoFreq, a widely used caller for sub-consensus variant calling in haploid genomes, was designed for short reads. This study assesses and benchmarks LoFreq variant identification accuracy on long-read data and proposes a calibration method to improve accuracy.
Methods: We created truth sets using three SARS-CoV-2 spike gene plasmids (7179 bases, 100 SNVs) and whole Escherichia coli genomes. Libraries were sequenced using R9.4.1 and R10.4.1 flow cells. LoFreq's recall was benchmarked across chemistries and library sizes. We also introduced a method to adjust Phred quality scores for enhanced accuracy.
Results: LoFreq showed high sensitivity, detecting variants at frequencies as low as 0.1%, with best performance on R10.4.1. However, false discovery rates (FDR) were chemistry- and depth-dependent. Our Phred score calibration reduced false positives while maintaining recall, though it was less effective below 10% frequency and not suitable for structural variant detection.
Conclusion: LoFreq can detect sub-consensus variants in ONT data, but high FDR limits direct use. Our calibration method improves accuracy, offering a simple Phred score adjustment approach that reduces false discovery rates in ONT data.