Abstract / Summary
Background: Amplicon-based next-generation sequencing is a cornerstone of viral genomic surveillance, but it often produces datasets with highly uneven read coverage due to PCR amplification bias. This leads to excessive data generation and substantial computational overhead for downstream bioinformatic analysis. Results: We evaluated the efficacy of two data reduction strategies on 220 SARS-CoV-2 clinical specimens and 12 wastewater samples: random subsampling with SeqKit and read depth normalization with BBNorm. FASTQ files were reduced to target depths from 10X to 10,000X to compare genome coverage, lineage assignment concordance, and computational resource usage. Both strategies markedly reduced total processing time. For clinical isolates, read depth normalization corrected for amplification bias resulting in highly concordant lineage assignments at a target depth of just 200X. Random subsampling preserved inherent bias but produced reliable assignments at depths over 500X. Reducing target depths below 500X in complex wastewater mixtures severely degraded variant deconvolution accuracy and heavily skewed lineage abundance profiles for both methods. Conclusion: Both random subsampling and read depth normalization successfully reduce computational overhead while preserving key clinical sequencing metrics at appropriate reduction settings. The optimal method depends on the analytical goal, ranging from computationally lightweight random subsampling for rapid lineage assignment, to read depth normalization for generating high-quality consensus sequences. However, aggressive data reduction should be avoided in environmental surveillance, where maintaining a high baseline sequencing depth is critical for accurate mixed-variant deconvolution.