Abstract / Summary
Background: Effective influenza molecular epidemiology depends on curated, temporally representative sequence datasets. The transition from raw GISAID or NCBI/GenBank batch downloads to analysis-ready FASTA often requires command-line pipeline configuration or custom scripting, which may be inaccessible to some field and clinical virologists. To our knowledge, a targeted, non-systematic comparison of widely used viral-genome tools did not identify a single browser-delivered workflow combining supported GISAID-style and NCBI/GenBank-compatible header parsing, interactive human-in-the-loop temporal subsampling, exact-sequence identity-group analysis, and six selectable language catalogues.
Methods: We developed VirSift (Viral Sequence Intelligence and Filtering Toolkit), a seven-page Streamlit web application implementing: (1) a custom parser handling five supported GISAID-style and NCBI/GenBank-compatible FASTA header variants with host inference and species normalization; (2) vectorized quality filtering and MD5-based identification of exact-sequence redundancy; (3) six documented human-in-the-loop (HITL) temporal sampling strategies with an operational date-span recommendation, plus a separate multi-attribute diversity filter for balanced host/location/clade subsampling; (4) ten Plotly visualization types, including a configurable N-level Sankey flow diagram and an exact-sequence persistence matrix; and (5) structured export to FASTA, CSV, and segment-organized ZIP bundles. The interface exposes six structurally aligned 921-key language catalogues. English and Russian contain complete or near-complete native-language values (100% and 99.0% respectively); French, Spanish, Arabic, and Chinese now also contain predominantly native-language values (94.1%, 98.0%, 98.9%, and 98.6% respectively) following a subsequent full-catalogue translation pass, with a small residual of proper nouns, format names, and language-specific cognates remaining in English. This tool-description paper documents implemented functions and includes a limited H3N2 demonstration and multi-file round-trip check; comprehensive analytical, usability, performance, and multilingual validation is reserved for separate work.
Results: Applied to 59 H3N2 haemagglutinin sequences from Novosibirsk, Russia, spanning October 2021 to April 2025, VirSift identified 12 records redundant by exact nucleotide identity (20.3% of records), retaining 47 unique sequences across 9 calendar months. The observed 1,279-day collection-date span was assigned to the software-defined long-span category (labelled "Endemic" in the version 1.0 interface), for which Custom Checkpoints was recommended. The Molecular Timeline reported 47 exact-sequence identity groups. Clade 3C.2a1b.2a.2a.3a.1 accounted for 69.5% of records. Multi-file ingestion and metadata-based split export were verified by reconciliation of record counts across five source files. All described operations were completed without programming.