Abstract / Summary
The Mycobacterium avium complex (MAC) comprises genetically close nontuberculous mycobacteria responsible for human and animal infections worldwide. Increasing availability of whole-genome sequencing (WGS) data offers new opportunities for genomic surveillance and lineage tracking of these pathogens. However, the taxonomy of the MAC remains unstable, with inconsistent species and subspecies boundaries that hinder reproducible genome-based identification complicating diagnostics and antimicrobial management. Furthermore, a major barrier to the clinical and epidemiological integration of WGS is the current lack of standardized, accessible identification tools capable of translating complex genomic similarity into harmonized taxonomic assignments. We performed a comprehensive in silico analysis of 2,737 MAC genomes using pairwise average nucleotide identity (ANI). Species and subspecies boundaries were statistically modeled at 95% and 98% ANI thresholds. A consensus genome-based framework combining FastANI and dRep clustering was developed. Machine learning classifiers based on k-mer frequency profiles were trained to reproduce genomic delineations. Fifteen species-level clusters were identified, with 99.9% concordance to GTDB-Tk reference classifications. Within M. avium , subsp. avium , subsp. hominissuis , and subsp. silvaticum formed a single genomic continuum, while the IS900-positive M. avium subsp. avium lineage paratuberculosis remained distinct but did not reach subspecies level divergence. Within M. intracellulare , subsp. chimaera showed deep internal structure, supporting recognition of M. intracellulare subsp. yongonense , whereas genomes assigned to M. paraintracellulare clustered within M. intracellulare subsp. intracellulare . Several nominal species ( M. bouchedurhonense , M. timonense ) lacked genomic coherence. The MAC displayed an open pangenome (191,904 gene clusters; 0.8% universal core). Machine learning classifiers trained on k-mer frequency profiles accurately reproduced whole-genome delineations, achieving near-perfect performances (Cohen’s kappa = 0.994). We provide a standardized, genome-based framework for classification of the MAC and implement it in MAC-Explorer (https://mycobacteriaexplorer.fr), an open-access web tool for automated MAC identification. By integrating k-mer -based classification and insertion sequence screening, MAC-Explorer delivers reproducible and harmonized species- and lineage-level assignments, providing a robust foundation for genome-based identification and lineage tracking of MAC isolates.