Abstract / Summary
Purpose: Early detection of glottic cancer is essential for better outcomes, but current diagnostic methods remain largely invasive, highlighting the need for reliable non-invasive alternatives. Acoustic voice signals offer a non-invasive screening input, but their clinical value depends on whether signal-processing features and learned models can distinguish malignancy from acoustically similar benign lesions rather than merely separate pathological from healthy voices. Despite growing research on acoustic analysis combined with machine learning (ML) techniques, no systematic synthesis of this evidence currently exists. Therefore, the present review aimed to identify acoustic features that differentiate early glottic malignancies from both benign vocal fold lesions and normal vocal function using various ML techniques. Method: Following PRISMA 2020 guidelines, a comprehensive literature search was conducted across Scopus, EBSCO, Ovid (Embase), and PubMed databases through November 2025. Eighteen studies met inclusion criteria, and PROBAST was used for quality appraisal of included studies. Data extraction focused on acoustic feature categories, ML techniques, classification performance metrics, and validation strategies. Results: Nine acoustic feature categories emerged, with cepstral coefficients (particularly MFCCs) most prevalent across 55.6% of studies. Binary pathological-versus-normal classification achieved consistently high accuracy (96-98.3%) across both traditional machine learning and deep learning approaches. However, malignant-versus-benign discrimination demonstrated substantially lower performance (81-87.88% accuracy, AUC 0.631-0.91), with voice-only models proving insufficient. Multimodal models incorporating clinical variables or laryngoscopic images achieved higher performance, but their results cannot be attributed to voice features alone. All 17 prediction studies were judged at high overall risk of bias, predominantly because of small effective sample sizes, model-selection optimism, possible data leakage, and limited independent validation. Conclusions: ML can detect broad voice pathology, but current evidence is insufficient to support stand-alone acoustic screening for early glottic cancer. Future studies require clearly defined cancer-specific targets, patient-level analysis, standardized acquisition, nested model development, calibration, and independent multicentervalidation.