Artificial Intelligence-Assisted Acoustic Voice Analysis for Differentiating Benign and Malignant Laryngeal Lesions: A Diagnostic Validation Study
Abstract
Background: Benign and malignant laryngeal lesions frequently present with dysphonia, making clinical differentiation based on voice characteristics alone challenging. Acoustic voice analysis provides objective measurements of vocal function, while artificial intelligence (AI) may identify complex acoustic patterns that are difficult to recognize using conventional parameters.
Objective: To evaluate the diagnostic performance of an AI-assisted acoustic voice analysis system in differentiating benign from malignant laryngeal lesions using clinically characterized voice recordings.
Materials and Methods: A prospective diagnostic validation study was conducted among adults with laryngeal lesions confirmed by endoscopic and histopathological assessment with ethics committee approval and written informed consent. Standardized sustained-vowel and connected-speech recordings were subjected to acoustic analysis. Fundamental frequency, jitter, shimmer, harmonics-to-noise ratio (HNR), cepstral peak prominence (CPP), spectral parameters and Mel-frequency cepstral coefficients were extracted. Machine-learning classifiers including logistic regression, random forest and extreme gradient boosting (XGBoost) .Histopathological diagnosis was the reference standard. Diagnostic performance was assessed using sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), F1 score and area under the receiver operating characteristic curve (AUC), with 95% confidence intervals (CIs) where estimable from the available model outputs.
Results: Among 120 participants, 78 had benign and 42 had malignant lesions. The XGBoost model yielded an AUC of 0.91. Based on the reported confusion matrix, sensitivity was 88.1% (95% CI 75.0–94.8%), specificity 84.6% (95% CI 75.0–91.0%), PPV 75.5% (95% CI 61.9–85.4%), NPV 93.0% (95% CI 84.6–97.0%), accuracy 85.8% (95% CI 78.5–91.0%) and F1 score 0.81. Random forest and logistic regression yielded AUCs of 0.86 and 0.82, respectively.
Conclusion: AI-assisted acoustic voice analysis has potential as a non-invasive adjunctive approach for differentiating benign and malignant laryngeal lesions. The reported performance reflects internal validation in a single-centre cohort and should not be regarded as evidence of clinical diagnostic effectiveness until confirmed in an external cohort. Its role should presently be considered complementary to laryngoscopic examination and histopathological diagnosis. Larger prospective, multicentre and externally validated datasets are required before routine clinical implementation.
Keywords:
Artificial intelligence, acoustic voice analysis, laryngeal lesion, laryngeal cancer, dysphonia, machine learning, voice biomarker, diagnostic accuracyDOI
https://doi.org/10.37022/wjcmpr.v8i3.445References
1. Maryn Y, Corthals P, Van Cauwenberge P, Roy N, De Bodt M. Toward improved ecological validity in the acoustic measurement of overall voice quality: combining continuous speech and sustained vowels. J Voice. 2010;24(5):540-555.
2. Baken RJ, Orlikoff RF. Clinical Measurement of Speaking and Voice. 2nd ed. San Diego: Singular Thomson Learning; 2000.
3. Hippargekar P, Bhise S, Kothule S, Shelke S. Acoustic Voice Analysis of Normal and Pathological Voices in Indian Population Using Praat Software. Indian J Otolaryngol Head Neck Surg. 2022;74(Suppl 3):5069-5074.
4. Deliyski DD, Shaw HS, Evans MK. Adverse effects of environmental noise on acoustic voice analysis. J Voice. 2005;19(1):15-28.
5. Kim HB, Song J, Park S, Lee YO. Classification of laryngeal diseases including laryngeal cancer, benign mucosal disease, and vocal cord paralysis by artificial intelligence using voice analysis. Sci Rep. 2024;14:9297.
6. Godino-Llorente JI, Gómez-Vilda P, Blanco-Velasco M. Dimensionality reduction of a pathological voice quality assessment system based on Gaussian mixture models and short-term cepstral parameters. IEEE Trans Biomed Eng. 2006;53(10):1943-1953.
7. Oates J. Auditory-perceptual evaluation of disordered voice quality: pros, cons and future directions. Folia Phoniatr Logop. 2009;61(1):49-56.
8. Ma K, Wang Y, Zhou Y, Chen L, Zhang T, Xu F, Peng X. Acoustic signatures of organic lesions and the role of artificial intelligence in voice disorder diagnostics. Digital Health. 2025;11:20552076251376264. doi:10.1177/20552076251376264.
9. Marrero-Gonzalez AR, Diemer TJ, Nguyen SA, Camilon TJM, Meenan K, O'Rourke A, et al. Application of artificial intelligence in laryngeal lesions: a systematic review and meta-analysis. Eur Arch Otorhinolaryngol. 2025;282:1543-1555.
10. Keung LC, Richardson K, Sharp Matheron D, Martel-Sauvageau V. A Comparison of Healthy and Disordered Voices Using Multi-Dimensional Voice Program, Praat, and TF32. J Voice. 2024;38(4):963.e23-963.e38.
11. Ozcelik Erdem R, Arbag H. Clinical utility of acoustic voice parameters and patient demographic variables in identifying vocal fold pathology and supporting referral for laryngeal evaluation. J Voice. Published online August 22, 2026. doi:10.1016/j.jvoice.2026.07.056.
12. Jenkins PD, Bedrick S, Karstens L, Hersh W, Bridge2AI-Voice Consortium, Dorr DA. From voice biomarkers to telemedicine screening: developing and evaluating a voice-based AI model for laryngeal lesion detection using the Bridge2AI-Voice dataset. Front Digit Health. 2026;8:1846369.
13. Boersma P, Weenink D. Praat: doing phonetics by computer [Computer program]. Version [insert version]. Amsterdam: University of Amsterdam; [year accessed]. Available from: https://www.praat.org.
14. Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825-2830.
15. Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York: ACM; 2016:785-794.
Published
Abstract Display: 0
PDF Downloads: 0 How to Cite
Issue
Section

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
