What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients

June 23, 2026 ยท Grace Period ยท ๐Ÿ› Interspeech 2026, ISCA, Sep 2026, Sydney, Australia

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Tuan Nguyen, Corinne Fredouille, Alain Ghio, Muriel Lalain, Virginie Woisard arXiv ID 2606.24949 Category cs.SD: Sound Cross-listed cs.AI, cs.NE Citations 0 Venue Interspeech 2026, ISCA, Sep 2026, Sydney, Australia
Abstract
This work investigates the interpretability of a Wav2Vec 2.0based speech intelligibility assessment model for oral and oropharyngeal cancer patients through canonical correlation analysis. By measuring the correlation between the model embeddings and eGeMAPS low-level descriptors (LLDs) as an interpretable reference, we analyze how acoustic information is encoded across the model layers. The analysis is conducted at two levels: individual LLDs layer-wise, and group-level: prosodic, spectral, and voice quality. Results show that the learned representations are most strongly correlated with spectral and prosodic features, with the first MFCC coefficient yielding the highest correlations across all layers. At the group level, spectral and prosodic groups achieve correlations of 0.77 and 0.71 respectively, while voice quality reaches 0.65. Beyond model interpretability, this work also offers practical guidance on acoustic feature selection for pathological speech assessment.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Sound