| 1 |
On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models
Shunsuke Kando, Wataru Nakata, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 2 |
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS
Seymanur Akti, Alexander Waibel
|
|
cs.SD
|
0 |
2 months ago |
| 3 |
From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection
Jan Jasiński, Mateusz Barański, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 4 |
HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
Mateusz Barański, Jan Jasiński, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 5 |
Cross-lingual Retrieval-Augmented Classification for Dysarthria Severity Assessment
Taeyoung Jeong, Insung Lee, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 6 |
Learning to Evade: Adaptive Attacks on Audio Watermarking
Weikang Ding, Hanqing Guo, ... (+5 more)
|
|
cs.SD
|
0 |
2 months ago |
| 7 |
What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study
Yaozhong Kang, Jiang Wang, ... (+4 more)
|
|
cs.SD
|
0 |
2 months ago |
| 8 |
Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study
Tomoki Koriyama
|
|
cs.CL
|
0 |
2 months ago |
| 9 |
Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR
Enes Yavuz Ugan, Alexander Waibel
|
|
cs.CL
|
0 |
2 months ago |
| 10 |
Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems
Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka
|
|
cs.HC
|
0 |
2 months ago |
| 11 |
Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead
Muyang Du, Jason Roche, Junjie Lai
|
|
cs.SD
|
0 |
2 months ago |
| 12 |
Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement
Nasser-Eddine Monir, Paul Magron, Romain Serizel
|
|
cs.SD
|
0 |
2 months ago |
| 13 |
Post-Training Speech Enhancement Language Models with Perceptual Rewards
Frédéric Berdoz, Luca A. Lanzendörfer, ... (+2 more)
|
|
cs.LG
|
0 |
2 months ago |
| 14 |
Sexualised synthetic personas encode and amplify gendered power asymmetries through voice
Alice Ross, Ariadna Sanchez, ... (+3 more)
|
|
eess.AS
|
0 |
2 months ago |
| 15 |
An Evaluation Framework for Text-to-Speech Voice Reconstruction
Ariadna Sanchez, Christoph Minixhofer, ... (+4 more)
|
|
eess.AS
|
0 |
2 months ago |
| 16 |
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
Raphaël Bagat, Zhe Zhang, ... (+3 more)
|
|
cs.CL
|
0 |
2 months ago |
| 17 |
Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment
Xiaoliang Wu, Qiyang Sun, ... (+4 more)
|
|
cs.AI
|
0 |
2 months ago |
| 18 |
LISE : Listenable Interpretable Speaker Embeddings
Xiaoliang Wu, Chongxin Gan, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 19 |
Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
Tzu-Chieh Wei, Yi-Cheng Lin, ... (+5 more)
|
|
eess.AS
|
0 |
2 months ago |
| 20 |
LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations
Younghan Park, Hoyeon Lee, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 21 |
Imitation Learning for Elder-Facing Speech Synthesis
Dongrui Han, Weidong Chen, ... (+4 more)
|
|
cs.SD
|
0 |
2 months ago |
| 22 |
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
Sameek Bhattacharya, Bharath Krishnamurthy, Ajita Rattani
|
|
cs.SD
|
0 |
2 months ago |
| 23 |
Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Rostislav Makarov, Timo Gerkmann
|
|
eess.AS
|
0 |
2 months ago |
| 24 |
PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors
Masaya Kawamura, Yuma Shirahata, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 25 |
Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations
Masato Takagi, Masaya Kawamura, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 26 |
Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal
Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury
|
|
cs.CL
|
0 |
2 months ago |
| 27 |
Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning
Satwinder Singh, Qianli Wang, ... (+5 more)
|
|
eess.AS
|
0 |
2 months ago |
| 28 |
VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition
Piyush Arora, Navlika Singh, ... (+3 more)
|
|
eess.AS
|
0 |
2 months ago |
| 29 |
Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs
Nithin Rao Koluguri, Sasha Meister, ... (+5 more)
|
|
cs.CL
|
0 |
2 months ago |
| 30 |
Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR
Yichi Wang, Junzhe Chen, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 31 |
wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2
James Tanner, Morgan Sonderegger, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 32 |
From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection
Yasaman Haghbin, Sina Rashidi, ... (+7 more)
|
|
cs.CL
|
0 |
2 months ago |
| 33 |
LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
Jonghyeon Park, Olivier Jiyoun Jung, Myungwoo Oh
|
|
cs.SD
|
0 |
2 months ago |
| 34 |
Do Speech Emphasis Models Generalize across Languages and Emotions?
Megan Wei, Deepali Aneja, ... (+4 more)
|
|
cs.CL
|
0 |
2 months ago |
| 35 |
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Artem Ploujnikov, Francesco Verdini, ... (+2 more)
|
|
cs.LG
|
0 |
2 months ago |
| 36 |
Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition
Zahra Omidi, John H. L. Hansen
|
|
cs.SD
|
0 |
2 months ago |
| 37 |
Elastic Time: Dynamic Frame Rate Bottlenecks for Neural Audio Coding
Dimitrios Bralios, Paris Smaragdis, Minje Kim
|
|
cs.SD
|
0 |
2 months ago |
| 38 |
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation
Mingda Lin, Lei Ding, ... (+7 more)
|
|
cs.SD
|
0 |
2 months ago |
| 39 |
VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation
Tianxin Xie, Chenxing Li, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 40 |
When Does Quality-Aware Multimodal Fusion Matter? A Leakage-Safe Diagnostic for Decision-Level Dependence
Jaden Moon, Arvind Pillai, Andrew Campbell
|
|
cs.LG
|
0 |
2 months ago |
| 41 |
AnySimLite: A Lightweight Few-Shot Similarity Encoder for On-Device Speech-Adjacent Classification
Sourav Ghosh, Yash Bhatia, ... (+4 more)
|
|
cs.CL
|
0 |
2 months ago |
| 42 |
Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
Neelam Saini, Sourav Ghosh
|
|
cs.SD
|
0 |
2 months ago |
| 43 |
SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
Jinming Zhang, Wei Rao, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 44 |
SFL-MTSC: Leveraging Semantic Frame-Level Multi-Task Self-Consistency for Robust Multi-Intent Spoken Language Understanding
Po-Yen Chen, Berlin Chen
|
|
cs.CL
|
0 |
2 months ago |
| 45 |
Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?
Tomoya Mizumoto, Yusuke Fujita
|
|
eess.AS
|
0 |
2 months ago |
| 46 |
Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS
Sandipan Dhar, Nirmesh J. Shah, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 47 |
CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations
Ram Annamdevula, Ankit Tatawat, ... (+3 more)
|
|
eess.AS
|
0 |
2 months ago |
| 48 |
Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
Milosz Dudek, Daria Hemmerling, ... (+10 more)
|
|
eess.AS
|
0 |
2 months ago |
| 49 |
ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
Jisu Jeon, Seungyeon Jwa, ... (+7 more)
|
|
cs.SD
|
0 |
2 months ago |
| 50 |
What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients
Tuan Nguyen, Corinne Fredouille, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |