| 651 |
Human Transcription Quality Improvement
Jian Gao, Hanbo Sun, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
2 years ago |
| 652 |
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
Jiamin Xie, John H. L. Hansen
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 653 |
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
Shivam Mehta, Harm Lameris, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 654 |
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Satyam Kumar, Sai Srujana Buddi, ... (+7 more)
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 655 |
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
Qifei Li, Yingming Gao, ... (+3 more)
|
👻
Ghosted
|
cs.MM
|
5 |
2 years ago |
| 656 |
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
Haoyang Li, Yuchen Hu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
5 |
1 year ago |
| 657 |
Efficient Segmental Cascades for Speech Recognition
Hao Tang, Weiran Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
10 years ago |
| 658 |
Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks
Di He, Boon Pang Lim, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
4 |
8 years ago |
| 659 |
Semi-tied Units for Efficient Gating in LSTM and Highway Networks
Chao Zhang, Philip Woodland
|
👻
Ghosted
|
cs.CL
|
4 |
8 years ago |
| 660 |
Combining Natural Gradient with Hessian Free Methods for Sequence Training
Adnan Haider, P. C. Woodland
|
👻
Ghosted
|
cs.LG
|
4 |
7 years ago |
| 661 |
Low-Dimensional Bottleneck Features for On-Device Continuous Speech Recognition
David B. Ramsay, Kevin Kilgour, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
7 years ago |
| 662 |
End-to-end Adaptation with Backpropagation through WFST for On-device Speech Recognition System
Emiru Tsunoo, Yosuke Kashiwagi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
7 years ago |
| 663 |
Deep Neural Baselines for Computational Paralinguistics
Daniel Elsner, Stefan Langer, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
7 years ago |
| 664 |
Acoustic Model Optimization Based On Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech Recognition
Xiaodong Cui, Michael Picheny
|
👻
Ghosted
|
cs.CL
|
4 |
7 years ago |
| 665 |
Self-Teaching Networks
Liang Lu, Eric Sun, Yifan Gong
|
👻
Ghosted
|
eess.AS
|
4 |
6 years ago |
| 666 |
Adapting a FrameNet Semantic Parser for Spoken Language Understanding Using Adversarial Learning
Gabriel Marzinotto, Geraldine Damnati, Frédéric Béchet
|
👻
Ghosted
|
cs.CL
|
4 |
6 years ago |
| 667 |
Statistical Testing on ASR Performance via Blockwise Bootstrap
Zhe Liu, Fuchun Peng
|
👻
Ghosted
|
stat.ML
|
4 |
6 years ago |
| 668 |
Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search
Mittul Singh, Sami Virpioja, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
6 years ago |
| 669 |
Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware Modeling
Siyuan Feng, Odette Scharenborg
|
👻
Ghosted
|
eess.AS
|
4 |
6 years ago |
| 670 |
An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances
Hu Hu, Sabato Marco Siniscalchi, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
4 |
6 years ago |
| 671 |
ICE-Talk: an Interface for a Controllable Expressive Talking Machine
Noé Tits, Kevin El Haddad, Thierry Dutoit
|
👻
Ghosted
|
eess.AS
|
4 |
6 years ago |
| 672 |
Stochastic Talking Face Generation Using Latent Distribution Matching
Ravindra Yadav, Ashish Sardana, ... (+2 more)
|
👻
Ghosted
|
cs.CV
|
4 |
5 years ago |
| 673 |
Investigating the Impact of Cross-lingual Acoustic-Phonetic Similarities on Multilingual Speech Recognition
Muhammad Umar Farooq, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
4 |
4 years ago |
| 674 |
Benchmarking Transformers-based models on French Spoken Language Understanding tasks
Oralie Cattan, Sahar Ghannay, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
4 years ago |
| 675 |
Towards Cross-speaker Reading Style Transfer on Audiobook Dataset
Xiang Li, Changhe Song, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
4 |
4 years ago |
| 676 |
Comparison and Analysis of New Curriculum Criteria for End-to-End ASR
Georgios Karakasidis, Tamás Grósz, Mikko Kurimo
|
👻
Ghosted
|
eess.AS
|
4 |
4 years ago |
| 677 |
Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition
Ye Bai, Jie Li, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
4 |
3 years ago |
| 678 |
Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks
Cassia Valentini-Botinhao, Manuel Sam Ribeiro, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 679 |
Deep LSTM Spoken Term Detection using Wav2Vec 2.0 Recognizer
Jan Švec, Jan Lehečka, Luboš Šmídl
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |
| 680 |
A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean Speech
Li-Wei Chen, Yao-Fei Cheng, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 681 |
Biased Self-supervised learning for ASR
Florian L. Kreyssig, Yangyang Shi, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |
| 682 |
Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations
Yoori Oh, Juheon Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
3 years ago |
| 683 |
Towards continually learning new languages
Ngoc-Quan Pham, Jan Niehues, Alexander Waibel
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |
| 684 |
Speaker-Aware Anti-Spoofing
Xuechen Liu, Md Sahidullah, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 685 |
HypR: A comprehensive study for ASR hypothesis revising with a reference corpus
Yi-Wei Wang, Ke-Han Lu, Kuan-Yu Chen
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 686 |
Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
Asad Ullah, Alessandro Ragano, Andrew Hines
|
👻
Ghosted
|
eess.AS
|
4 |
2 years ago |
| 687 |
Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification
Spandan Dey, Premjeet Singh, Goutam Saha
|
👻
Ghosted
|
eess.AS
|
4 |
2 years ago |
| 688 |
How Much Context Does My Attention-Based ASR System Need?
Robert Flynn, Anton Ragni
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 689 |
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li, ... (+2 more)
|
👻
Ghosted
|
cs.MM
|
4 |
2 years ago |
| 690 |
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
Lun Wang, Om Thakkar, ... (+4 more)
|
👻
Ghosted
|
cs.CR
|
4 |
2 years ago |
| 691 |
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Oggi Rudovic, ... (+8 more)
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 692 |
Rapport-Driven Virtual Agent: Rapport Building Dialogue Strategy for Improving User Experience at First Meeting
Muhammad Yeza Baihaqi, Angel García Contreras, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 693 |
Prosody-Driven Privacy-Preserving Dementia Detection
Dominika Woszczyk, Ranya Aloufi, Soteris Demetriou
|
👻
Ghosted
|
cs.SD
|
4 |
2 years ago |
| 694 |
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
Margaret Kroll, Kelsey Kraus
|
👻
Ghosted
|
cs.AI
|
4 |
1 year ago |
| 695 |
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
Anton Firc, Manasi Chhibber, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
4 |
1 year ago |
| 696 |
Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi
|
👻
Ghosted
|
cs.SD
|
4 |
1 year ago |
| 697 |
Plagiarism Detection in Polyphonic Music using Monaural Signal Separation
Soham De, Indradyumna Roy, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
3 |
11 years ago |
| 698 |
Joint Sound Source Separation and Speaker Recognition
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.SD
|
3 |
10 years ago |
| 699 |
A real-time framework for visual feedback of articulatory data using statistical shape models
Kristy James, Alexander Hewer, ... (+2 more)
|
👻
Ghosted
|
cs.HC
|
3 |
9 years ago |
| 700 |
Learning Similarity Functions for Pronunciation Variations
Einat Naaman, Yossi Adi, Joseph Keshet
|
👻
Ghosted
|
cs.CL
|
3 |
9 years ago |