| 451 |
Towards Universal Dialogue Act Tagging for Task-Oriented Dialogues
Shachi Paul, Rahul Goel, Dilek Hakkani-Tür
|
👻
Ghosted
|
cs.CL
|
13 |
7 years ago |
| 452 |
Analyzing the Quality and Stability of a Streaming End-to-End On-Device Speech Recognizer
Yuan Shangguan, Kate Knister, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
13 |
6 years ago |
| 453 |
Transfer Learning for Robust Low-Resource Children's Speech ASR with Transformers and Source-Filter Warping
Jenthe Thienpondt, Kris Demuynck
|
👻
Ghosted
|
eess.AS
|
13 |
4 years ago |
| 454 |
Improving Transformer-based Conversational ASR by Inter-Sentential Attention Mechanism
Kun Wei, Pengcheng Guo, Ning Jiang
|
👻
Ghosted
|
cs.SD
|
13 |
4 years ago |
| 455 |
ASR Error Correction with Constrained Decoding on Operation Prediction
Jingyuan Yang, Rongjun Li, Wei Peng
|
👻
Ghosted
|
cs.CL
|
13 |
4 years ago |
| 456 |
A Language Agnostic Multilingual Streaming On-Device ASR System
Bo Li, Tara N. Sainath, ... (+10 more)
|
👻
Ghosted
|
eess.AS
|
13 |
4 years ago |
| 457 |
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition
Xiaohuan Zhou, Jiaming Wang, ... (+5 more)
|
👻
Ghosted
|
cs.MM
|
13 |
3 years ago |
| 458 |
4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders
Yui Sudo, Muhammad Shakeel, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
13 |
3 years ago |
| 459 |
Leveraging Word Embeddings for Spoken Document Summarization
Kuan-Yu Chen, Shih-Hung Liu, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
11 years ago |
| 460 |
The IBM Speaker Recognition System: Recent Advances and Error Analysis
Seyed Omid Sadjadi, Jason Pelecanos, Sriram Ganapathy
|
👻
Ghosted
|
cs.CL
|
12 |
10 years ago |
| 461 |
Automatic Speech Recognition and Topic Identification for Almost-Zero-Resource Languages
Matthew Wiesner, Chunxi Liu, ... (+7 more)
|
👻
Ghosted
|
cs.CL
|
12 |
8 years ago |
| 462 |
Gated Recurrent Unit Based Acoustic Modeling with Future Context
Jie Li, Xiaorui Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
8 years ago |
| 463 |
Semi-supervised and Active-learning Scenarios: Efficient Acoustic Model Refinement for a Low Resource Indian Language
Maharajan Chellapriyadharshini, Anoop Toffy, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 464 |
Enriching Rare Word Representations in Neural Language Models by Embedding Matrix Augmentation
Yerbolat Khassanov, Zhiping Zeng, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 465 |
Self-imitating Feedback Generation Using GAN for Computer-Assisted Pronunciation Training
Seung Hee Yang, Minhwa Chung
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 466 |
On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech Retrieval
Ankita Pasad, Bowen Shi, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 467 |
A computational model of early language acquisition from audiovisual experiences of young infants
Okko Räsänen, Khazar Khorrami
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 468 |
Predicting Behavior in Cancer-Afflicted Patient and Spouse Interactions using Speech and Language
Sandeep Nallan Chakravarthula, Haoqi Li, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 469 |
Early Stage LM Integration Using Local and Global Log-Linear Combination
Wilfried Michel, Ralf Schlüter, Hermann Ney
|
👻
Ghosted
|
eess.AS
|
12 |
6 years ago |
| 470 |
Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement
Jun Qi, Hu Hu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
12 |
6 years ago |
| 471 |
Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification
Hu Hu, Sabato Marco Siniscalchi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
6 years ago |
| 472 |
Audio Dequantization for High Fidelity Audio Generation in Flow-based Neural Vocoder
Hyun-Wook Yoon, Sang-Hoon Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
6 years ago |
| 473 |
On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer
Liang Lu, Zhong Meng, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
5 years ago |
| 474 |
Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion
Disong Wang, Songxiang Liu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 475 |
Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech Recognition
Sashi Novitasari, Andros Tjandra, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
5 years ago |
| 476 |
Bootstrap an end-to-end ASR system by multilingual training, transfer learning, text-to-text mapping and synthetic audio
Manuel Giollo, Deniz Gunceler, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 477 |
Detecting Unintended Memorization in Language-Model-Fused ASR
W. Ronny Huang, Steve Chien, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 478 |
Towards End-to-End Private Automatic Speaker Recognition
Francisco Teixeira, Alberto Abad, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
4 years ago |
| 479 |
Improving Deliberation by Text-Only and Semi-Supervised Training
Ke Hu, Tara N. Sainath, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 480 |
When Is TTS Augmentation Through a Pivot Language Useful?
Nathaniel Robinson, Perez Ogayo, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 481 |
Visual Transformers for Primates Classification and Covid Detection
Steffen Illium, Robert Müller, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
12 |
3 years ago |
| 482 |
Parameter-Efficient Learning for Text-to-Speech Accent Adaptation
Li-Jen Yang, Chao-Han Huck Yang, Jen-Tzung Chien
|
👻
Ghosted
|
cs.SD
|
12 |
3 years ago |
| 483 |
Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Weiran Wang, Zelin Wu, ... (+11 more)
|
👻
Ghosted
|
cs.CL
|
12 |
2 years ago |
| 484 |
Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
12 |
2 years ago |
| 485 |
Learning Speech Rate in Speech Recognition
Xiangyu Zeng, Shi Yin, Dong Wang
|
👻
Ghosted
|
cs.CL
|
11 |
11 years ago |
| 486 |
An Empirical Analysis of the Correlation of Syntax and Prosody
Arne Köhn, Timo Baumann, Oskar Dörfler
|
👻
Ghosted
|
cs.CL
|
11 |
8 years ago |
| 487 |
Keyword Spotting for Hearing Assistive Devices Robust to External Speakers
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen
|
👻
Ghosted
|
cs.SD
|
11 |
7 years ago |
| 488 |
The NTNU System at the Interspeech 2020 Non-Native Children's Speech ASR Challenge
Tien-Hong Lo, Fu-An Chao, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 489 |
A non-causal FFTNet architecture for speech enhancement
Muhammed PV Shifas, Nagaraj Adiga, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 490 |
Black-box Adaptation of ASR for Accented Speech
Kartik Khandelwal, Preethi Jyothi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 491 |
Compact Speaker Embedding: lrx-vector
Munir Georges, Jonathan Huang, Tobias Bocklet
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 492 |
Evolutionary Algorithm Enhanced Neural Architecture Search for Text-Independent Speaker Verification
Xiaoyang Qu, Jianzong Wang, Jing Xiao
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 493 |
Extracting Targeted Training Data from ASR Models, and How to Mitigate It
Ehsan Amid, Om Thakkar, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
11 |
4 years ago |
| 494 |
Automatic Prosody Annotation with Pre-Trained Text-Speech Model
Ziqian Dai, Jianwei Yu, ... (+6 more)
|
👻
Ghosted
|
cs.SD
|
11 |
4 years ago |
| 495 |
QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using MLPMixer
Jinmiao Huang, Waseem Gharbieh, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
11 |
4 years ago |
| 496 |
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model
Brooke Stephenson, Laurent Besacier, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
11 |
4 years ago |
| 497 |
Unsupervised domain adaptation for speech recognition with unsupervised error correction
Long Mai, Julie Carson-Berndsen
|
👻
Ghosted
|
cs.SD
|
11 |
3 years ago |
| 498 |
VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
Junjie Li, Meng Ge, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
11 |
3 years ago |
| 499 |
Minimum Latency Training of Sequence Transducers for Streaming End-to-End Speech Recognition
Yusuke Shinohara, Shinji Watanabe
|
👻
Ghosted
|
eess.AS
|
11 |
3 years ago |
| 500 |
Bayesian Networks for the robust and unbiased prediction of depression and its symptoms utilizing speech and multimodal data
Salvatore Fara, Orlaith Hickey, ... (+4 more)
|
👻
Ghosted
|
cs.LG
|
11 |
3 years ago |