| 551 |
Self-supervised pre-training with acoustic configurations for replay spoofing detection
Hye-jin Shim, Hee-Soo Heo, ... (+2 more)
|
👻
Ghosted
|
cs.LG
|
8 |
6 years ago |
| 552 |
CNN-LSTM models for Multi-Speaker Source Separation using Bayesian Hyper Parameter Optimization
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
8 |
6 years ago |
| 553 |
Exploring TTS without T Using Biologically/Psychologically Motivated Neural Network Modules (ZeroSpeech 2020)
Takashi Morita, Hiroki Koda
|
👻
Ghosted
|
cs.CL
|
8 |
6 years ago |
| 554 |
Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition
Kohei Matsuura, Masato Mimura, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
8 |
6 years ago |
| 555 |
Analysis of Disfluency in Children's Speech
Trang Tran, Morgan Tinkler, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
8 |
5 years ago |
| 556 |
Auxiliary Sequence Labeling Tasks for Disfluency Detection
Dongyub Lee, Byeongil Ko, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
8 |
5 years ago |
| 557 |
A low latency ASR-free end to end spoken language understanding system
Mohamed Mhiri, Samuel Myer, Vikrant Singh Tomar
|
👻
Ghosted
|
cs.CV
|
8 |
5 years ago |
| 558 |
Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
Yuto Nishimura, Yuki Saito, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
8 |
4 years ago |
| 559 |
Bottleneck Low-rank Transformers for Low-resource Spoken Language Understanding
Pu Wang, Hugo Van hamme
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 560 |
A Multi-Task BERT Model for Schema-Guided Dialogue State Tracking
Eleftherios Kapelonis, Efthymios Georgiou, Alexandros Potamianos
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 561 |
Compute Cost Amortized Transformer for Streaming ASR
Yi Xie, Jonathan Macoskey, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 562 |
Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation
Vikramjit Mitra, Hsiang-Yun Sherry Chien, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
8 |
4 years ago |
| 563 |
Explicit Intensity Control for Accented Text-to-speech
Rui Liu, Haolin Zuo, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
8 |
3 years ago |
| 564 |
SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
Mirazul Haque, Rutvij Shah, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
8 |
3 years ago |
| 565 |
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
Ming-Hao Hsu, Kai-Wei Chang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
8 |
2 years ago |
| 566 |
Prompt Tuning for Audio Deepfake Detection: Computationally Efficient Test-time Domain Adaptation with Limited Target Dataset
Hideyuki Oiso, Yuto Matsunaga, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
8 |
1 year ago |
| 567 |
Contrastive Entropy: A new evaluation metric for unnormalized language models
Kushal Arora, Anand Rangarajan
|
👻
Ghosted
|
cs.CL
|
7 |
10 years ago |
| 568 |
A Generative Model for Score Normalization in Speaker Recognition
Albert Swart, Niko Brummer
|
👻
Ghosted
|
stat.ML
|
7 |
8 years ago |
| 569 |
Modeling Interpersonal Influence of Verbal Behavior in Couples Therapy Dyadic Interactions
Sandeep Nallan Chakravarthula, Brian Baucom, Panayiotis Georgiou
|
👻
Ghosted
|
cs.CL
|
7 |
8 years ago |
| 570 |
Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSR
Wilfried Michel, Ralf Schlüter, Hermann Ney
|
👻
Ghosted
|
eess.AS
|
7 |
7 years ago |
| 571 |
SANTLR: Speech Annotation Toolkit for Low Resource Languages
Xinjian Li, Zhong Zhou, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
7 |
7 years ago |
| 572 |
Practical applicability of deep neural networks for overlapping speaker separation
Pieter Appeltans, Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
7 |
6 years ago |
| 573 |
CTC-synchronous Training for Monotonic Attention Model
Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara
|
👻
Ghosted
|
cs.CL
|
7 |
6 years ago |
| 574 |
Utterance-Wise Meeting Transcription System Using Asynchronous Distributed Microphones
Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu
|
👻
Ghosted
|
eess.AS
|
7 |
6 years ago |
| 575 |
ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
Valentin Pelloin, Franck Dary, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
7 |
4 years ago |
| 576 |
Improving Data Driven Inverse Text Normalization using Data Augmentation
Laxmi Pandey, Debjyoti Paul, ... (+7 more)
|
👻
Ghosted
|
cs.CL
|
7 |
4 years ago |
| 577 |
Thutmose Tagger: Single-pass neural model for Inverse Text Normalization
Alexandra Antonova, Evelina Bakhturina, Boris Ginsburg
|
👻
Ghosted
|
cs.CL
|
7 |
4 years ago |
| 578 |
ESSumm: Extractive Speech Summarization from Untranscribed Meeting
Jun Wang
|
👻
Ghosted
|
eess.AS
|
7 |
3 years ago |
| 579 |
Spoken Term Detection and Relevance Score Estimation using Dot-Product of Pronunciation Embeddings
Jan Švec, Luboš Šmídl, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
7 |
3 years ago |
| 580 |
A Compact End-to-End Model with Local and Global Context for Spoken Language Identification
Fei Jia, Nithin Rao Koluguri, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
7 |
3 years ago |
| 581 |
Single-channel speech enhancement using learnable loss mixup
Oscar Chang, Dung N. Tran, Kazuhito Koishida
|
👻
Ghosted
|
eess.AS
|
7 |
2 years ago |
| 582 |
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
Yun Chen, Lingxiao Yang, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
7 |
2 years ago |
| 583 |
Missingness-resilient Video-enhanced Multimodal Disfluency Detection
Payal Mohapatra, Shamika Likhite, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
7 |
2 years ago |
| 584 |
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
Han EunGi, Oh Hyun-Bin, ... (+5 more)
|
👻
Ghosted
|
cs.CV
|
7 |
2 years ago |
| 585 |
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
Honglie Chen, Rodrigo Mira, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
7 |
2 years ago |
| 586 |
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
Yifei Xin, Xuxin Cheng, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
7 |
1 year ago |
| 587 |
Audio Content based Geotagging in Multimedia
Anurag Kumar, Benjamin Elizalde, Bhiksha Raj
|
👻
Ghosted
|
cs.SD
|
6 |
10 years ago |
| 588 |
Sequential Recurrent Neural Networks for Language Modeling
Youssef Oualil, Clayton Greenberg, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
9 years ago |
| 589 |
A Batch Noise Contrastive Estimation Approach for Training Large Vocabulary Language Models
Youssef Oualil, Dietrich Klakow
|
👻
Ghosted
|
cs.CL
|
6 |
9 years ago |
| 590 |
State Gradients for RNN Memory Analysis
Lyan Verwimp, Hugo Van hamme, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
8 years ago |
| 591 |
Neural Named Entity Recognition from Subword Units
Abdalghani Abujabal, Judith Gaspers
|
👻
Ghosted
|
cs.CL
|
6 |
8 years ago |
| 592 |
Large-scale Speaker Retrieval on Random Speaker Variability Subspace
Suwon Shon, Younggun Lee, Taesu Kim
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 593 |
Multi-lingual Dialogue Act Recognition with Deep Learning Methods
Jiří Martínek, Pavel Král, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 594 |
Code-Switching Detection Using ASR-Generated Language Posteriors
Qinyi Wang, Emre Yılmaz, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 595 |
Attention model for articulatory features detection
Ievgen Karaulov, Dmytro Tkanov
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 596 |
NIESR: Nuisance Invariant End-to-end Speech Recognition
I-Hung Hsu, Ayush Jaiswal, Premkumar Natarajan
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 597 |
Large-Scale Mixed-Bandwidth Deep Neural Network Acoustic Modeling for Automatic Speech Recognition
Khoi-Nguyen C. Mac, Xiaodong Cui, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 598 |
The phonetic bases of vocal expressed emotion: natural versus acted
Hira Dhamyal, Shahan Ali Memon, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
6 years ago |
| 599 |
Chirp Complex Cepstrum-based Decomposition for Asynchronous Glottal Analysis
Thomas Drugman, Thierry Dutoit
|
👻
Ghosted
|
cs.SD
|
6 |
6 years ago |
| 600 |
Applying GPGPU to Recurrent Neural Network Language Model based Fast Network Search in the Real-Time LVCSR
Kyungmin Lee, Chiyoun Park, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
6 |
6 years ago |