| 601 |
FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics
Katri Leino, Juho Leinonen, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
6 |
6 years ago |
| 602 |
Reduce and Reconstruct: ASR for Low-Resource Phonetic Languages
Anuj Diwan, Preethi Jyothi
|
👻
Ghosted
|
eess.AS
|
6 |
5 years ago |
| 603 |
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
Wubo Li, Dongwei Jiang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
5 years ago |
| 604 |
Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models
Vineet Garg, Ognjen Rudovic, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 605 |
The Emotion is Not One-hot Encoding: Learning with Grayscale Label for Emotion Recognition in Conversation
Joosung Lee
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 606 |
Accelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit Quantization
Andrea Fasoli, Chia-Yu Chen, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 607 |
Nonwords Pronunciation Classification in Language Development Tests for Preschool Children
Ilja Baumann, Dominik Wagner, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 608 |
Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS
Kenta Udagawa, Yuki Saito, Hiroshi Saruwatari
|
👻
Ghosted
|
cs.SD
|
6 |
4 years ago |
| 609 |
Toward Low-Cost End-to-End Spoken Language Understanding
Marco Dinarelli, Marco Naguib, François Portet
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 610 |
GlowVC: Mel-spectrogram space disentangling model for language-independent text-free voice conversion
Magdalena Proszewska, Grzegorz Beringer, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 611 |
End-to-end speech recognition modeling from de-identified data
Martin Flechl, Shou-Chun Yin, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 612 |
PoeticTTS -- Controllable Poetry Reading for Literary Studies
Julia Koch, Florian Lux, ... (+7 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 613 |
Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free
M. Iftekhar Tanveer, Diego Casabuena, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 614 |
A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis
Qibing Bai, Tom Ko, Yu Zhang
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 615 |
Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition
Kartik Audhkhasi, Yinghui Huang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
3 years ago |
| 616 |
CCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment
Yuchen Liu, Li-Chia Yang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
3 years ago |
| 617 |
Relationship between auditory and semantic entrainment using Deep Neural Networks (DNN)
Jay Kejriwal, Štefan Beňuš
|
👻
Ghosted
|
cs.CL
|
6 |
2 years ago |
| 618 |
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
Xinlei Niu, Jing Zhang, Charles Patrick Martin
|
👻
Ghosted
|
cs.SD
|
6 |
2 years ago |
| 619 |
Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
6 |
2 years ago |
| 620 |
Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
Shuai Wang, Dehao Zhang, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
6 |
2 years ago |
| 621 |
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
6 |
2 years ago |
| 622 |
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Heng-Jui Chang, Hongyu Gong, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
6 |
1 year ago |
| 623 |
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Umberto Cappellazzo, Minsu Kim, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
6 |
1 year ago |
| 624 |
Recognize Foreign Low-Frequency Words with Similar Pairs
Xi Ma, Xiaoxi Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
11 years ago |
| 625 |
Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
Wenpeng Li, BinBin Zhang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
9 years ago |
| 626 |
Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal Classification
Bo-Ru Lu, Frank Shyu, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
5 |
8 years ago |
| 627 |
Empirical Evaluation of Speaker Adaptation on DNN based Acoustic Model
Ke Wang, Junbo Zhang, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
8 years ago |
| 628 |
Unsupervised and Efficient Vocabulary Expansion for Recurrent Neural Network Language Models in ASR
Yerbolat Khassanov, Eng Siong Chng
|
👻
Ghosted
|
cs.CL
|
5 |
8 years ago |
| 629 |
Memory Time Span in LSTMs for Multi-Speaker Source Separation
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
5 |
8 years ago |
| 630 |
Synchronising audio and ultrasound by learning cross-modal embeddings
Aciel Eshky, Manuel Sam Ribeiro, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
7 years ago |
| 631 |
Ultrasound tongue imaging for diarization and alignment of child speech therapy sessions
Manuel Sam Ribeiro, Aciel Eshky, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
7 years ago |
| 632 |
Latent Dirichlet Allocation Based Acoustic Data Selection for Automatic Speech Recognition
Mortaza, Doulaty, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
5 |
7 years ago |
| 633 |
Bandwidth Embeddings for Mixed-bandwidth Speech Recognition
Gautam Mantena, Ozlem Kalinli, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 634 |
Speaker Re-identification with Speaker Dependent Speech Enhancement
Yanpei Shi, Qiang Huang, Thomas Hain
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 635 |
Data balancing for boosting performance of low-frequency classes in Spoken Language Understanding
Judith Gaspers, Quynh Do, Fabian Triefenbach
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 636 |
Do face masks introduce bias in speech technologies? The case of automated scoring of speaking proficiency
Anastassia Loukina, Keelan Evanini, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 637 |
RECOApy: Data recording, pre-processing and phonetic transcription for end-to-end speech-based applications
Adriana Stan
|
👻
Ghosted
|
eess.AS
|
5 |
5 years ago |
| 638 |
Perceptimatic: A human speech perception benchmark for unsupervised subword modelling
Juliette Millet, Ewan Dunbar
|
👻
Ghosted
|
cs.CL
|
5 |
5 years ago |
| 639 |
Knowledge Distillation for Singing Voice Detection
Soumava Paul, Gurunath Reddy M, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
5 years ago |
| 640 |
Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE
Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber
|
👻
Ghosted
|
cs.CL
|
5 |
4 years ago |
| 641 |
Towards Green ASR: Lossless 4-bit Quantization of a Hybrid TDNN System on the 300-hr Switchboard Corpus
Junhao Xu, Shoukang Hu, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
4 years ago |
| 642 |
Low-resource Accent Classification in Geographically-proximate Settings: A Forensic and Sociophonetics Perspective
Qingcheng Zeng, Dading Chong, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
4 years ago |
| 643 |
Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)
Ziyao Zhang, Alessio Falai, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
5 |
4 years ago |
| 644 |
A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese
Song Zhang, Ken Zheng, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
4 years ago |
| 645 |
Streaming Intended Query Detection using E2E Modeling for Continued Conversation
Shuo-yiin Chang, Guru Prakash, ... (+8 more)
|
👻
Ghosted
|
cs.CL
|
5 |
4 years ago |
| 646 |
AdaMS: Deep Metric Learning with Adaptive Margin and Adaptive Scale for Acoustic Word Discrimination
Myunghun Jung, Hoirin Kim
|
👻
Ghosted
|
eess.AS
|
5 |
3 years ago |
| 647 |
Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages
Anusha Prakash, Arun Kumar, ... (+25 more)
|
👻
Ghosted
|
eess.AS
|
5 |
3 years ago |
| 648 |
Multitask Learning for Low Resource Spoken Language Understanding
Quentin Meeus, Marie-Francine Moens, Hugo Van hamme
|
👻
Ghosted
|
cs.CL
|
5 |
3 years ago |
| 649 |
Efficient Multimodal Neural Networks for Trigger-less Voice Assistants
Sai Srujana Buddi, Utkarsh Oggy Sarawgi, ... (+5 more)
|
👻
Ghosted
|
cs.LG
|
5 |
3 years ago |
| 650 |
A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao, Xiuping Yang, ... (+7 more)
|
👻
Ghosted
|
cs.SD
|
5 |
3 years ago |