| 751 |
Streaming Audio-Visual Speech Recognition with Alignment Regularization
Pingchuan Ma, Niko Moritz, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
2 |
3 years ago |
| 752 |
On-Device Speaker Anonymization of Acoustic Embeddings for ASR based onFlexible Location Gradient Reversal Layer
Md Asif Jalal, Pablo Peso Parada, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
2 |
3 years ago |
| 753 |
PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network
Qinghua Liu, Meng Ge, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
2 |
2 years ago |
| 754 |
Improved Factorized Neural Transducer Model For text-only Domain Adaptation
Junzhe Liu, Jianwei Yu, Xie Chen
|
👻
Ghosted
|
cs.CL
|
2 |
2 years ago |
| 755 |
RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios
Yiwen Shao, Shi-Xiong Zhang, Dong Yu
|
👻
Ghosted
|
eess.AS
|
2 |
2 years ago |
| 756 |
FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
Swarup Ranjan Behera, Abhishek Dhiman, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
2 |
2 years ago |
| 757 |
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
Adriana Fernandez-Lopez, Honglie Chen, ... (+6 more)
|
👻
Ghosted
|
cs.CV
|
2 |
2 years ago |
| 758 |
SecureSpectra: Safeguarding Digital Identity from Deep Fake Threats via Intelligent Signatures
Oguzhan Baser, Kaan Kale, Sandeep P. Chinchali
|
👻
Ghosted
|
cs.CR
|
2 |
2 years ago |
| 759 |
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
Jizhen Li, Xinmeng Xu, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
2 |
2 years ago |
| 760 |
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
Haoxiang Shi, Ziqi Liang, Jun Yu
|
👻
Ghosted
|
cs.MM
|
2 |
2 years ago |
| 761 |
Turbo your multi-modal classification with contrastive learning
Zhiyu Zhang, Da Liu, ... (+4 more)
|
👻
Ghosted
|
cs.LG
|
2 |
1 year ago |
| 762 |
Song Form-aware Full-Song Text-to-Lyrics Generation with Multi-Level Granularity Syllable Count Control
Yunkee Chae, Eunsik Shin, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
2 |
1 year ago |
| 763 |
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
Youjun Chen, Xurong Xie, ... (+7 more)
|
👻
Ghosted
|
cs.SD
|
2 |
1 year ago |
| 764 |
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Wenxuan Wu, Shuai Wang, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
2 |
1 year ago |
| 765 |
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Ruofan Hu, Yan Xia, ... (+6 more)
|
👻
Ghosted
|
cs.IR
|
2 |
1 year ago |
| 766 |
Automatic Measurement of Pre-aspiration
Yaniv Sheena, Míša Hejná, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
1 |
9 years ago |
| 767 |
Joint Learning of Interactive Spoken Content Retrieval and Trainable User Simulator
Pei-Hung Chung, Kuan Tung, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
1 |
8 years ago |
| 768 |
Play Duration based User-Entity Affinity Modeling in Spoken Dialog System
Bo Xiao, Nicholas Monath, ... (+2 more)
|
👻
Ghosted
|
cs.IR
|
1 |
8 years ago |
| 769 |
Multiple topic identification in telephone conversations
Xavier Bost, Marc El Bèze, Renato De Mori
|
👻
Ghosted
|
cs.CL
|
1 |
7 years ago |
| 770 |
Real to H-space Encoder for Speech Recognition
Titouan Parcollet, Mohamed Morchid, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
1 |
7 years ago |
| 771 |
Exploration of Audio Quality Assessment and Anomaly Localisation Using Attention Models
Qiang Huang, Thomas Hain
|
👻
Ghosted
|
eess.AS
|
1 |
6 years ago |
| 772 |
Improving Unsupervised Sparsespeech Acoustic Models with Categorical Reparameterization
Benjamin Milde, Chris Biemann
|
👻
Ghosted
|
eess.AS
|
1 |
6 years ago |
| 773 |
Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions
Santi Prieto, Alfonso Ortega, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
1 |
6 years ago |
| 774 |
Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
Masood S. Mortazavi
|
👻
Ghosted
|
cs.LG
|
1 |
5 years ago |
| 775 |
Incremental Machine Speech Chain Towards Enabling Listening while Speaking in Real-time
Sashi Novitasari, Andros Tjandra, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
1 |
5 years ago |
| 776 |
The THUEE System Description for the IARPA OpenASR21 Challenge
Jing Zhao, Haoyu Wang, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
1 |
4 years ago |
| 777 |
Graph-based Multi-View Fusion and Local Adaptation: Mitigating Within-Household Confusability for Speaker Identification
Long Chen, Yixiong Meng, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
1 |
4 years ago |
| 778 |
Attention Enhanced Citrinet for Speech Recognition
Xianchao Wu
|
👻
Ghosted
|
cs.CL
|
1 |
4 years ago |
| 779 |
A Universal Identity Backdoor Attack against Speaker Verification based on Siamese Network
Haodong Zhao, Wei Du, ... (+2 more)
|
👻
Ghosted
|
cs.CR
|
1 |
3 years ago |
| 780 |
PoCaPNet: A Novel Approach for Surgical Phase Recognition Using Speech and X-Ray Images
Kubilay Can Demir, Tobias Weise, ... (+4 more)
|
👻
Ghosted
|
cs.HC
|
1 |
3 years ago |
| 781 |
CASEIN: Cascading Explicit and Implicit Control for Fine-grained Emotion Intensity Regulation
Yuhao Cui, Xiongwei Wang, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
1 |
3 years ago |
| 782 |
Getting More for Less: Using Weak Labels and AV-Mixup for Robust Audio-Visual Speaker Verification
Anith Selvakumar, Homa Fashandi
|
👻
Ghosted
|
cs.SD
|
1 |
2 years ago |
| 783 |
Unsupervised Auditory and Semantic Entrainment Models with Deep Neural Networks
Jay Kejriwal, Stefan Benus, Lina M. Rojas-Barahona
|
👻
Ghosted
|
cs.CL
|
1 |
2 years ago |
| 784 |
Compositional Generalization in Spoken Language Understanding
Avik Ray, Yilin Shen, Hongxia Jin
|
👻
Ghosted
|
cs.CL
|
1 |
2 years ago |
| 785 |
Sequence-to-Sequence Multi-Modal Speech In-Painting
Mahsa Kadkhodaei Elyaderani, Shahram Shirani
|
👻
Ghosted
|
cs.SD
|
1 |
2 years ago |
| 786 |
MMSD-Net: Towards Multi-modal Stuttering Detection
Liangyu Nie, Sudarsana Reddy Kadiri, Ruchit Agrawal
|
👻
Ghosted
|
cs.SD
|
1 |
2 years ago |
| 787 |
Neuromorphic Keyword Spotting with Pulse Density Modulation MEMS Microphones
Sidi Yaya Arnaud Yarga, Sean U. N. Wood
|
👻
Ghosted
|
cs.NE
|
1 |
2 years ago |
| 788 |
Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation
Myeonghoon Ryu, Hongseok Oh, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
1 |
1 year ago |
| 789 |
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren, Chenxing Li, ... (+8 more)
|
👻
Ghosted
|
cs.MM
|
1 |
1 year ago |
| 790 |
Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations
Seongsil Heo, Calvin Murdock, ... (+2 more)
|
👻
Ghosted
|
cs.HC
|
1 |
1 year ago |
| 791 |
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
Elvir Karimov, Alexander Varlamov, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
1 |
1 year ago |
| 792 |
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Le Xu, Chenxing Li, ... (+6 more)
|
👻
Ghosted
|
cs.MM
|
1 |
1 year ago |
| 793 |
Children's Voice Privacy: First Steps And Emerging Challenges
Ajinkya Kulkarni, Francisco Teixeira, ... (+4 more)
|
👻
Ghosted
|
cs.CY
|
1 |
1 year ago |
| 794 |
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Orchid Chetia Phukan, Girish, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
1 |
1 year ago |
| 795 |
SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
Orchid Chetia Phukan, Mohd Mujtaba Akhtar, ... (+9 more)
|
👻
Ghosted
|
eess.AS
|
1 |
1 year ago |
| 796 |
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
Fang Kang, Yin Cao, Haoyu Chen
|
👻
Ghosted
|
cs.SD
|
1 |
1 year ago |
| 797 |
Blind score normalization method for PLDA based speaker recognition
Danila Doroshin, Nikolay Lubimov, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
0 |
10 years ago |
| 798 |
TheanoLM - An Extensible Toolkit for Neural Network Language Modeling
Seppo Enarvi, Mikko Kurimo
|
👻
Ghosted
|
cs.CL
|
0 |
10 years ago |
| 799 |
Generation and Pruning of Pronunciation Variants to Improve ASR Accuracy
Zhenhao Ge, Aravind Ganapathiraju, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
0 |
10 years ago |
| 800 |
Entity-Aware Language Model as an Unsupervised Reranker
Mohammad Sadegh Rasooli, Sarangarajan Parthasarathy
|
👻
Ghosted
|
cs.CL
|
0 |
8 years ago |