Speaker Diarization using Deep Recurrent Convolutional Neural Networks for Speaker Embeddings

August 09, 2017 ยท Declared Dead ยท ๐Ÿ› International Conference on Information Systems Architecture and Technology

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Pawel Cyrta, Tomasz Trzciล„ski, Wojciech Stokowiec arXiv ID 1708.02840 Category cs.SD: Sound Cross-listed cs.MM, cs.NE Citations 37 Venue International Conference on Information Systems Architecture and Technology Last Checked 6 months ago
Abstract
In this paper we propose a new method of speaker diarization that employs a deep learning architecture to learn speaker embeddings. In contrast to the traditional approaches that build their speaker embeddings using manually hand-crafted spectral features, we propose to train for this purpose a recurrent convolutional neural network applied directly on magnitude spectrograms. To compare our approach with the state of the art, we collect and release for the public an additional dataset of over 6 hours of fully annotated broadcast material. The results of our evaluation on the new dataset and three other benchmark datasets show that our proposed method significantly outperforms the competitors and reduces diarization error rate by a large margin of over 30% with respect to the baseline.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Sound

Died the same way โ€” ๐Ÿ‘ป Ghosted