DEPA: Self-Supervised Audio Embedding for Depression Detection

October 29, 2019 · Declared Dead · 🏛 ACM Multimedia

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Pingyue Zhang, Mengyue Wu, Heinrich Dinkel, Kai Yu arXiv ID 1910.13028 Category cs.HC: Human-Computer Interaction Cross-listed cs.SD, eess.AS Citations 74 Venue ACM Multimedia Last Checked 3 months ago

Abstract

Depression detection research has increased over the last few decades, one major bottleneck of which is the limited data availability and representation learning. Recently, self-supervised learning has seen success in pretraining text embeddings and has been applied broadly on related tasks with sparse data, while pretrained audio embeddings based on self-supervised learning are rarely investigated. This paper proposes DEPA, a self-supervised, pretrained depression audio embedding method for depression detection. An encoder-decoder network is used to extract DEPA on in-domain depressed datasets (DAIC and MDD) and out-domain (Switchboard, Alzheimer's) datasets. With DEPA as the audio embedding extracted at response-level, a significant performance gain is achieved on downstream tasks, evaluated on both sparse datasets like DAIC and large major depression disorder dataset (MDD). This paper not only exhibits itself as a novel embedding extracting method capturing response-level representation for depression detection but more significantly, is an exploration of self-supervised learning in a specific task within audio processing.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Human-Computer Interaction

R.I.P. 👻 Ghosted

Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy

Ben Shneiderman

cs.HC 🏛 International journal of human computer interactions 📚 974 cites 6 years ago

R.I.P. 👻 Ghosted

Improving fairness in machine learning systems: What do industry practitioners need?

Kenneth Holstein, Jennifer Wortman Vaughan, ... (+3 more)

cs.HC 🏛 CHI 📚 919 cites 7 years ago

R.I.P. 👻 Ghosted

Identifying Stable Patterns over Time for Emotion Recognition from EEG

Wei-Long Zheng, Jia-Yi Zhu, Bao-Liang Lu

cs.HC 🏛 IEEE TAC 📚 837 cites 10 years ago

R.I.P. 👻 Ghosted

Questioning the AI: Informing Design Practices for Explainable AI User Experiences

Q. Vera Liao, Daniel Gruen, Sarah Miller

cs.HC 🏛 CHI 📚 835 cites 6 years ago

R.I.P. 👻 Ghosted

Deep Learning for Sensor-based Human Activity Recognition: Overview, Challenges and Opportunities

Kaixuan Chen, Dalin Zhang, ... (+4 more)

cs.HC 🏛 ACM CSUR 📚 788 cites 6 years ago

R.I.P. 👻 Ghosted

Educational data mining and learning analytics: An updated survey

C. Romero, S. Ventura

cs.HC 🏛 WIREs Data Mining Knowl. Discov. 📚 787 cites 2 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, ... (+29 more)

cs.CL 🏛 NeurIPS 📚 54.2K cites 6 years ago

R.I.P. 👻 Ghosted

PyTorch: An Imperative Style, High-Performance Deep Learning Library

Adam Paszke, Sam Gross, ... (+19 more)

cs.LG 🏛 NeurIPS 📚 49.7K cites 6 years ago

R.I.P. 👻 Ghosted

XGBoost: A Scalable Tree Boosting System

Tianqi Chen, Carlos Guestrin

cs.LG 🏛 KDD 📚 49.2K cites 10 years ago

R.I.P. 👻 Ghosted

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, Christian Szegedy

cs.LG 🏛 ICML 📚 46.0K cites 11 years ago