CLAMShell: Speeding up Crowds for Low-latency Data Labeling

September 20, 2015 · Declared Dead · 🏛 Proceedings of the VLDB Endowment

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Daniel Haas, Jiannan Wang, Eugene Wu, Michael J. Franklin arXiv ID 1509.05969 Category cs.DB: Databases Citations 64 Venue Proceedings of the VLDB Endowment Last Checked 3 months ago

Abstract

Data labeling is a necessary but often slow process that impedes the development of interactive systems for modern data analysis. Despite rising demand for manual data labeling, there is a surprising lack of work addressing its high and unpredictable latency. In this paper, we introduce CLAMShell, a system that speeds up crowds in order to achieve consistently low-latency data labeling. We offer a taxonomy of the sources of labeling latency and study several large crowd-sourced labeling deployments to understand their empirical latency profiles. Driven by these insights, we comprehensively tackle each source of latency, both by developing novel techniques such as straggler mitigation and pool maintenance and by optimizing existing methods such as crowd retainer pools and active learning. We evaluate CLAMShell in simulation and on live workers on Amazon's Mechanical Turk, demonstrating that our techniques can provide an order of magnitude speedup and variance reduction over existing crowdsourced labeling strategies.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Databases

R.I.P. 👻 Ghosted

Datasheets for Datasets

Timnit Gebru, Jamie Morgenstern, ... (+5 more)

cs.DB 🏛 CACM 📚 2.6K cites 8 years ago

R.I.P. 👻 Ghosted

The Case for Learned Index Structures

Tim Kraska, Alex Beutel, ... (+3 more)

cs.DB 🏛 SIGMOD 📚 1.2K cites 8 years ago

R.I.P. 👻 Ghosted

Untangling Blockchain: A Data Processing View of Blockchain Systems

Tien Tuan Anh Dinh, Rui Liu, ... (+4 more)

cs.DB 🏛 IEEE TKDE 📚 997 cites 8 years ago

R.I.P. 👻 Ghosted

Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades

Garrick Orchard, Ajinkya Jayawant, ... (+2 more)

cs.DB 🏛 Frontiers in Neuroscience 📚 905 cites 10 years ago

R.I.P. 👻 Ghosted

BLOCKBENCH: A Framework for Analyzing Private Blockchains

Tien Tuan Anh Dinh, Ji Wang, ... (+4 more)

cs.DB 🏛 SIGMOD 📚 872 cites 9 years ago

R.I.P. 👻 Ghosted

Data Synthesis based on Generative Adversarial Networks

Noseong Park, Mahmoud Mohammadi, ... (+4 more)

cs.DB 🏛 VLDB 📚 568 cites 7 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, ... (+29 more)

cs.CL 🏛 NeurIPS 📚 54.2K cites 6 years ago

R.I.P. 👻 Ghosted

PyTorch: An Imperative Style, High-Performance Deep Learning Library

Adam Paszke, Sam Gross, ... (+19 more)

cs.LG 🏛 NeurIPS 📚 49.7K cites 6 years ago

R.I.P. 👻 Ghosted

XGBoost: A Scalable Tree Boosting System

Tianqi Chen, Carlos Guestrin

cs.LG 🏛 KDD 📚 49.2K cites 10 years ago

R.I.P. 👻 Ghosted

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, Christian Szegedy

cs.LG 🏛 ICML 📚 46.0K cites 11 years ago