Performance Characterization of In-Memory Data Analytics on a Modern Cloud Server

June 25, 2015 Β· Declared Dead Β· πŸ› 2015 IEEE Fifth International Conference on Big Data and Cloud Computing

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Ahsan Javed Awan, Mats Brorsson, Vladimir Vlassov, Eduard Ayguade arXiv ID 1506.07742 Category cs.DC: Distributed Computing Cross-listed cs.PF Citations 51 Venue 2015 IEEE Fifth International Conference on Big Data and Cloud Computing Last Checked 5 months ago
Abstract
In last decade, data analytics have rapidly progressed from traditional disk-based processing to modern in-memory processing. However, little effort has been devoted at enhancing performance at micro-architecture level. This paper characterizes the performance of in-memory data analytics using Apache Spark framework. We use a single node NUMA machine and identify the bottlenecks hampering the scalability of workloads. We also quantify the inefficiencies at micro-architecture level for various data analysis workloads. Through empirical evaluation, we show that spark workloads do not scale linearly beyond twelve threads, due to work time inflation and thread level load imbalance. Further, at the micro-architecture level, we observe memory bound latency to be the major cause of work time inflation.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Distributed Computing

Died the same way β€” πŸ‘» Ghosted