PARALLEL DATA LAB 

PDL Abstract

Hadoop's Adolescence: An Analysis of Hadoop Usage in Scientific Workloads

Very Large Data Bases 2013 (VLDB13), Riva del Garda, Italy, Aug 26-30, 2013.

Kai Ren1, YongChul Kwon2, Magdalena Balazinska3, Bill Howe3

1School of Computer Science, Carnegie Mellon University
2 Microsoft,
3 University of Washington

http://www.pdl.cmu.edu

We analyze Hadoop workloads from three different research clusters from a user-centric perspective. The goal is to better understand data scientists' use of the system and how well the use of the system matches its design. Our analysis suggests that Hadoop usage is still in its adolescence. We see underuse of Hadoop features, extensions, and tools. We see significant diversity in resource usage and application styles, including some interactive and iterative workloads, motivating new tools in the ecosystem. We also observe significant opportunities for optimizations of these workloads. We find that job customization and configuration are used in a narrow scope, suggesting the future pursuit of automatic tuning systems. Overall, we present the first user-centered measurement study of Hadoop and find significant opportunities for improving its efficient use for data scientists.

FULL PAPER: pdf