Community resources for enabling research in distributed scientific workflows

Rafael Ferreira Da Silva, Weiwei Chen, Gideon Juve, Karan Vahi, Ewa Deelman

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

87 Scopus citations

Abstract

A significant amount of recent research in scientific workflows aims to develop new techniques, algorithms and systems that can overcome the challenges of efficient and robust execution of ever larger workflows on increasingly complex distributed infrastructures. Since the infrastructures, systems and applications are complex, and their behavior is difficult to reproduce using physical experiments, much of this research is based on simulation. However, there exists a shortage of realistic datasets and tools that can be used for such studies. In this paper we describe a collection of tools and data that have enabled research in new techniques, algorithms, and systems for scientific workflows. These resources include: 1) execution traces of real workflow applications from which workflow and system characteristics such as resource usage and failure profiles can be extracted, 2) a synthetic workflow generator that can produce realistic synthetic workflows based on profiles extracted from execution traces, and 3) a simulator framework that can simulate the execution of synthetic workflows on realistic distributed infrastructures. This paper describes how we have used these resources to investigate new techniques for efficient and robust workflow execution, as well as to provide improvements to the Pegasus Workflow Management System or other workflow tools. Our goal in describing these resources is to share them with other researchers in the workflow research community. All of the tools and data are freely available online for the community at http://www.workflowarchive.org. These data have already been leveraged for a number of studies.

Original languageEnglish
Title of host publicationProceedings - 2014 IEEE 10th International Conference on eScience, eScience 2014
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages177-184
Number of pages8
ISBN (Electronic)9781479942886
DOIs
StatePublished - Dec 2 2014
Externally publishedYes
Event10th IEEE International Conference on eScience, eScience 2014 - Guaruja, Brazil
Duration: Oct 20 2014Oct 24 2014

Publication series

NameProceedings - 2014 IEEE 10th International Conference on eScience, eScience 2014
Volume1

Conference

Conference10th IEEE International Conference on eScience, eScience 2014
Country/TerritoryBrazil
CityGuaruja
Period10/20/1410/24/14

Keywords

  • Scientific Workflows
  • Workflow Simulation
  • Workload Archive
  • Workload Profiling and Characterization

Fingerprint

Dive into the research topics of 'Community resources for enabling research in distributed scientific workflows'. Together they form a unique fingerprint.

Cite this