Skip to main navigation Skip to search Skip to main content

Towards Predicting the Impact of Roll-Forward Failure Recovery for HPC Applications

  • Bo Fang
  • , Jieyang Chen
  • , Karthik Pattabiraman
  • , Matei Ripeanu
  • , Sriram Krishnamoorthy

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

The roll-forward recovery schemes on HPC systems implicitly trade off faster time to solution for higher risk: as it usually performs a probabilistic repair, this may cause further failures such as SDCs. It is essential for users to be able to reason about the impact of a particular repair exercised by the scheme. Towards this goal, we identify two research questions aiming to determine the outcome of a repair either at the failure point or at the end of the execution. For the former, we propose a promising hybrid approach that combines machine learning and error propagation analysis techniques.

Original languageEnglish
Title of host publicationProceedings - 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks - Supplemental Volume, DSN-S 2019
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages13-14
Number of pages2
ISBN (Electronic)9781728130286
DOIs
StatePublished - Jun 2019
Externally publishedYes
Event49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN-S 2019 - Portland, United States
Duration: Jun 24 2019Jun 27 2019

Publication series

NameProceedings - 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks - Supplemental Volume, DSN-S 2019

Conference

Conference49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN-S 2019
Country/TerritoryUnited States
CityPortland
Period06/24/1906/27/19

Funding

ACKNOWLEDGEMENT This work was supported in part by the U.S. Department of Energy‘s (DOE) Office of Science, the National Sciences and Engineering Research Council of Canada (NSERC), Office of AdvancedScientific Computing Research, under award 66905. Pacific Northwest National Laboratory is operated by Battelle for DOE under Contract DE-AC05-76RL01830. This work was supported in part by the U.S. Department of Energy's (DOE) Office of Science, the National Sciences and Engineering Research Council of Canada (NSERC), Office of Advanced Scientific Computing Research, under award 66905. Pacific Northwest National Laboratory is operated by Battelle for DOE under Contract DE-AC05-76RL01830.

Keywords

  • Checkpoint/restart
  • Fault tolerance
  • HPC
  • Machine learning
  • Roll forward recovery

Fingerprint

Dive into the research topics of 'Towards Predicting the Impact of Roll-Forward Failure Recovery for HPC Applications'. Together they form a unique fingerprint.

Cite this