在非同步的多任务方案编制模式中努力实现分布式软件复原力 (Towards Distributed Software Resilience in Asynchronous Many-Task Programming Models) - 专知论文

会员服务 ·

0

评论员 · 缩放 · 样板代码 · 容差 · MoDELS ·

2020 年 10 月 19 日

Towards Distributed Software Resilience in Asynchronous Many-Task Programming Models

翻译：在非同步的多任务方案编制模式中努力实现分布式软件复原力

Nikunj Gupta,Jackson R. Mayo,Adrian S. Lemoine,Hartmut Kaiser

from arxiv, arXiv admin note: text overlap with arXiv:2004.07203

Exceptions and errors occurring within mission critical applications due to hardware failures have a high cost. With the emerging Next Generation Platforms (NGPs), the rate of hardware failures will likely increase. Therefore, designing our applications to be resilient is a critical concern in order to retain the reliability of results while meeting the constraints on power budgets. In this paper, we discuss software resilience in AMTs at both local and distributed scale. We choose HPX to prototype our resiliency designs. We implement two resiliency APIs that we expose to the application developers, namely task replication and task replay. Task replication repeats a task n-times and executes them asynchronously. Task replay reschedules a task up to n-times until a valid output is returned. Furthermore, we expose algorithm based fault tolerance (ABFT) using user provided predicates (e.g., checksums) to validate the returned results. We benchmark the resiliency scheme for both synthetic and real world applications at local and distributed scale and show that most of the added execution time arises from the replay, replication or data movement of the tasks and not the boilerplate code added to achieve resilience.

翻译：由于硬件故障,在任务关键应用程序中出现的例外和错误成本很高。随着下一代平台(NGPs)的出现,硬件故障率可能会上升。因此,设计我们的应用程序以保持弹性是一个关键问题,目的是在满足对电力预算的限制的同时保持结果的可靠性。在本文中,我们讨论本地和分布规模的AMT的软件复原力。我们选择HPX来原型我们的恢复能力设计。我们实施了两种弹性API,我们暴露在应用程序开发者面前,即任务复制和任务重播。任务复制重复了任务正时,并且不同步地执行它们。任务重播任务重排一个任务,直到n-时间,直到返回有效产出。此外,我们用用户提供的前提(例如校验和)来披露基于算法的过失容忍(ABFT)来验证返回的结果。我们为本地和分布规模的合成和真实世界应用的弹性计划设定基准,并显示大部分增加的执行时间来自任务的重置、复制或数据移动,而不是锅炉板的代码来实现恢复能力。

0

相关内容

评论员

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

79+阅读 · 2020年7月26日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

35+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

180+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

104+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

人工智能 | ISAIR 2019诚邀稿件（推荐SCI期刊）

人工智能 | ISAIR 2019诚邀稿件（推荐SCI期刊）

Call4Papers

6+阅读 · 2019年4月1日

人工智能 | SCI期刊专刊信息3条

人工智能 | SCI期刊专刊信息3条

Call4Papers

5+阅读 · 2019年1月10日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

17+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

ssd-sgd: communication sparsification for distributed deep learning training

Arxiv

0+阅读 · 2020年12月10日

Semi-Dynamic Load Balancing: Efficient Distributed Learning in Non-Dedicated Environments

Arxiv

0+阅读 · 2020年12月8日

Odyssey: Creation, Analysis and Detection of Trojan Models

Arxiv

0+阅读 · 2020年12月8日

Transition-Oriented Programming: Developing Verifiable Systems

Arxiv

0+阅读 · 2020年12月8日

Cost-effective Machine Learning Inference Offload for Edge Computing

Arxiv

0+阅读 · 2020年12月7日

Transdisciplinary AI Observatory -- Retrospective Analyses and Future-Oriented Contradistinctions

Arxiv

0+阅读 · 2020年12月7日

Low-Latency Asynchronous Logic Design for Inference at the Edge

Arxiv

0+阅读 · 2020年12月7日

Optimal Caching for Low Latency in Distributed Coded Storage Systems

Arxiv

0+阅读 · 2020年12月5日

Asynchronous Federated Optimization

Arxiv

0+阅读 · 2020年12月5日

Distributed Constraint Optimization Problems and Applications: A Survey

Arxiv

5+阅读 · 2018年1月11日

VIP会员

文章信息

相关主题

相关VIP内容

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

79+阅读 · 2020年7月26日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

35+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

180+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

104+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

中文版 | 美国会在最终的25财年国防协议中缩减了预算目标

《发展“敏捷战斗部署”所需的作战支援任务就绪空勤人员》最新102页报告

《人工智能在决策中角色的演变》最新278页

中文版 | 近程防空系统的必要性日益凸显

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

人工智能 | ISAIR 2019诚邀稿件（推荐SCI期刊）

人工智能 | ISAIR 2019诚邀稿件（推荐SCI期刊）

Call4Papers

6+阅读 · 2019年4月1日

人工智能 | SCI期刊专刊信息3条

人工智能 | SCI期刊专刊信息3条

Call4Papers

5+阅读 · 2019年1月10日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

17+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

强化学习 cartpole_a3c

强化学习 cartpole_a3c

CreateAMind

9+阅读 · 2017年7月21日

相关论文

ssd-sgd: communication sparsification for distributed deep learning training

Arxiv

0+阅读 · 2020年12月10日

Semi-Dynamic Load Balancing: Efficient Distributed Learning in Non-Dedicated Environments

Arxiv

0+阅读 · 2020年12月8日

Odyssey: Creation, Analysis and Detection of Trojan Models

Arxiv

0+阅读 · 2020年12月8日

Transition-Oriented Programming: Developing Verifiable Systems

Arxiv

0+阅读 · 2020年12月8日

Cost-effective Machine Learning Inference Offload for Edge Computing

Arxiv

0+阅读 · 2020年12月7日

Transdisciplinary AI Observatory -- Retrospective Analyses and Future-Oriented Contradistinctions

Arxiv

0+阅读 · 2020年12月7日

Low-Latency Asynchronous Logic Design for Inference at the Edge

Arxiv

0+阅读 · 2020年12月7日

Optimal Caching for Low Latency in Distributed Coded Storage Systems

Arxiv

0+阅读 · 2020年12月5日

Asynchronous Federated Optimization

Arxiv

0+阅读 · 2020年12月5日

Distributed Constraint Optimization Problems and Applications: A Survey

Arxiv

5+阅读 · 2018年1月11日

微信扫码咨询专知VIP会员