Details
-
Sub-task
-
Status: Resolved
-
Major
-
Resolution: Fixed
-
None
-
None
-
None
Description
I'd like to add a field(retry-times) in ContainerLaunchContext. When AM launches containers, it could specify the value. Then NM will re-launch the container 'retry-times' times when it fails to run(e.g.exit code is not 0).
It will save a lot of time. It avoids container localization. RM does not need to re-schedule the container. And local files in container's working directory will be left for re-use.(If container have downloaded some big files, it does not need to re-download them when running again.)
We find it is useful in systems like Storm.
Attachments
Attachments
Issue Links
- contains
-
YARN-8609 NM oom because of large container statuses
- Resolved
- is related to
-
YARN-5547 NMLeveldbStateStore should be more tolerant of unknown keys
- Resolved
- is required by
-
SLIDER-1174 Support Tensorflow on Slider
- Resolved