[YARN-7913] Improve error handling when application recovery fails with exception - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Resolved
Priority: Major
Resolution: Fixed
Affects Version/s: 3.0.0
Fix Version/s: 3.3.0, 3.2.2, 3.1.4
Component/s: resourcemanager
Labels:
None

Hadoop Flags:

Reviewed

Description

There are edge cases when the application recovery fails with an exception.

Example failure scenario:

setup: a queue is a leaf queue in the primary RM's config and the same queue is a parent queue in the secondary RM's config.
When failover happens with this setup, the recovery will fail for applications on this queue, and an APP_REJECTED event will be dispatched to the async dispatcher. On the same thread (that handles the recovery), a NullPointerException is thrown when the applicationAttempt is tried to be recovered (https://github.com/apache/hadoop/blob/55066cc53dc22b68f9ca55a0029741d6c846be0a/hadoop-yarn-project/hadoop-yarn/hadoop-yarn-server/hadoop-yarn-server-resourcemanager/src/main/java/org/apache/hadoop/yarn/server/resourcemanager/scheduler/fair/FairScheduler.java#L494). I don't see a good way to avoid the NPE in this scenario, because when the NPE occurs the APP_REJECTED has not been processed yet, and we don't know that the application recovery failed.

Currently the first exception will abort the recovery, and if there are X applications, there will be ~X passive -> active RM transition attempts - the passive -> active RM transition will only succeed when the last APP_REJECTED event is processed on the async dispatcher thread.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

YARN-7913-branch-3.2.001.patch
21/Jan/20 13:52
12 kB
wilfreds#1
YARN-7913-branch-3.1.001.patch
21/Jan/20 13:52
12 kB
wilfreds#1
YARN-7913-branch-3.1.001.patch
22/Jan/20 11:57
12 kB
Szilard Nemeth
YARN-7913.003.patch
07/Jan/20 03:08
12 kB
wilfreds#1
YARN-7913.002.patch
06/Jan/20 23:54
12 kB
wilfreds#1
YARN-7913.001.patch
06/Jan/20 10:42
12 kB
wilfreds#1
YARN-7913.000.poc.patch
09/Feb/18 11:05
2 kB
Gergo Repas

Issue Links

is duplicated by

YARN-10290 Resourcemanager recover failed when fair scheduler queue acl changed

Resolved

YARN-7998 RM crashes with NPE during recovering if ACL configuration was changed

Resolved

relates to

YARN-10046 RM failed to transition to Active because of App recovery throwing java.lang.NullPointerException

Open

Activity

People

Assignee:: Wilfred Spiegelenburg

Reporter:: Gergo Repas

Votes:: 0 Vote for this issue

Watchers:: 7 Start watching this issue

Dates

Created:: 09/Feb/18 10:59

Updated:: 14/Jul/21 20:30

Resolved:: 22/Jan/20 15:52