[YARN-7720] Race condition between second app attempt and UAM timeout when first attempt node is down - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Sub-task
Status: Resolved
Priority: Major
Resolution: Fixed
Affects Version/s: 3.4.0
Fix Version/s: 3.4.0
Component/s: federation
Labels:
- pull-request-available

Target Version/s:

3.4.0
Hadoop Flags:

Reviewed

Description

In Federation, multiple attempts of an application share the same UAM in each secondary sub-cluster. When first attempt fails, we reply on the fact that secondary RM won't kill the existing UAM before the AM heartbeat timeout (default at 10 min). When second attempt comes up in the home sub-cluster, it will pick up the UAM token from Yarn Registry and resume the UAM heartbeat to secondary RMs.

The default heartbeat timeout for NM and AM are both 10 mins. The problem is that when the first attempt node goes down or out of connection, only after 10 mins will the home RM mark the first attempt as failed, and then schedule the 2nd attempt in some other node. By then the UAMs in secondaries are already timing out, and they might not survive until the second attempt comes up.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

YARN-7720.v1.patch
30/Nov/18 20:26
4 kB
Botong Huang
YARN-7720.v2.patch
01/Dec/18 00:42
4 kB
Botong Huang

Issue Links

links to

GitHub Pull Request #5672

Activity

People

Assignee:: Shilun Fan

Reporter:: Botong Huang

Votes:: 0 Vote for this issue

Watchers:: 9 Start watching this issue

Dates

Created:: 09/Jan/18 01:19

Updated:: 12/Feb/24 06:49

Resolved:: 29/May/23 17:37