[HDFS-5016] Deadlock in pipeline recovery causes Datanode to be marked dead - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Blocker
Resolution: Fixed
Affects Version/s: None
Fix Version/s: 2.1.0-beta
Component/s: None
Labels:
None

Target Version/s:

2.1.0-beta
Hadoop Flags:

Reviewed

Description

In the testing of some failure scenarios for HBase MTTR, we have been simulating node failures via firewalling of nodes (where all communication ports would be firewalled except ssh's port). We have noticed that when a (data)node is firewalled, we lose certain other datanodes - those that were involved in some communication with the firewalled node before the latter was firewalled. Will attach jstack output from one of the lost datanodes. The heartbeating thread seems to be locked up.
This jira is to track a fix for the problem.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

jstack1.txt
21/Jul/13 07:37
213 kB
Devaraj Das
HDFS-5016.patch
21/Jul/13 23:18
2 kB
Suresh Srinivas
HDFS-5016.3.patch
26/Jul/13 03:27
17 kB
Suresh Srinivas
HDFS-5016.2.patch
26/Jul/13 00:46
16 kB
Suresh Srinivas
HDFS-5016.1.patch
25/Jul/13 02:06
2 kB
Suresh Srinivas

Issue Links

duplicates

HDFS-3655 Datanode recoverRbw could hang sometime

Resolved

Activity

People

Assignee:: Suresh Srinivas

Reporter:: Devaraj Das

Votes:: 0 Vote for this issue

Watchers:: 14 Start watching this issue

Dates

Created:: 21/Jul/13 07:36

Updated:: 27/Aug/13 22:07

Resolved:: 26/Jul/13 05:58