Hadoop HDFS
  1. Hadoop HDFS
  2. HDFS-5016

Deadlock in pipeline recovery causes Datanode to be marked dead

    Details

    • Type: Bug Bug
    • Status: Closed
    • Priority: Blocker Blocker
    • Resolution: Fixed
    • Affects Version/s: None
    • Fix Version/s: 2.1.0-beta
    • Component/s: None
    • Labels:
      None
    • Target Version/s:
    • Hadoop Flags:
      Reviewed

      Description

      In the testing of some failure scenarios for HBase MTTR, we have been simulating node failures via firewalling of nodes (where all communication ports would be firewalled except ssh's port). We have noticed that when a (data)node is firewalled, we lose certain other datanodes - those that were involved in some communication with the firewalled node before the latter was firewalled. Will attach jstack output from one of the lost datanodes. The heartbeating thread seems to be locked up.
      This jira is to track a fix for the problem.

      1. HDFS-5016.1.patch
        2 kB
        Suresh Srinivas
      2. HDFS-5016.2.patch
        16 kB
        Suresh Srinivas
      3. HDFS-5016.3.patch
        17 kB
        Suresh Srinivas
      4. HDFS-5016.patch
        2 kB
        Suresh Srinivas
      5. jstack1.txt
        213 kB
        Devaraj Das

        Issue Links

          Activity

          No work has yet been logged on this issue.

            People

            • Assignee:
              Suresh Srinivas
              Reporter:
              Devaraj Das
            • Votes:
              0 Vote for this issue
              Watchers:
              15 Start watching this issue

              Dates

              • Created:
                Updated:
                Resolved:

                Development