Uploaded image for project: 'Hadoop HDFS'
  1. Hadoop HDFS
  2. HDFS-5016

Deadlock in pipeline recovery causes Datanode to be marked dead

    Details

    • Type: Bug
    • Status: Closed
    • Priority: Blocker
    • Resolution: Fixed
    • Affects Version/s: None
    • Fix Version/s: 2.1.0-beta
    • Component/s: None
    • Labels:
      None
    • Target Version/s:
    • Hadoop Flags:
      Reviewed

      Description

      In the testing of some failure scenarios for HBase MTTR, we have been simulating node failures via firewalling of nodes (where all communication ports would be firewalled except ssh's port). We have noticed that when a (data)node is firewalled, we lose certain other datanodes - those that were involved in some communication with the firewalled node before the latter was firewalled. Will attach jstack output from one of the lost datanodes. The heartbeating thread seems to be locked up.
      This jira is to track a fix for the problem.

        Attachments

        1. jstack1.txt
          213 kB
          Devaraj Das
        2. HDFS-5016.patch
          2 kB
          Suresh Srinivas
        3. HDFS-5016.3.patch
          17 kB
          Suresh Srinivas
        4. HDFS-5016.2.patch
          16 kB
          Suresh Srinivas
        5. HDFS-5016.1.patch
          2 kB
          Suresh Srinivas

          Issue Links

            Activity

              People

              • Assignee:
                sureshms Suresh Srinivas
                Reporter:
                devaraj Devaraj Das
              • Votes:
                0 Vote for this issue
                Watchers:
                14 Start watching this issue

                Dates

                • Created:
                  Updated:
                  Resolved: