Uploaded image for project: 'Hadoop HDFS'
  1. Hadoop HDFS
  2. HDFS-5016

Deadlock in pipeline recovery causes Datanode to be marked dead

    XMLWordPrintableJSON

Details

    • Bug
    • Status: Closed
    • Blocker
    • Resolution: Fixed
    • None
    • 2.1.0-beta
    • None
    • None
    • Reviewed

    Description

      In the testing of some failure scenarios for HBase MTTR, we have been simulating node failures via firewalling of nodes (where all communication ports would be firewalled except ssh's port). We have noticed that when a (data)node is firewalled, we lose certain other datanodes - those that were involved in some communication with the firewalled node before the latter was firewalled. Will attach jstack output from one of the lost datanodes. The heartbeating thread seems to be locked up.
      This jira is to track a fix for the problem.

      Attachments

        1. HDFS-5016.3.patch
          17 kB
          Suresh Srinivas
        2. HDFS-5016.2.patch
          16 kB
          Suresh Srinivas
        3. HDFS-5016.1.patch
          2 kB
          Suresh Srinivas
        4. HDFS-5016.patch
          2 kB
          Suresh Srinivas
        5. jstack1.txt
          213 kB
          Devaraj Das

        Issue Links

          Activity

            People

              sureshms Suresh Srinivas
              ddas Devaraj Das
              Votes:
              0 Vote for this issue
              Watchers:
              15 Start watching this issue

              Dates

                Created:
                Updated:
                Resolved: