HBase
  1. HBase
  2. HBASE-5081

Distributed log splitting deleteNode races against splitLog retry

    Details

    • Type: Bug Bug
    • Status: Resolved
    • Priority: Major Major
    • Resolution: Fixed
    • Affects Version/s: 0.92.0, 0.94.0
    • Fix Version/s: 0.92.0
    • Component/s: wal
    • Labels:
      None
    • Hadoop Flags:
      Reviewed

      Description

      Recently, during 0.92 rc testing, we found distributed log splitting hangs there forever. Please see attached screen shot.
      I looked into it and here is what happened I think:

      1. One rs died, the servershutdownhandler found it out and started the distributed log splitting;
      2. All three tasks failed, so the three tasks were deleted, asynchronously;
      3. Servershutdownhandler retried the log splitting;
      4. During the retrial, it created these three tasks again, and put them in a hashmap (tasks);
      5. The asynchronously deletion in step 2 finally happened for one task, in the callback, it removed one
      task in the hashmap;
      6. One of the newly submitted tasks' zookeeper watcher found out that task is unassigned, and it is not
      in the hashmap, so it created a new orphan task.
      7. All three tasks failed, but that task created in step 6 is an orphan so the batch.err counter was one short,
      so the log splitting hangs there and keeps waiting for the last task to finish which is never going to happen.

      So I think the problem is step 2. The fix is to make deletion sync, instead of async, so that the retry will have
      a clean start.

      Async deleteNode will mess up with split log retrial. In extreme situation, if async deleteNode doesn't happen
      soon enough, some node created during the retrial could be deleted.

      deleteNode should be sync.

      1. distributed-log-splitting-screenshot.png
        58 kB
        Jimmy Xiang
      2. patch_for_92.txt
        1 kB
        Jimmy Xiang
      3. patch_for_92_v2.txt
        2 kB
        Jimmy Xiang
      4. patch_for_92_v3.txt
        2 kB
        Jimmy Xiang
      5. hbase-5081_patch_for_92_v4.txt
        2 kB
        Jimmy Xiang
      6. hbase-5081_patch_v5.txt
        2 kB
        Jimmy Xiang
      7. hbase-5081-patch-v6.txt
        2 kB
        Jimmy Xiang
      8. hbase-5081-patch-v7.txt
        4 kB
        Jimmy Xiang
      9. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        29 kB
        Prakash Khemani
      10. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        29 kB
        Prakash Khemani
      11. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        29 kB
        Prakash Khemani
      12. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        32 kB
        Prakash Khemani
      13. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        33 kB
        Prakash Khemani
      14. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        34 kB
        Prakash Khemani
      15. 5081-deleteNode-with-while-loop.txt
        31 kB
        Ted Yu
      16. 0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        38 kB
        Prakash Khemani
      17. HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
        37 kB
        Ted Yu
      18. distributed_log_splitting_screen_shot2.png
        125 kB
        Jimmy Xiang
      19. distributed_log_splitting_screenshot3.png
        286 kB
        Jimmy Xiang

        Issue Links

          Activity

          Jimmy Xiang created issue -
          Jimmy Xiang made changes -
          Field Original Value New Value
          Attachment distributed-log-splitting-screenshot.png [ 12508194 ]
          Jimmy Xiang made changes -
          Attachment patch_for_92.txt [ 12508259 ]
          Jimmy Xiang made changes -
          Status Open [ 1 ] Patch Available [ 10002 ]
          Jimmy Xiang made changes -
          Attachment patch_for_92_v2.txt [ 12508264 ]
          Jimmy Xiang made changes -
          Attachment patch_for_92_v3.txt [ 12508267 ]
          Jimmy Xiang made changes -
          Attachment hbase-5081_patch_for_92_v4.txt [ 12508272 ]
          Jimmy Xiang made changes -
          Attachment hbase-5081_patch_v5.txt [ 12508273 ]
          Jimmy Xiang made changes -
          Status Patch Available [ 10002 ] Open [ 1 ]
          Jimmy Xiang made changes -
          Attachment hbase-5081-patch-v6.txt [ 12508297 ]
          Jimmy Xiang made changes -
          Status Open [ 1 ] Patch Available [ 10002 ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12508264/patch_for_92_v2.txt
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              -1 tests included. The patch doesn't appear to include any new or modified tests.
                                  Please justify why no new tests are needed for this patch.
                                  Also please list what manual steps were performed to verify this patch.

              -1 javadoc. The javadoc tool appears to have generated -152 warning messages.

              +1 javac. The applied patch does not increase the total number of javac compiler warnings.

              -1 findbugs. The patch appears to introduce 76 new Findbugs (version 1.3.9) warnings.

              +1 release audit. The applied patch does not increase the total number of release audit warnings.

               -1 core tests. The patch failed these unit tests:
                                 org.apache.hadoop.hbase.replication.TestReplication
                            org.apache.hadoop.hbase.replication.TestMultiSlaveReplication
                            org.apache.hadoop.hbase.replication.TestMasterReplication

          Test results: https://builds.apache.org/job/PreCommit-HBASE-Build/570//testReport/
          Findbugs warnings: https://builds.apache.org/job/PreCommit-HBASE-Build/570//artifact/trunk/patchprocess/newPatchFindbugsWarnings.html
          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/570//console

          This message is automatically generated. ]
          Ted Yu made changes -
          Comment [
          -----------------------------------------------------------
          This is an automatically generated e-mail. To reply, visit:
          https://reviews.apache.org/r/3292/
          -----------------------------------------------------------

          (Updated 2011-12-21 17:01:22.901024)


          Review request for hbase, Ted Yu, Michael Stack, and Lars Hofhansl.


          Changes
          -------

          Updated the comments.


          Summary
          -------

          In this patch, after a task is done, we don't delete the node if the task is failed. So that when it's retried later on, there won't be race problem.

          It used to delete the node always.


          This addresses bug HBASE-5081.
              https://issues.apache.org/jira/browse/HBASE-5081


          Diffs (updated)
          -----

            src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java 7b7316f

          Diff: https://reviews.apache.org/r/3292/diff


          Testing
          -------

          mvn -Dtest=TestDistributedLogSplitting clean test


          Thanks,

          Jimmy

          ]
          Ted Yu made changes -
          Comment [

          bq. On 2011-12-21 17:06:22, Michael Stack wrote:
          bq. > src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java, line 358
          bq. > <https://reviews.apache.org/r/3292/diff/2/?file=65660#file65660line358>
          bq. >
          bq. > White space on end and 'is' should be 'if'

          Let me fix it.


          - Jimmy


          -----------------------------------------------------------
          This is an automatically generated e-mail. To reply, visit:
          https://reviews.apache.org/r/3292/#review4046
          -----------------------------------------------------------


          On 2011-12-21 17:01:22, Jimmy Xiang wrote:
          bq.
          bq. -----------------------------------------------------------
          bq. This is an automatically generated e-mail. To reply, visit:
          bq. https://reviews.apache.org/r/3292/
          bq. -----------------------------------------------------------
          bq.
          bq. (Updated 2011-12-21 17:01:22)
          bq.
          bq.
          bq. Review request for hbase, Ted Yu, Michael Stack, and Lars Hofhansl.
          bq.
          bq.
          bq. Summary
          bq. -------
          bq.
          bq. In this patch, after a task is done, we don't delete the node if the task is failed. So that when it's retried later on, there won't be race problem.
          bq.
          bq. It used to delete the node always.
          bq.
          bq.
          bq. This addresses bug HBASE-5081.
          bq. https://issues.apache.org/jira/browse/HBASE-5081
          bq.
          bq.
          bq. Diffs
          bq. -----
          bq.
          bq. src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java 7b7316f
          bq.
          bq. Diff: https://reviews.apache.org/r/3292/diff
          bq.
          bq.
          bq. Testing
          bq. -------
          bq.
          bq. mvn -Dtest=TestDistributedLogSplitting clean test
          bq.
          bq.
          bq. Thanks,
          bq.
          bq. Jimmy
          bq.
          bq.

          ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12508273/hbase-5081_patch_v5.txt
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              -1 tests included. The patch doesn't appear to include any new or modified tests.
                                  Please justify why no new tests are needed for this patch.
                                  Also please list what manual steps were performed to verify this patch.

              -1 javadoc. The javadoc tool appears to have generated -152 warning messages.

              +1 javac. The applied patch does not increase the total number of javac compiler warnings.

              -1 findbugs. The patch appears to introduce 76 new Findbugs (version 1.3.9) warnings.

              +1 release audit. The applied patch does not increase the total number of release audit warnings.

               -1 core tests. The patch failed these unit tests:
                                 org.apache.hadoop.hbase.coprocessor.TestCoprocessorEndpoint
                            org.apache.hadoop.hbase.mapred.TestTableMapReduce
                            org.apache.hadoop.hbase.mapreduce.TestHFileOutputFormat
                            org.apache.hadoop.hbase.master.TestSplitLogManager

          Test results: https://builds.apache.org/job/PreCommit-HBASE-Build/572//testReport/
          Findbugs warnings: https://builds.apache.org/job/PreCommit-HBASE-Build/572//artifact/trunk/patchprocess/newPatchFindbugsWarnings.html
          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/572//console

          This message is automatically generated. ]
          Jimmy Xiang made changes -
          Status Patch Available [ 10002 ] Open [ 1 ]
          Jimmy Xiang made changes -
          Attachment hbase-5081-patch-v7.txt [ 12508328 ]
          Jimmy Xiang made changes -
          Status Open [ 1 ] Patch Available [ 10002 ]
          stack made changes -
          Status Patch Available [ 10002 ] Resolved [ 5 ]
          Hadoop Flags Reviewed [ 10343 ]
          Fix Version/s 0.92.0 [ 12314223 ]
          Resolution Fixed [ 1 ]
          stack made changes -
          Resolution Fixed [ 1 ]
          Status Resolved [ 5 ] Reopened [ 4 ]
          Ted Yu made changes -
          Comment [ Integrated in HBase-0.92 #208 (See [https://builds.apache.org/job/HBase-0.92/208/])
              HBASE-5081 Distributed log splitting deleteNode races againsth splitLog retry; REVERT -- COMMITTED BEFORE REVIEW FINISHED
          HBASE-5081 Distributed log splitting deleteNode races againsth splitLog retry

          stack :
          Files :
          * /hbase/branches/0.92/CHANGES.txt
          * /hbase/branches/0.92/src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java
          * /hbase/branches/0.92/src/test/java/org/apache/hadoop/hbase/master/TestSplitLogManager.java

          stack :
          Files :
          * /hbase/branches/0.92/CHANGES.txt
          * /hbase/branches/0.92/src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java
          * /hbase/branches/0.92/src/test/java/org/apache/hadoop/hbase/master/TestSplitLogManager.java
          ]
          Ted Yu made changes -
          Comment [ Integrated in HBase-TRUNK #2567 (See [https://builds.apache.org/job/HBase-TRUNK/2567/])
              HBASE-5081 Distributed log splitting deleteNode races againsth splitLog retry; REVERT -- COMMITTED BEFORE REVIEW FINISHED
          HBASE-5081 Distributed log splitting deleteNode races againsth splitLog retry

          stack :
          Files :
          * /hbase/trunk/CHANGES.txt
          * /hbase/trunk/src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java
          * /hbase/trunk/src/test/java/org/apache/hadoop/hbase/master/TestSplitLogManager.java

          stack :
          Files :
          * /hbase/trunk/CHANGES.txt
          * /hbase/trunk/src/main/java/org/apache/hadoop/hbase/master/SplitLogManager.java
          * /hbase/trunk/src/test/java/org/apache/hadoop/hbase/master/TestSplitLogManager.java
          ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12508259/patch_for_92.txt
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              -1 tests included. The patch doesn't appear to include any new or modified tests.
                                  Please justify why no new tests are needed for this patch.
                                  Also please list what manual steps were performed to verify this patch.

              -1 javadoc. The javadoc tool appears to have generated -152 warning messages.

              +1 javac. The applied patch does not increase the total number of javac compiler warnings.

              -1 findbugs. The patch appears to introduce 76 new Findbugs (version 1.3.9) warnings.

              +1 release audit. The applied patch does not increase the total number of release audit warnings.

               -1 core tests. The patch failed these unit tests:
                                 org.apache.hadoop.hbase.client.TestInstantSchemaChange
                            org.apache.hadoop.hbase.mapred.TestTableMapReduce
                            org.apache.hadoop.hbase.mapreduce.TestHFileOutputFormat
                            org.apache.hadoop.hbase.master.TestSplitLogManager

          Test results: https://builds.apache.org/job/PreCommit-HBASE-Build/569//testReport/
          Findbugs warnings: https://builds.apache.org/job/PreCommit-HBASE-Build/569//artifact/trunk/patchprocess/newPatchFindbugsWarnings.html
          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/569//console

          This message is automatically generated. ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12508297/hbase-5081-patch-v6.txt
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              +1 tests included. The patch appears to include 3 new or modified tests.

              -1 javadoc. The javadoc tool appears to have generated -152 warning messages.

              +1 javac. The applied patch does not increase the total number of javac compiler warnings.

              -1 findbugs. The patch appears to introduce 76 new Findbugs (version 1.3.9) warnings.

              +1 release audit. The applied patch does not increase the total number of release audit warnings.

               -1 core tests. The patch failed these unit tests:
                                 org.apache.hadoop.hbase.replication.TestReplication

          Test results: https://builds.apache.org/job/PreCommit-HBASE-Build/573//testReport/
          Findbugs warnings: https://builds.apache.org/job/PreCommit-HBASE-Build/573//artifact/trunk/patchprocess/newPatchFindbugsWarnings.html
          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/573//console

          This message is automatically generated. ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12508328/hbase-5081-patch-v7.txt
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              +1 tests included. The patch appears to include 3 new or modified tests.

              -1 javadoc. The javadoc tool appears to have generated -152 warning messages.

              +1 javac. The applied patch does not increase the total number of javac compiler warnings.

              -1 findbugs. The patch appears to introduce 76 new Findbugs (version 1.3.9) warnings.

              +1 release audit. The applied patch does not increase the total number of release audit warnings.

               -1 core tests. The patch failed these unit tests:
                                 org.apache.hadoop.hbase.replication.TestReplication

          Test results: https://builds.apache.org/job/PreCommit-HBASE-Build/575//testReport/
          Findbugs warnings: https://builds.apache.org/job/PreCommit-HBASE-Build/575//artifact/trunk/patchprocess/newPatchFindbugsWarnings.html
          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/575//console

          This message is automatically generated. ]
          Jimmy Xiang made changes -
          Assignee Jimmy Xiang [ jxiang ] Prakash Khemani [ khemani ]
          Prakash Khemani made changes -
          Ted Yu made changes -
          Status Reopened [ 4 ] Patch Available [ 10002 ]
          Jean-Daniel Cryans made changes -
          Summary Distributed log splitting deleteNode races againsth splitLog retry Distributed log splitting deleteNode races against splitLog retry
          Prakash Khemani made changes -
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12509344/0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              +1 tests included. The patch appears to include 8 new or modified tests.

              -1 patch. The patch command could not apply the patch.

          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/656//console

          This message is automatically generated. ]
          Prakash Khemani made changes -
          Prakash Khemani made changes -
          Prakash Khemani made changes -
          Prakash Khemani made changes -
          Ted Yu made changes -
          Attachment 5081-deleteNode-with-while-loop.txt [ 12509487 ]
          stack made changes -
          Status Patch Available [ 10002 ] Open [ 1 ]
          stack made changes -
          Status Open [ 1 ] Patch Available [ 10002 ]
          Ted Yu made changes -
          Attachment 5081-deleteNode-with-while-loop.txt [ 12509487 ]
          Ted Yu made changes -
          Attachment 5081-deleteNode-with-while-loop.txt [ 12509507 ]
          Prakash Khemani made changes -
          Ted Yu made changes -
          Jimmy Xiang made changes -
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12509653/distributed_log_splitting_screen_shot2.png
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              -1 tests included. The patch doesn't appear to include any new or modified tests.
                                  Please justify why no new tests are needed for this patch.
                                  Also please list what manual steps were performed to verify this patch.

              -1 patch. The patch command could not apply the patch.

          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/682//console

          This message is automatically generated. ]
          Ted Yu made changes -
          Comment [ -1 overall. Here are the results of testing the latest attachment
            http://issues.apache.org/jira/secure/attachment/12509352/0001-HBASE-5081-jira-Distributed-log-splitting-deleteNode.patch
            against trunk revision .

              +1 @author. The patch does not contain any @author tags.

              +1 tests included. The patch appears to include 8 new or modified tests.

              -1 patch. The patch command could not apply the patch.

          Console output: https://builds.apache.org/job/PreCommit-HBASE-Build/657//console

          This message is automatically generated. ]
          Ted Yu made changes -
          Link This issue is related to HBASE-5136 [ HBASE-5136 ]
          Jimmy Xiang made changes -
          Ted Yu made changes -
          Status Patch Available [ 10002 ] Resolved [ 5 ]
          Resolution Fixed [ 1 ]

            People

            • Assignee:
              Prakash Khemani
              Reporter:
              Jimmy Xiang
            • Votes:
              0 Vote for this issue
              Watchers:
              6 Start watching this issue

              Dates

              • Created:
                Updated:
                Resolved:

                Development