[HADOOP-2606] Namenode unstable when replicating 500k blocks at once - ASF JIRA

Voters

Watch issue

Watchers

Create sub-task

Link

Clone

Update Comment Author

Replace String in Comment

Update Comment Visibility

Delete Comments

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 0.14.3
Fix Version/s: 0.17.0
Component/s: None
Labels:
None

Description

We tried to decommission about 40 nodes at once, each containing 12k blocks. (about 500k total)
(This also happened when we first tried to decommission 2 million blocks)

Clients started experiencing "java.lang.RuntimeException: java.net.SocketTimeoutException: timed out waiting for rpc
response" and namenode was in 100% cpu state.

It was spending most of its time on one thread,

"org.apache.hadoop.dfs.FSNamesystem$ReplicationMonitor@7f401d28" daemon prio=10 tid=0x0000002e10702800 nid=0x6718
runnable [0x0000000041a42000..0x0000000041a42a30]
java.lang.Thread.State: RUNNABLE
at org.apache.hadoop.dfs.FSNamesystem.containingNodeList(FSNamesystem.java:2766)
at org.apache.hadoop.dfs.FSNamesystem.pendingTransfers(FSNamesystem.java:2870)

locked <0x0000002aa3cef720> (a org.apache.hadoop.dfs.UnderReplicatedBlocks)
locked <0x0000002aa3c42e28> (a org.apache.hadoop.dfs.FSNamesystem)
at org.apache.hadoop.dfs.FSNamesystem.computeDatanodeWork(FSNamesystem.java:1928)
at org.apache.hadoop.dfs.FSNamesystem$ReplicationMonitor.run(FSNamesystem.java:1868)
at java.lang.Thread.run(Thread.java:619)

We confirmed that Namenode was not in the fullGC states when these problem happened.

Also, dfsadmin -metasave was showing "Blocks waiting for replication" was decreasing very slowly.

I believe this is not specific to decommission and same problem would happen if we lose one rack.