[HADOOP-2606] Namenode unstable when replicating 500k blocks at once - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 0.14.3
Fix Version/s: 0.17.0
Component/s: None
Labels:
None

Description

We tried to decommission about 40 nodes at once, each containing 12k blocks. (about 500k total)
(This also happened when we first tried to decommission 2 million blocks)

Clients started experiencing "java.lang.RuntimeException: java.net.SocketTimeoutException: timed out waiting for rpc
response" and namenode was in 100% cpu state.

It was spending most of its time on one thread,

"org.apache.hadoop.dfs.FSNamesystem$ReplicationMonitor@7f401d28" daemon prio=10 tid=0x0000002e10702800 nid=0x6718
runnable [0x0000000041a42000..0x0000000041a42a30]
java.lang.Thread.State: RUNNABLE
at org.apache.hadoop.dfs.FSNamesystem.containingNodeList(FSNamesystem.java:2766)
at org.apache.hadoop.dfs.FSNamesystem.pendingTransfers(FSNamesystem.java:2870)

locked <0x0000002aa3cef720> (a org.apache.hadoop.dfs.UnderReplicatedBlocks)
locked <0x0000002aa3c42e28> (a org.apache.hadoop.dfs.FSNamesystem)
at org.apache.hadoop.dfs.FSNamesystem.computeDatanodeWork(FSNamesystem.java:1928)
at org.apache.hadoop.dfs.FSNamesystem$ReplicationMonitor.run(FSNamesystem.java:1868)
at java.lang.Thread.run(Thread.java:619)

We confirmed that Namenode was not in the fullGC states when these problem happened.

Also, dfsadmin -metasave was showing "Blocks waiting for replication" was decreasing very slowly.

I believe this is not specific to decommission and same problem would happen if we lose one rack.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

ReplicatorNew.patch
14/Mar/08 11:49
43 kB
Konstantin Shvachko
ReplicatorNew1.patch
18/Mar/08 03:12
47 kB
Konstantin Shvachko
ReplicatorNew2.patch
18/Mar/08 23:33
47 kB
Konstantin Shvachko
ReplicatorTestOld.patch
14/Mar/08 11:37
38 kB
Konstantin Shvachko

Issue Links

duplicates

HDFS-150 Replication should be decoupled from heartbeat

Resolved

is related to

HDFS-373 Name node should notify administrator if when struggling with replication

Open

relates to

HDFS-150 Replication should be decoupled from heartbeat

Resolved

HADOOP-2649 The ReplicationMonitor sleep period should be configurable

Closed

HADOOP-2755 dfs fsck extremely slow, dfs ls times out

Closed

Activity

People

Assignee:: Konstantin Shvachko

Reporter:: Koji Noguchi

Votes:: 0 Vote for this issue

Watchers:: 3 Start watching this issue

Dates

Created:: 14/Jan/08 23:58

Updated:: 17/Jul/14 18:37

Resolved:: 19/Mar/08 00:16