Hadoop Map/Reduce
  1. Hadoop Map/Reduce
  2. MAPREDUCE-2450

Calls from running tasks to TaskTracker methods sometimes fail and incur a 60s timeout

    Details

    • Type: Bug Bug
    • Status: Closed
    • Priority: Major Major
    • Resolution: Fixed
    • Affects Version/s: 0.23.0, 2.0.0-alpha
    • Fix Version/s: 0.23.1
    • Component/s: None
    • Labels:
      None
    • Tags:
      mrv2

      Description

      I'm seeing some map tasks in my jobs take 1 minute to commit after they finish the map computation. On the map side, the output looks like this:

      <code>
      2009-03-02 21:30:54,384 INFO org.apache.hadoop.metrics.jvm.JvmMetrics: Cannot initialize JVM Metrics with processName=MAP, sessionId= - already initialized
      2009-03-02 21:30:54,437 INFO org.apache.hadoop.mapred.MapTask: numReduceTasks: 800
      2009-03-02 21:30:54,437 INFO org.apache.hadoop.mapred.MapTask: io.sort.mb = 300
      2009-03-02 21:30:55,493 INFO org.apache.hadoop.mapred.MapTask: data buffer = 239075328/298844160
      2009-03-02 21:30:55,494 INFO org.apache.hadoop.mapred.MapTask: record buffer = 786432/983040
      2009-03-02 21:31:00,381 INFO org.apache.hadoop.mapred.MapTask: Starting flush of map output
      2009-03-02 21:31:07,892 INFO org.apache.hadoop.mapred.MapTask: Finished spill 0
      2009-03-02 21:31:07,951 INFO org.apache.hadoop.mapred.TaskRunner: Task:attempt_200903022127_0001_m_003163_0 is done. And is in the process of commiting
      2009-03-02 21:32:07,949 INFO org.apache.hadoop.mapred.TaskRunner: Communication exception: java.io.IOException: Call to /127.0.0.1:50311 failed on local exception: java.nio.channels.ClosedChannelException
      at org.apache.hadoop.ipc.Client.wrapException(Client.java:765)
      at org.apache.hadoop.ipc.Client.call(Client.java:733)
      at org.apache.hadoop.ipc.RPC$Invoker.invoke(RPC.java:220)
      at org.apache.hadoop.mapred.$Proxy0.ping(Unknown Source)
      at org.apache.hadoop.mapred.Task$TaskReporter.run(Task.java:525)
      at java.lang.Thread.run(Thread.java:619)
      Caused by: java.nio.channels.ClosedChannelException
      at java.nio.channels.spi.AbstractSelectableChannel.register(AbstractSelectableChannel.java:167)
      at java.nio.channels.SelectableChannel.register(SelectableChannel.java:254)
      at org.apache.hadoop.net.SocketIOWithTimeout$SelectorPool.select(SocketIOWithTimeout.java:331)
      at org.apache.hadoop.net.SocketIOWithTimeout.doIO(SocketIOWithTimeout.java:157)
      at org.apache.hadoop.net.SocketInputStream.read(SocketInputStream.java:155)
      at org.apache.hadoop.net.SocketInputStream.read(SocketInputStream.java:128)
      at java.io.FilterInputStream.read(FilterInputStream.java:116)
      at org.apache.hadoop.ipc.Client$Connection$PingInputStream.read(Client.java:276)
      at java.io.BufferedInputStream.fill(BufferedInputStream.java:218)
      at java.io.BufferedInputStream.read(BufferedInputStream.java:237)
      at java.io.DataInputStream.readInt(DataInputStream.java:370)
      at org.apache.hadoop.ipc.Client$Connection.receiveResponse(Client.java:501)
      at org.apache.hadoop.ipc.Client$Connection.run(Client.java:446)

      2009-03-02 21:32:07,953 INFO org.apache.hadoop.mapred.TaskRunner: Task 'attempt_200903022127_0001_m_003163_0' done.
      </code>

      In the TaskTracker log, it looks like this:

      <code>
      2009-03-02 21:31:08,110 WARN org.apache.hadoop.ipc.Server: IPC Server Responder, call ping(attempt_200903022127_0001_m_003163_0) from 127.0.0.1:56884: output error
      2009-03-02 21:31:08,111 INFO org.apache.hadoop.ipc.Server: IPC Server handler 10 on 50311 caught: java.nio.channels.ClosedChannelException
      at sun.nio.ch.SocketChannelImpl.ensureWriteOpen(SocketChannelImpl.java:126)
      at sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:324) at org.apache.hadoop.ipc.Server.channelWrite(Server.java:1195)
      at org.apache.hadoop.ipc.Server.access$1900(Server.java:77)
      at org.apache.hadoop.ipc.Server$Responder.processResponse(Server.java:613)
      at org.apache.hadoop.ipc.Server$Responder.doRespond(Server.java:677)
      at org.apache.hadoop.ipc.Server$Handler.run(Server.java:981)
      </code>

      Note that the task actually seemed to commit - it didn't get speculatively executed or anything. However, the job wasn't able to continue until this one task was done. Both parties seem to think the channel was closed. How does the channel get closed externally? If closing it from outside is unavoidable, maybe the right thing to do is to set a much lower timeout, because 1 minute delay can be pretty significant for a small job.

      1. HADOOP_5380-Y.0.20.20x.patch
        2 kB
        Rajesh Balamohan
      2. HADOOP-5380.patch
        2 kB
        Rajesh Balamohan
      3. HADOOP-5380.Y.20.branch.patch
        0.6 kB
        Rajesh Balamohan
      4. mapreduce-2450.patch
        2 kB
        Rajesh Balamohan
      5. MAPREDUCE-2450.patch
        3 kB
        Arun C Murthy

        Activity

        Matei Zaharia created issue -
        Matei Zaharia made changes -
        Field Original Value New Value
        Affects Version/s 0.20.1 [ 12313866 ]
        Rajesh Balamohan made changes -
        Attachment HADOOP-5380.Y.20.branch.patch [ 12466783 ]
        Rajesh Balamohan made changes -
        Attachment HADOOP_5380-Y.0.20.20x.patch [ 12467854 ]
        Rajesh Balamohan made changes -
        Attachment HADOOP-5380.patch [ 12477210 ]
        Rajesh Balamohan made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Rajesh Balamohan made changes -
        Status Patch Available [ 10002 ] Open [ 1 ]
        Fix Version/s 0.23.0 [ 12315569 ]
        Rajesh Balamohan made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Affects Version/s 0.23.0 [ 12315569 ]
        Affects Version/s 0.20.1 [ 12313866 ]
        Rajesh Balamohan made changes -
        Project Hadoop Common [ 12310240 ] Hadoop Map/Reduce [ 12310941 ]
        Key HADOOP-5380 MAPREDUCE-2450
        Affects Version/s 0.23.0 [ 12315570 ]
        Affects Version/s 0.23.0 [ 12315569 ]
        Fix Version/s 0.23.0 [ 12315570 ]
        Fix Version/s 0.23.0 [ 12315569 ]
        Rajesh Balamohan made changes -
        Status Patch Available [ 10002 ] Open [ 1 ]
        Rajesh Balamohan made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Rajesh Balamohan made changes -
        Status Patch Available [ 10002 ] Open [ 1 ]
        Rajesh Balamohan made changes -
        Attachment mapreduce-2450.patch [ 12477611 ]
        Rajesh Balamohan made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Todd Lipcon made changes -
        Comment [ hi
        this is the used cars.
        [url=http://www.sexyeditor.com/]used cars[/url] ]
        Todd Lipcon made changes -
        Comment [ hi
        this is the used cars.
        [link=used cars]"http://www.sexyeditor.com" rel="dofollow"[/link] ]
        Todd Lipcon made changes -
        Comment [ hi
        this is the used cars.
        <a href="http://www.sexyeditor.com" rel="dofollow">used cars</a> ]
        Todd Lipcon made changes -
        Comment [ hi
        this is the used cars.
        [url="http://www.sexyeditor.com" rel="dofollow"]used cars[/url] ]
        Arun C Murthy made changes -
        Status Patch Available [ 10002 ] Open [ 1 ]
        Tags mrv2
        Arun C Murthy made changes -
        Fix Version/s 0.24.0 [ 12317654 ]
        Fix Version/s 0.23.0 [ 12315570 ]
        Arun C Murthy made changes -
        Attachment MAPREDUCE-2450.patch [ 12510796 ]
        Arun C Murthy made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Target Version/s 0.23.1 [ 12318883 ]
        Assignee Rajesh Balamohan [ rajesh.balamohan ]
        Fix Version/s 0.24.0 [ 12317654 ]
        Arun C Murthy made changes -
        Status Patch Available [ 10002 ] Resolved [ 5 ]
        Target Version/s 0.23.1 [ 12318883 ]
        Fix Version/s 1.0.0 [ 12318240 ]
        Resolution Fixed [ 1 ]
        Vinod Kumar Vavilapalli made changes -
        Fix Version/s 0.23.1 [ 12318883 ]
        Fix Version/s 0.24.0 [ 12317654 ]
        Affects Version/s 0.24.0 [ 12317654 ]
        Vinod Kumar Vavilapalli made changes -
        Fix Version/s 0.24.0 [ 12317654 ]
        Fix Version/s 1.0.0 [ 12318240 ]
        Arun C Murthy made changes -
        Status Resolved [ 5 ] Closed [ 6 ]
        Allen Wittenauer made changes -
        Affects Version/s 2.0.0-alpha [ 12320354 ]
        Affects Version/s 0.24.0 [ 12317654 ]

          People

          • Assignee:
            Rajesh Balamohan
            Reporter:
            Matei Zaharia
          • Votes:
            2 Vote for this issue
            Watchers:
            13 Start watching this issue

            Dates

            • Created:
              Updated:
              Resolved:

              Development