Hadoop Common
  1. Hadoop Common
  2. HADOOP-5861

s3n files are not getting split by default

    Details

    • Type: Bug Bug
    • Status: Closed
    • Priority: Major Major
    • Resolution: Fixed
    • Affects Version/s: 0.19.1
    • Fix Version/s: 0.21.0
    • Component/s: fs/s3
    • Labels:
      None
    • Environment:

      ec2

    • Hadoop Flags:
      Incompatible change, Reviewed
    • Release Note:
      Files stored on the native S3 filesystem (s3n:// URIs) now report a block size determined by the fs.s3n.block.size property (default 64MB).

      Description

      running with stock ec2 scripts against hadoop-19 - i tried to run a job against a directory with 4 text files - each about 2G in size. These were not split (only 4 mappers were run).

      The reason seems to have two parts - primarily that S3N files report a block size of 5G. This causes FileInputFormat.getSplits to fall back on goal size (which is totalsize/conf.get("mapred.map.tasks")).Goal Size in this case was 4G - hence the files were not split. This is not an issue with other file systems since the block size reported is much smaller and the splits get based on block size (not goal size).

      can we make the S3N files report a more reasonable block size?

      1. hadoop-5861-v2.patch
        4 kB
        Tom White
      2. hadoop-5861.patch
        4 kB
        Tom White

        Activity

        Joydeep Sen Sarma created issue -
        Tom White made changes -
        Field Original Value New Value
        Attachment hadoop-5861.patch [ 12409188 ]
        Tom White made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Assignee Tom White [ tomwhite ]
        Tom White made changes -
        Attachment hadoop-5861-v2.patch [ 12409550 ]
        Tom White made changes -
        Status Patch Available [ 10002 ] Open [ 1 ]
        Tom White made changes -
        Status Open [ 1 ] Patch Available [ 10002 ]
        Tom White made changes -
        Status Patch Available [ 10002 ] Resolved [ 5 ]
        Hadoop Flags [Incompatible change, Reviewed]
        Release Note Files stored on the native S3 filesystem (s3n:// URIs) now report a block size determined by the fs.s3n.block.size property (default 64MB).
        Fix Version/s 0.21.0 [ 12313563 ]
        Resolution Fixed [ 1 ]
        Tom White made changes -
        Status Resolved [ 5 ] Closed [ 6 ]

          People

          • Assignee:
            Tom White
            Reporter:
            Joydeep Sen Sarma
          • Votes:
            0 Vote for this issue
            Watchers:
            4 Start watching this issue

            Dates

            • Created:
              Updated:
              Resolved:

              Development