[HADOOP-3646] Providing bzip2 as codec - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 0.19.0
Fix Version/s: 0.19.0
Component/s: conf, io
Labels:
None

Hadoop Flags:

Reviewed
Release Note:
Introduced support for bzip2 compressed files.

Description

Hadoop recognizes gzip compressed input and automatically decompresses the data before providing it to the mapper. But Hadoop can not split a gzip stream due to the very nature of the gzip compression. Consequently one gzip stream (e.g a whole file) can go to only one mapper. On the contrary Bzip2 compressed stream can be split across its block delimiters.

We are interested in extending Hadoop to support splittable bzip2 with a codec. (https://issues.apache.org/jira/browse/HADOOP-1823 uses input reader to split the bzip2 files, which must be provided by the user and can handle FileInputFormat. If a user wants to use some other input format or wants to do custom record handling, he must write a new input reader!)

We have a patch now that provides a basic bzip2 codec equivalent to the current gzip codec. We are in the process of extending that to support splitting.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

HADOOP-3646.patch
11/Jul/08 19:04
130 kB
Abdul Qadeer
HADOOP-3646.patch
03/Jul/08 07:30
134 kB
Abdul Qadeer
HADOOP-3646version3.patch
17/Jul/08 09:25
107 kB
Abdul Qadeer
HADOOP-3646-version4.patch
18/Jul/08 07:29
109 kB
Abdul Qadeer
HADOOP-3646-version5.patch
29/Jul/08 05:57
108 kB
Abdul Qadeer

Issue Links

is blocked by

HADOOP-5326 bzip2 codec (CBZip2OutputStream) creates corrupted output file for some inputs

Closed

is related to

HADOOP-4918 Fix bzip2 work with SequenceFile

Closed

HADOOP-1823 want InputFormat for bzip2 files

Resolved

HADOOP-4012 Providing splitting support for bzip2 compressed files

Closed

MAPREDUCE-772 Chaging LineRecordReader algo so that it does not need to skip backwards in the stream

Closed

HADOOP-5379 Throw exception instead of writing to System.err when there is a CRC error on CBZip2InputStream

Closed

(1 is related to)

Activity

People

Assignee:: Abdul Qadeer

Reporter:: Abdul Qadeer

Votes:: 0 Vote for this issue

Watchers:: 10 Start watching this issue

Dates

Created:: 25/Jun/08 22:41

Updated:: 10/Mar/09 02:27

Resolved:: 29/Jul/08 19:01

Time Tracking

Estimated:

1,008h

Remaining:

1,008h

Logged:

Not Specified