[HADOOP-5539] o.a.h.mapred.Merger not maintaining map out compression on intermediate files - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Bug
Status: Resolved
Priority: Blocker
Resolution: Fixed
Affects Version/s: 0.19.1
Fix Version/s: 0.20.1
Component/s: None
Labels:
None
Environment:

0.19.2-dev, r753365

Hadoop Flags:

Reviewed

Description

hadoop-site.xml :
mapred.compress.map.output = true

map output files are compressed but when the in memory merger closes
on the reduce the on disk merger runs to reduce input files to <= io.sort.factor if needed.

when this happens it outputs files called intermediate.x files these
do not maintain compression setting the writer (o.a.h.mapred.Merger.class line 432)
passes the codec but I added some logging and its always null map output compression set true or false.

This causes task to fail if they can not hold the uncompressed size of the data of the reduce its holding
I thank this is just and oversight of the codec not getting set correctly for the on disk merges.

2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: Merging 30 intermediate segments out of a total of 3000
2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: intermediate.1 used codec: null

I added

          // added my me
	   if (codec != null){
	     LOG.info("intermediate." + passNo + " used codec: " + codec.toString());
	   } else {
	     LOG.info("intermediate." + passNo + " used codec: Null");
	   }
	   // end added by me

Just before the creation of the writer o.a.h.mapred.Merger.class line 432
and it outputs the second line above.

I have confirmed this with the logging and I have looked at the files on the disk of the tasktracker. I can read the data in
the intermediate files clearly telling me that there not compressed but I can not read the map.out files direct from the map output
telling me the compression is working on the map end but not on the on disk merge that produces the intermediate.

I can see no benefit for these not maintaining the compression setting and as it looks they where intended to maintain it.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

5539.patch
22/Mar/09 08:59
5 kB
Billy Pearson
hadoop-5539.patch
15/May/09 07:18
6 kB
Jothi Padmanabhan
hadoop-5539-v1.patch
21/May/09 03:30
5 kB
Jothi Padmanabhan
hadoop-5539-branch20.patch
27/May/09 09:39
4 kB
Jothi Padmanabhan

Activity

People

Assignee:: Jothi Padmanabhan

Reporter:: Billy Pearson

Votes:: 1 Vote for this issue

Watchers:: 7 Start watching this issue

Dates

Created:: 20/Mar/09 06:51

Updated:: 08/Jul/09 16:53

Resolved:: 28/May/09 11:37