[NUTCH-1308] Add main() to ZipParser - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Closed
Priority: Minor
Resolution: Fixed
Affects Version/s: 1.4, nutchgora
Fix Version/s: 1.13
Component/s: parser
Labels:
None

Description

Two issues here...

1) Recently ferdy committed ~~NUTCH-965~~ which skips parsing of truncated documents. Parse zip has it's own implementation for the same when it should really draw on the aforementioned implementation.
2) If (in the offending piece of code mentioned above) truncation occurs, we get an incorrect log message the "Parser can't handle incomplete pdf files"!!! This is incorrect, shouldn't be there, and should be removed.

72      if (contentLen != null && contentInBytes.length != len) {
73 	return new ParseStatus(ParseStatus.FAILED,
74 	ParseStatus.FAILED_TRUNCATED, "Content truncated at "
75 	+ contentInBytes.length
76 	+ " bytes. Parser can't handle incomplete pdf file.")
77 	.getEmptyParseResult(content.getUrl(), getConf());
78 	}

For clarity, the issue is present in both Nutchgora branch[1] and Nutch trunk[2]

[1] https://svn.apache.org/viewvc/nutch/branches/nutchgora/src/plugin/parse-zip/src/java/org/apache/nutch/parse/zip/ZipParser.java?diff_format=h&view=markup
[2] https://svn.apache.org/viewvc/nutch/trunk/src/plugin/parse-zip/src/java/org/apache/nutch/parse/zip/ZipParser.java?diff_format=h&view=markup
[2]

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

NUTCH-1308-ZipParser-main-trunk.patch
16/Apr/14 22:38
2 kB
Sebastian Nagel

Issue Links

is duplicated by

NUTCH-1603 ZIP parser complains about truncated PDF file

Closed

Activity

People

Assignee:: Sebastian Nagel

Reporter:: Lewis John McGibbney

Votes:: 0 Vote for this issue

Watchers:: 3 Start watching this issue

Dates

Created:: 09/Mar/12 16:38

Updated:: 13/Mar/24 14:50

Resolved:: 02/Jul/16 10:46