nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Sebastian Nagel (JIRA)" <j...@apache.org>
Subject [jira] [Updated] (NUTCH-1308) Unnecessary truncate content configuration, and logging in parse-zip/ZipParser
Date Wed, 16 Apr 2014 22:40:15 GMT

     [ https://issues.apache.org/jira/browse/NUTCH-1308?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Sebastian Nagel updated NUTCH-1308:
-----------------------------------

    Attachment: NUTCH-1308-ZipParser-main-trunk.patch

Hi [~lewismc], is this fixed with NUTCH-1603?
Attached a minimalist main for ZipParser (definitely useful).

> Unnecessary truncate content configuration, and logging in parse-zip/ZipParser  
> --------------------------------------------------------------------------------
>
>                 Key: NUTCH-1308
>                 URL: https://issues.apache.org/jira/browse/NUTCH-1308
>             Project: Nutch
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 1.4, nutchgora
>            Reporter: Lewis John McGibbney
>            Assignee: Lewis John McGibbney
>             Fix For: 2.4
>
>         Attachments: NUTCH-1308-ZipParser-main-trunk.patch
>
>
> Two issues here...
> 1) Recently ferdy committed NUTCH-965 which skips parsing of truncated documents. Parse
zip has it's own implementation for the same when it should really draw on the aforementioned
implementation.
> 2) If (in the offending piece of code mentioned above) truncation occurs, we get an incorrect
log message the "Parser can't handle incomplete pdf files"!!! This is incorrect, shouldn't
be there, and should be removed.
> {code}
> 72      if (contentLen != null && contentInBytes.length != len) {
> 73 	return new ParseStatus(ParseStatus.FAILED,
> 74 	ParseStatus.FAILED_TRUNCATED, "Content truncated at "
> 75 	+ contentInBytes.length
> 76 	+ " bytes. Parser can't handle incomplete pdf file.")
> 77 	.getEmptyParseResult(content.getUrl(), getConf());
> 78 	}
> {code}
> For clarity, the issue is present in both Nutchgora branch[1] and Nutch trunk[2]
> [1] https://svn.apache.org/viewvc/nutch/branches/nutchgora/src/plugin/parse-zip/src/java/org/apache/nutch/parse/zip/ZipParser.java?diff_format=h&view=markup
> [2] https://svn.apache.org/viewvc/nutch/trunk/src/plugin/parse-zip/src/java/org/apache/nutch/parse/zip/ZipParser.java?diff_format=h&view=markup
> [2] 



--
This message was sent by Atlassian JIRA
(v6.2#6252)

Mime
View raw message