nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "ASF GitHub Bot (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (NUTCH-2442) Injector to stop if job fails to avoid loss of CrawlDb
Date Sat, 04 Nov 2017 17:38:00 GMT

    [ https://issues.apache.org/jira/browse/NUTCH-2442?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16239121#comment-16239121
] 

ASF GitHub Bot commented on NUTCH-2442:
---------------------------------------

Omkar20895 opened a new pull request #239: NUTCH-2442 Injector to stop if job fails to avoid
loss of CrawlDb
URL: https://github.com/apache/nutch/pull/239
 
 
   - Added Job status checks in the classes: Injector, ReadHostDb, CrawlCompletionStats, ProtocolStatusStatistics,
SitemapProcessor and DomainStatistics. 

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
users@infra.apache.org


> Injector to stop if job fails to avoid loss of CrawlDb
> ------------------------------------------------------
>
>                 Key: NUTCH-2442
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2442
>             Project: Nutch
>          Issue Type: Bug
>          Components: injector
>    Affects Versions: 1.13
>            Reporter: Sebastian Nagel
>            Priority: Critical
>             Fix For: 1.14
>
>
> Injector does not check whether the MapReduce job is successful. Even if the job fails
> - installs the CrawlDb
> -- move current/ to old/
> -- replace current/ with an empty or potentially incomplete version
> - exits with code 0 so that scripts running the crawl workflow cannot detect the failure
-- if Injector is run a second time the CrawlDb is lost (both current/ and old/ are empty
or corrupted)



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

Mime
View raw message