nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Chris A. Mattmann (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (NUTCH-2038) Naive Bayes classifier based html Parse filter (for filtering outlinks)
Date Wed, 01 Jul 2015 13:13:05 GMT

    [ https://issues.apache.org/jira/browse/NUTCH-2038?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14610201#comment-14610201
] 

Chris A. Mattmann commented on NUTCH-2038:
------------------------------------------

Hey [~markus.jelsma@openindex.io] yeah we tried to insulate the dependencies to being a plugin,
but for whatever reason when doing so, it doesn't seem to work? Seb and Asitang and I tried
it - I think it has to do with the PluginClasspathLoading. Anyways, this is something we should
try and figure out for 1.11 release to see if we can get out of there, but not a blocker now
since the functionality is pretty neat and since we're just talking about more jar files (that
don't conflict with anything) in the lib directory.

[~asitang] can you open a new issue to try and figure out how to get the dependencies into
the plugin's ivy and to still make it work?

> Naive Bayes classifier based html Parse filter (for filtering outlinks)
> -----------------------------------------------------------------------
>
>                 Key: NUTCH-2038
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2038
>             Project: Nutch
>          Issue Type: New Feature
>          Components: fetcher, injector, parser
>            Reporter: Asitang Mishra
>            Assignee: Chris A. Mattmann
>              Labels: memex, nutch
>             Fix For: 1.11
>
>
> A html parse filter that will filter out the outlinks in two stages. 
> Classify the parse text and decide if the parent page is relevant. If relevant then don't
filter the outlinks. If irrelevant then go thru each outlink and see if the url contains any
of the important words from a list. If it does then let it pass.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Mime
View raw message