nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Hudson (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (NUTCH-2038) Naive Bayes classifier based html Parse filter (for filtering outlinks)
Date Mon, 29 Jun 2015 05:50:04 GMT

    [ https://issues.apache.org/jira/browse/NUTCH-2038?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14605169#comment-14605169
] 

Hudson commented on NUTCH-2038:
-------------------------------

FAILURE: Integrated in Nutch-trunk #3181 (See [https://builds.apache.org/job/Nutch-trunk/3181/])
fix for NUTCH-2038: Naive Bayes classifier based html Parse filter (for filtering outlinks)
contributed by Asitang Mishra <asitang@gmail.com> this closes #39 (mattmann: http://svn.apache.org/viewvc/nutch/trunk/?view=rev&rev=1688085)
* /nutch/trunk/CHANGES.txt
fix for NUTCH-2038: Naive Bayes classifier based html Parse filter (for filtering outlinks)
contributed by Asitang Mishra <asitang@gmail.com> this closes #39 (mattmann: http://svn.apache.org/viewvc/nutch/trunk/?view=rev&rev=1688084)
* /nutch/trunk/.gitignore
* /nutch/trunk/build.xml
* /nutch/trunk/conf/nutch-default.xml
* /nutch/trunk/ivy/ivy.xml
* /nutch/trunk/src/plugin/build.xml
* /nutch/trunk/src/plugin/parsefilter-naivebayes
* /nutch/trunk/src/plugin/parsefilter-naivebayes/build.xml
* /nutch/trunk/src/plugin/parsefilter-naivebayes/ivy.xml
* /nutch/trunk/src/plugin/parsefilter-naivebayes/plugin.xml
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch/parsefilter
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch/parsefilter/naivebayes
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch/parsefilter/naivebayes/NaiveBayesClassifier.java
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch/parsefilter/naivebayes/NaiveBayesParseFilter.java
* /nutch/trunk/src/plugin/parsefilter-naivebayes/src/java/org/apache/nutch/parsefilter/naivebayes/package-info.java


> Naive Bayes classifier based html Parse filter (for filtering outlinks)
> -----------------------------------------------------------------------
>
>                 Key: NUTCH-2038
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2038
>             Project: Nutch
>          Issue Type: New Feature
>          Components: fetcher, injector, parser
>            Reporter: Asitang Mishra
>            Assignee: Chris A. Mattmann
>              Labels: memex, nutch
>             Fix For: 1.11
>
>
> A html parse filter that will filter out the outlinks in two stages. 
> Classify the parse text and decide if the parent page is relevant. If relevant then don't
filter the outlinks. If irrelevant then go thru each outlink and see if the url contains any
of the important words from a list. If it does then let it pass.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Mime
View raw message