tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Ken Krugler (JIRA)" <j...@apache.org>
Subject [jira] [Comment Edited] (TIKA-1723) Integrate language-detector into Tika
Date Tue, 01 Sep 2015 21:46:46 GMT

    [ https://issues.apache.org/jira/browse/TIKA-1723?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14726205#comment-14726205
] 

Ken Krugler edited comment on TIKA-1723 at 9/1/15 9:46 PM:
-----------------------------------------------------------

Version 2 of my patch (not be be confused with Tim's patch, which is about moving this code
into a new tika-langdetect module)


was (Author: kkrugler):
Version 2 of my patch (not be be confused with Tim's patch, which is about moving this code
into a new tika-language module)

> Integrate language-detector into Tika
> -------------------------------------
>
>                 Key: TIKA-1723
>                 URL: https://issues.apache.org/jira/browse/TIKA-1723
>             Project: Tika
>          Issue Type: Improvement
>          Components: languageidentifier
>    Affects Versions: 1.11
>            Reporter: Ken Krugler
>            Assignee: Ken Krugler
>            Priority: Minor
>         Attachments: TIKA-1723-2.patch, TIKA-1723.patch, TIKA-1723v2.patch
>
>
> The language-detector project at https://github.com/optimaize/language-detector is faster,
has more languages (70 vs 13) and better accuracy than the built-in language detector.
> This is a stab at integrating it, with some initial findings. There are a number of issues
this raises, especially if [~chrismattmann] moves forward with turning language detection
into a pluggable extension point.
> I'll add comments with results below.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Mime
View raw message