tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "ASF GitHub Bot (Jira)" <j...@apache.org>
Subject [jira] [Commented] (TIKA-2224) OneNote formats support - Mime Magic and Parser
Date Tue, 10 Dec 2019 16:54:00 GMT

    [ https://issues.apache.org/jira/browse/TIKA-2224?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16992736#comment-16992736
] 

ASF GitHub Bot commented on TIKA-2224:
--------------------------------------

tballison commented on issue #300: TIKA-2224 - OneNote parser
URL: https://github.com/apache/tika/pull/300#issuecomment-564126279
 
 
   Need to review statics to make sure this parser will be thread safe.
   Remove json unless critical.
   
   I'm working on these and a few of the above now.
 
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
users@infra.apache.org


> OneNote formats support - Mime Magic and Parser
> -----------------------------------------------
>
>                 Key: TIKA-2224
>                 URL: https://issues.apache.org/jira/browse/TIKA-2224
>             Project: Tika
>          Issue Type: Improvement
>          Components: mime
>    Affects Versions: 1.14
>            Reporter: Nick Burch
>            Priority: Major
>         Attachments: Sample1.json, Sample1.one, note-ssn-test-mmmm.one
>
>
> As raised at http://stackoverflow.com/questions/41272195/onenote-support-for-apache-tika-parsers,
we don't have any magic for the OneNote formats. Several years ago we dug out the file format
specs (see http://lucene.472066.n3.nabble.com/Tika-OneNote-Support-td4020393.html), but didn't
have volunteer energy to implement a parser. However, armed with those specs, we should be
able to come up with some mime magic for detection



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

Mime
View raw message