tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Marc Teutelink (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (TIKA-1199) Tika extracts weird signs instead of text
Date Thu, 21 Nov 2013 14:57:35 GMT

    [ https://issues.apache.org/jira/browse/TIKA-1199?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13828990#comment-13828990
] 

Marc Teutelink commented on TIKA-1199:
--------------------------------------

I did the same thing just now. Maybe the PDF file uses a specific font that is embedded in
the file an maps to images of karakters instead of ASCI or Unicode characters. I tried selecting
and copying it and that doesn't work either, which more or less proofs my theory.

> Tika extracts weird signs instead of text
> -----------------------------------------
>
>                 Key: TIKA-1199
>                 URL: https://issues.apache.org/jira/browse/TIKA-1199
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 1.4
>         Environment: MacOSX, Linux
>            Reporter: Marc Teutelink
>         Attachments: gaat fout.pdf, plain_text_tika_output_from_gaat_fout_pdf.txt, structured_text_tika_output_from_gaat_fout_pdf.xml
>
>
> Tika extracts complete bogus text from the attached document. I have attached the .PDF
in question and also added the plain and structured text output from Tika.



--
This message was sent by Atlassian JIRA
(v6.1#6144)

Mime
View raw message