tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Alberto Ornaghi (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (TIKA-1080) Arabic characters under windows
Date Thu, 07 Feb 2013 14:07:13 GMT

    [ https://issues.apache.org/jira/browse/TIKA-1080?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13573512#comment-13573512

Alberto Ornaghi commented on TIKA-1080:

perfect. problem solved. thank you.
> Arabic characters under windows
> -------------------------------
>                 Key: TIKA-1080
>                 URL: https://issues.apache.org/jira/browse/TIKA-1080
>             Project: Tika
>          Issue Type: Bug
>          Components: parser, server
>    Affects Versions: 1.3
>         Environment: Windows 2003 or Windows 2008
>            Reporter: Alberto Ornaghi
>         Attachments: arabic.docx
> If tika is executed under windows the text mode (--text) is failing to extract arabic
chars and outputs only question marks. The same behaviour occurs if tika is executed as a
server. The issue is not present in the GUI, only commandline. The issue is not present if
the output is html.

This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira

View raw message