tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Md (JIRA)" <j...@apache.org>
Subject [jira] [Updated] (TIKA-2593) docx with track change producing incorrect output
Date Thu, 01 Mar 2018 12:35:00 GMT

     [ https://issues.apache.org/jira/browse/TIKA-2593?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Md updated TIKA-2593:
---------------------
    Description: 
I am using following code to extract text from docx file 
{code:java}
AutoDetectParser parser = new AutoDetectParser();
ContentHandler contentHandler = new BodyContentHandler();
inputStream = new BufferedInputStream(new FileInputStream(inputFileName));
Metadata metadata = new Metadata();

OfficeParserConfig officeParserConfig = new OfficeParserConfig();
officeParserConfig.setIncludeDeletedContent(false);
parseContext.set(OfficeParserConfig.class, officeParserConfig);

parser.parse(inputStream, contentHandler, metadata, parseContext);
System.out.println(contentHandler.toString());
{code}
When I am sending track revised files it's adding all the text deleted with the actual text
and inserted text. Is there a way to tell parser to exclude the deleted text?

Here is an example 

input Text: This is a sample text. -This part will- be deleted. +This is inserted.+

outputText: This is a sample text. This part will be deleted. This is inserted.

Desired output: This is a sample text.  be deleted. This is inserted.

  was:
I am using following code to extract text from docx file 
{code:java}
contentHandler = new BodyContentHandler();
inputStream = new BufferedInputStream(new FileInputStream(inputFileName));
Metadata metadata = new Metadata();
StringBuilder fileContent = new StringBuilder();
recursiveParserWrapper.parse(inputStream, contentHandler, metadata, parseContext);
System.out.println("Metadata WordCount Value: "+contentHandler.toString());

{code}
When I am sending track revised files it's adding all the text deleted with the actual text
and inserted text. Is there a way to tell parser to exclude the deleted text?

Here is an example 

input Text: This is a sample text. -This part will- be deleted. +This is inserted.+

outputText: This is a sample text. This part will be deleted. This is inserted.

Desired output: This is a sample text.  be deleted. This is inserted.


> docx with track change producing incorrect output
> -------------------------------------------------
>
>                 Key: TIKA-2593
>                 URL: https://issues.apache.org/jira/browse/TIKA-2593
>             Project: Tika
>          Issue Type: Bug
>          Components: core, handler
>    Affects Versions: 1.17
>            Reporter: Md
>            Priority: Major
>         Attachments: sample.docx
>
>
> I am using following code to extract text from docx file 
> {code:java}
> AutoDetectParser parser = new AutoDetectParser();
> ContentHandler contentHandler = new BodyContentHandler();
> inputStream = new BufferedInputStream(new FileInputStream(inputFileName));
> Metadata metadata = new Metadata();
> OfficeParserConfig officeParserConfig = new OfficeParserConfig();
> officeParserConfig.setIncludeDeletedContent(false);
> parseContext.set(OfficeParserConfig.class, officeParserConfig);
> parser.parse(inputStream, contentHandler, metadata, parseContext);
> System.out.println(contentHandler.toString());
> {code}
> When I am sending track revised files it's adding all the text deleted with the actual
text and inserted text. Is there a way to tell parser to exclude the deleted text?
> Here is an example 
> input Text: This is a sample text. -This part will- be deleted. +This is inserted.+
> outputText: This is a sample text. This part will be deleted. This is inserted.
> Desired output: This is a sample text.  be deleted. This is inserted.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Mime
View raw message