tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Dave Meikle (JIRA)" <j...@apache.org>
Subject [jira] [Resolved] (TIKA-973) PDF form data isn't included in extracted content.
Date Tue, 04 Feb 2014 23:18:10 GMT

     [ https://issues.apache.org/jira/browse/TIKA-973?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Dave Meikle resolved TIKA-973.
------------------------------

    Resolution: Fixed

> PDF form data isn't included in extracted content.
> --------------------------------------------------
>
>                 Key: TIKA-973
>                 URL: https://issues.apache.org/jira/browse/TIKA-973
>             Project: Tika
>          Issue Type: Bug
>          Components: general
>    Affects Versions: 1.2
>            Reporter: Michael Graessle
>            Assignee: Tim Allison
>            Priority: Minor
>             Fix For: 1.5
>
>         Attachments: TIKA-973-patch.tar.gz, TIKA-973.patch.tar.gz, i-9_screenshot.png
>
>
> When extracting content from PDFs, PDF form data isn't extracted. 
> The following code extracts this data via PDF box, but it seems like something Tika should
be doing.
> PDDocumentCatalog docCatalog = load.getDocumentCatalog();
> if (docCatalog != null) {
>   PDAcroForm acroForm = docCatalog.getAcroForm();
>   if (acroForm != null) {
> 	@SuppressWarnings("unchecked")
> 	List<PDField> fields = acroForm.getFields();
> 	if (fields != null && fields.size() > 0) {
> 	  documentContent.append(" ");
> 	  for (PDField field : fields) {
> 		if (field.getValue()!=null) {
> 		  documentContent.append(field.getValue());
> 		  documentContent.append(" ");
> 		}
> 	  }
> 	}
>   }
> }



--
This message was sent by Atlassian JIRA
(v6.1.5#6160)

Mime
View raw message