tika-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "ASF GitHub Bot (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (TIKA-2100) Html Parser does not keep the html tag attributes
Date Sat, 26 May 2018 19:23:00 GMT

    [ https://issues.apache.org/jira/browse/TIKA-2100?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16491794#comment-16491794
] 

ASF GitHub Bot commented on TIKA-2100:
--------------------------------------

tballison commented on issue #238: TIKA-2100 extract content language from html lang attribute
URL: https://github.com/apache/tika/pull/238#issuecomment-392282862
 
 
   It feels weird to me to allow such special handling of lang, but if you and
   fellow devs don’t mind, go for it.
   
   On Sat, May 26, 2018 at 2:49 PM Chris Mattmann <notifications@github.com>
   wrote:
   
   > this LGTM - @tballison <https://github.com/tballison> are we good to
   > commit this?
   >
   > —
   > You are receiving this because you were mentioned.
   >
   >
   > Reply to this email directly, view it on GitHub
   > <https://github.com/apache/tika/pull/238#issuecomment-392280901>, or mute
   > the thread
   > <https://github.com/notifications/unsubscribe-auth/AGbWvldSbrJarXkAigHsQiL5z7LgSguGks5t2aOqgaJpZM4UN5BA>
   > .
   >
   

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
users@infra.apache.org


> Html Parser does not keep the html tag attributes
> -------------------------------------------------
>
>                 Key: TIKA-2100
>                 URL: https://issues.apache.org/jira/browse/TIKA-2100
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 1.13
>            Reporter: Gerard Bouchar
>            Priority: Major
>
> Parsing a very simple html like 
>  <!DOCTYPE html>
> <html lang="en">
> <head>
> <title>Page Title</title>
> </head>
> <body>
> <h1 align="left">My First Heading</h1>
> <p>My first paragraph.</p>
> </body>
> </html> 
> you won't be able to access the html tag's attributes (here lang="en") in the ContentHandler
: 
> *in the method startElement(String ns, String localName, String name,
>       Attributes atts), atts is empty.
> *Moreover it seems that the html tag's attributes are not passed trough the HtmlMapper.mapSafeAttribute
method too.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Mime
View raw message