lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Yonik Seeley (JIRA)" <j...@apache.org>
Subject [jira] Commented: (LUCENE-1799) Unicode compression
Date Wed, 28 Jul 2010 20:57:16 GMT

    [ https://issues.apache.org/jira/browse/LUCENE-1799?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12893362#action_12893362
] 

Yonik Seeley commented on LUCENE-1799:
--------------------------------------

bq. so if you want to improve indexing speed, i think you will be more successful looking
at other parts of the code.

I have only been measuring performance at this point, and I haven't expressed an option about
what defaults should be used.
If we convert to BOCU-1 as a default, and if UTF-8 remains an option, then I'd at least want
to be able to document any trade-offs and when people should consider setting the encoding
back to UTF-8.


> Unicode compression
> -------------------
>
>                 Key: LUCENE-1799
>                 URL: https://issues.apache.org/jira/browse/LUCENE-1799
>             Project: Lucene - Java
>          Issue Type: New Feature
>          Components: Store
>    Affects Versions: 2.4.1
>            Reporter: DM Smith
>            Priority: Minor
>         Attachments: Benchmark.java, Benchmark.java, Benchmark.java, LUCENE-1779.patch,
LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch,
LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch,
LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799_big.patch
>
>
> In lucene-1793, there is the off-topic suggestion to provide compression of Unicode data.
The motivation was a custom encoding in a Russian analyzer. The original supposition was that
it provided a more compact index.
> This led to the comment that a different or compressed encoding would be a generally
useful feature. 
> BOCU-1 was suggested as a possibility. This is a patented algorithm by IBM with an implementation
in ICU. If Lucene provide it's own implementation a freely avIlable, royalty-free license
would need to be obtained.
> SCSU is another Unicode compression algorithm that could be used. 
> An advantage of these methods is that they work on the whole of Unicode. If that is not
needed an encoding such as iso8859-1 (or whatever covers the input) could be used.    

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: dev-help@lucene.apache.org


Mime
View raw message