Translating with DrWatson… this can take a few seconds the first time.
This is a genuine, complex translation. DrWatson protects commands, error codes, and log output while naturally translating the surrounding text. It’s translated once and saved.
A user needed a multilingual UTF-8 database and found that the shipped Croatian (sh_hr) UTF-8 locale's LC_COLLATE section is byte-identical to en_US, so Croatian sorting is wrong; the same appears true for other locales such as cs_CZ. He asked whether documentation exists for writing/compiling a custom .lc/.lco locale with a proper collation sequence, and how to handle characters outside the target language. Another poster suggested logging it with support, but IBM reportedly classed it as a feature request rather than a bug. No resolution or workaround is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Hello everyone.
Here is the problem: I have a customer who wants to use a multilingual web
app on Informix. Therefore they have to use utf8 encoding so they can have
n(var)chars from different character sets in the same database (latin,
cyrillic and possibly even others).
My main concern is UTF-8 support for Croatian language. It's shipped on
International Language Support CD, all letters go in and out just fine, but
the collation does not work.
Of course it doesn't since it's defined identicaly to en_US collation, which
is not the way it is supposed to be (I checked e01c.lco files both in en_us
and sh_hr directory, LC_COLLATION part of them is identical in every byte!).
So I was thinking of writing a new locale with correct collation sequence.
(Let's for a minute imagine I find someone in IBM to compile my new source
'lc' file into an object 'lco' file). It's boring but otherwise not much of
a trouble since everything else is in place to do it.
Of course, there is a problem: I know how to specify collation sequence for
Croatian letters, but what about all the rest? How do I "ignore" others? Is
it possible to do?
So the real question would be: is there any documentation on this topic out
there? I searched everywhere without success...
Thanks everyone.
you said
"My main concern is UTF-8 support for Croatian language. It's shipped
on
International Language Support CD, all letters go in and out just fine,
but
the collation does not work. "
What file in the ILS is the Croatian UTF-8 support provided through? I
can see Russian, Czech and Slovak but not Croatian. maybe I have an old
ILS?
I'm therefore concerned that Croatian language is not shipped with the
ILS. In that case the sort order being wrong would not be a bug.
Yes, my colleague did.
They say it's not a bug, but a feature request.
"scottishpoet" <dryburghj@yahoo.com> wrote in message
news:1137673311.807644.4580@o13g2000cwo.googlegroups.com...
> havbe you contacted your Informix support provider to log a bug with
> the exisiting collation sequence?
>
It's under sh_hr directory (aka Serbo-Croatian, Croatian territory) in
gls/lc11/sh_hr/e01c.lc(o) file(s).
"scottishpoet" <dryburghj@yahoo.com> wrote in message
news:1137676548.588271.309640@g14g2000cwa.googlegroups.com...
> you said
>
> "My main concern is UTF-8 support for Croatian language. It's shipped
> on
> International Language Support CD, all letters go in and out just fine,
> but
> the collation does not work. "
>
> What file in the ILS is the Croatian UTF-8 support provided through? I
> can see Russian, Czech and Slovak but not Croatian. maybe I have an old
> ILS?
>
> I'm therefore concerned that Croatian language is not shipped with the
> ILS. In that case the sort order being wrong would not be a bug.
>
I've just checked, if you compare e01c.lco files in en_us and (for example)
cs_cz directory you'll find that LC_COLLATE definitions are identical. Same
for sh_hr.
Also there is no collationg definition in corresponding lc files.
The GLS manual says:
---------------------------------------------
GLS locales that use the Unicode code set (UIF-8) support Unicode collation
of NCHAR and NVARCHAR data by the ICU Unicode Collation Algorithm. For more
information about this algorithm, see the Unicode website at
http://www.unicode.org/unicode/reports/tr10.
---------------------------------------------
Documentation for algorithm is quite long, I didn't go through it but for me
it doesn't give the expected results.
"scottishpoet" <dryburghj@yahoo.com> wrote in message
news:1137676548.588271.309640@g14g2000cwa.googlegroups.com...
> you said
>
> "My main concern is UTF-8 support for Croatian language. It's shipped
> on
> International Language Support CD, all letters go in and out just fine,
> but
> the collation does not work. "
>
> What file in the ILS is the Croatian UTF-8 support provided through? I
> can see Russian, Czech and Slovak but not Croatian. maybe I have an old
> ILS?
>
> I'm therefore concerned that Croatian language is not shipped with the
> ILS. In that case the sort order being wrong would not be a bug.
>
Your privacy choices
We use strictly necessary cookies to make this site work. With your
consent we’d also use optional cookies for analytics and marketing. You can accept all,
reject all, or choose. Read our Cookie Policy.