error while converting from en_us to utf8
Posted in 2011
A user got "data is corrupt"/unknown errors running dbimport of an en_US.819 (ISO 8859-1) export into an en_us.utf8 database, having set CLIENT_LOCALE equal to DB_LOCALE (utf8). Respondents explained the cause: bytes in the 0x80-0xFF range are valid 8859-1 but not valid UTF-8 sequences, so the data fails validation. The fix offered is to set CLIENT_LOCALE=en_us.8859-1 with DB_LOCALE=en_us.utf8 so GLS performs the codeset conversion during import (Leffler reported testing this), or to pre-convert the file with a tool such as his sbcs2utf8. The slow procedure import was not addressed.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Data Types & Schema Design, Migration, Import/Export & Data Conversion
I am getting error while importing from en_US.819 to en_us.utf8.
while doing a dbimport i am getting data is corrupt message and unknown
error.
but when i removed the data file it got imported. The tables only contained
char, varchar ,date fields.
is there any specific reason why it should fail like this.
Also the import was very very slow on procedures
Thanks & Regards
Debadatta Mishra
ph -+918015907520
--00248c0eec0c4d138b049f21b87f
On Wed, Mar 23, 2011 at 00:58, Debadatta Mishra <mishra.dd@gmail.com> wrote:
> I am getting error while importing from en_US.819 to en_us.utf8.
> while doing a dbimport i am getting data is corrupt message and unknown
> error.
>
> but when i removed the data file it got imported. The tables only contained
> char, varchar ,date fields.
>
> is there any specific reason why it should fail like this.
>
> Also the import was very very slow on procedures
>
Various questions arise, such as:
* Version and platform?
* Setting of CLIENT_LOCALE?
* What characters appeared in the data file from the range 0x80..0xFF?
--
Jonathan Leffler <jonathan.leffler@gmail.com> #include <disclaimer.h>
Guardian of DBD::Informix - v2008.0513 - http://dbi.perl.org
"Blessed are we who can laugh at ourselves, for we shall never cease to be
amused."
--001517447aa42aec5a049f22b071
i set the same client_locale as db_locale.
version is 11.5
character in data files are mostly ascii charcters
Thanks & Regards
Debadatta Mishra
ph -+918015907520
On Wed, Mar 23, 2011 at 2:37 PM, Jonathan Leffler <
jonathan.leffler@gmail.com> wrote:
> On Wed, Mar 23, 2011 at 00:58, Debadatta Mishra <mishra.dd@gmail.com>
> wrote:
>
> > I am getting error while importing from en_US.819 to en_us.utf8.
> > while doing a dbimport i am getting data is corrupt message and unknown
> > error.
> >
> > but when i removed the data file it got imported. The tables only
> contained
> > char, varchar ,date fields.
> >
> > is there any specific reason why it should fail like this.
> >
> > Also the import was very very slow on procedures
> >
>
> Various questions arise, such as:
> * Version and platform?
> * Setting of CLIENT_LOCALE?
> * What characters appeared in the data file from the range 0x80..0xFF?
>
> --
> Jonathan Leffler <jonathan.leffler@gmail.com> #include <disclaimer.h>
> Guardian of DBD::Informix - v2008.0513 - http://dbi.perl.org
> "Blessed are we who can laugh at ourselves, for we shall never cease to be
> amused."
>
> --001517447aa42aec5a049f22b071
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--20cf300255d61a460d049f245ae1
I'd guess for the import you need CLIENT_LOCALE set as en_US.819 so that the
characters are transcoded as they load...
On 23 March 2011 11:06, Debadatta Mishra <mishra.dd@gmail.com> wrote:
> i set the same client_locale as db_locale.
> version is 11.5
> character in data files are mostly ascii charcters
>
> Thanks & Regards
> Debadatta Mishra
> ph -+918015907520
>
> On Wed, Mar 23, 2011 at 2:37 PM, Jonathan Leffler <
> jonathan.leffler@gmail.com> wrote:
>
> > On Wed, Mar 23, 2011 at 00:58, Debadatta Mishra <mishra.dd@gmail.com>
> > wrote:
> >
> > > I am getting error while importing from en_US.819 to en_us.utf8.
> > > while doing a dbimport i am getting data is corrupt message and unknown
> > > error.
> > >
> > > but when i removed the data file it got imported. The tables only
> > contained
> > > char, varchar ,date fields.
> > >
> > > is there any specific reason why it should fail like this.
> > >
> > > Also the import was very very slow on procedures
> > >
> >
> > Various questions arise, such as:
> > * Version and platform?
> > * Setting of CLIENT_LOCALE?
> > * What characters appeared in the data file from the range 0x80..0xFF?
> >
> > --
> > Jonathan Leffler <jonathan.leffler@gmail.com> #include <disclaimer.h>
> > Guardian of DBD::Informix - v2008.0513 - http://dbi.perl.org
> > "Blessed are we who can laugh at ourselves, for we shall never cease to
> be
> > amused."
> >
> > --001517447aa42aec5a049f22b071
> >
> >
> >
> >
>
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --20cf300255d61a460d049f245ae1
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Nick Lello | Web Architect
o +44 (0) 8433309374 | m +44 (0) 7917 138319
Email: nick.lello at rentrak.com
RENTRAK | www.rentrak.com | NASDAQ: RENT
--0016363b8044cb8483049f257360
"mostly" is probably the problem....
From: "Debadatta Mishra" <mishra.dd@gmail.com>
To: ids@iiug.org
Date: 03/23/2011 06:08 AM
Subject: Re: error while converting from en_us to utf8 [23176]
Sent by: ids-bounces@iiug.org
i set the same client_locale as db_locale.
version is 11.5
character in data files are mostly ascii charcters
Thanks & Regards
Debadatta Mishra
ph -+918015907520
On Wed, Mar 23, 2011 at 2:37 PM, Jonathan Leffler <
jonathan.leffler@gmail.com> wrote:
> On Wed, Mar 23, 2011 at 00:58, Debadatta Mishra <mishra.dd@gmail.com>
> wrote:
>
> > I am getting error while importing from en_US.819 to en_us.utf8.
> > while doing a dbimport i am getting data is corrupt message and unknown
> > error.
> >
> > but when i removed the data file it got imported. The tables only
> contained
> > char, varchar ,date fields.
> >
> > is there any specific reason why it should fail like this.
> >
> > Also the import was very very slow on procedures
> >
>
> Various questions arise, such as:
> * Version and platform?
> * Setting of CLIENT_LOCALE?
> * What characters appeared in the data file from the range 0x80..0xFF?
>
> --
> Jonathan Leffler <jonathan.leffler@gmail.com> #include <disclaimer.h>
> Guardian of DBD::Informix - v2008.0513 - http://dbi.perl.org
> "Blessed are we who can laugh at ourselves, for we shall never cease to
be
> amused."
>
> --001517447aa42aec5a049f22b071
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--20cf300255d61a460d049f245ae1
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
On Wed, Mar 23, 2011 at 04:06, Debadatta Mishra <mishra.dd@gmail.com> wrote:
> I set the same client_locale as db_locale.
> version is 11.5
> character in data files are mostly ascii charcters
>
As others have noted, your problem then comes from the implied claim that
all the 8859-1 data is in fact encoded as valid UTF8. For characters in the
range 0x00..0x7F (or U+0000..U+007F), there is no problem. For characters
from the 8859-1 range 0x80..0xFF, the value is not a simple valid UTF8
encoding and you will end up with errors. UTF8 does not allow character
codes 0xC0, 0xC1, 0xF5..0xFF at all, for example, and requires certain
sequences of characters (a byte in the range 0xC2..0xDF must be followed by
one byte in the range 0x80..0xBF, for example) which random 8859-1 data is
unlikely to satisfy. (This is fundamentally the point I raised in my
original response.)
You could arrange to map the 8859-1 to UTF8 characters; I have a program
sbcs2utf8 that could be used to convert any single-byte code set (SBCS) to
UTF8 (given an appropriate mapping file), and there are plenty of other
(better) programs to do the same job.
Alternatively, you can do the import with CLIENT_LOCALE=en_us.8859-1 and
DB_LOCALE=en_us.utf8. Then GLS does the codeset conversion automatically.
[Tested on MacOS X 10.6.7 with IDS 11.70.FC1. Database created with local
en_us.8859-1; table contained single VARCHAR(255) field; data with one byte
for each of 0x80..0xFF loaded into table. Database exported; export hacked
to change DB name (from iso8859_1) to utf8. Database imported with
environment as suggested; data selected is demonstrably UTF8 when the
CLIENT_LOCALE is set to en_us.utf8. And the last character, LATIN SMALL
LETTER Y WITH DIAERESIS (ÿ) was truncated because the VARCHAR(255) can only
hold 127 2-byte characters.]
> On Wed, Mar 23, 2011 at 2:37 PM, Jonathan Leffler wrote:
> > On Wed, Mar 23, 2011 at 00:58, Debadatta Mishra wrote:
> > > I am getting error while importing from en_US.819 to en_us.utf8.
> > > while doing a dbimport i am getting data is corrupt message and unknown
> > > error.
> > >
> > > but when i removed the data file it got imported. The tables only
> > > contained char, varchar ,date fields.
> > >
> > > is there any specific reason why it should fail like this.
> > >
> > > Also the import was very very slow on procedures
> >
> > Various questions arise, such as:
> > * Version and platform?
> > * Setting of CLIENT_LOCALE?
> > * What characters appeared in the data file from the range 0x80..0xFF?
>
--
Jonathan Leffler <jonathan.leffler@gmail.com> #include <disclaimer.h>
Guardian of DBD::Informix - v2008.0513 - http://dbi.perl.org
"Blessed are we who can laugh at ourselves, for we shall never cease to be
amused."
--00151747be16a4afd5049f32f1de