Cannot insert the utf8 data into database
Posted in 2013
User attempted to migrate Big5 Chinese character data to UTF-8 database. Non-standard Big5 characters (not in standard zone) converted to random Unicode values. Informix rejected inserts with "Invalid byte in codeset conversion input" error. Martin Fuerderer (IBM Informix developer) explained that random-generated character codes may not be valid UTF-8 code points, and Informix validates UTF-8 input.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Platform-Specific Issues, Internationalization & Character Sets
I have created the database of DB_LOCALE = en_us.utf8 on the AIX system. My DB is informix dynamic server 10.00.FC5. My job is to convert the Big5 tables of a old database into a new database of UTF8. I used the c# program to convert the Big5 chinese characters into unicode (UTF8) and then inserted into new database. If the Big5 chinese characters are in standard zone, I can convert into corresponding unicode (UTF8), and then can be inserted into new database. If the Big5 chinese characters not in the standard zone, it cannot be converted to corresponding unicode (it is obvious that those characters are created by myself, they of-cause have no corresponding unicode and it returned a random unicode code). When I inserted those converted data, it showed the error message "Invalid byte in codeset conversion input". I am so strange that because : - i have converted the Big5 chinese characters into unicode (UTF8) even though some not have the corresponding unicode. - in my client side, DB_LOCALE = CLIENT_LOCALE = en.us_utf8. It should be no conversion performed. Why ????
Hi, before trying to give a possible explanation, I've to state, that I'm not an expert for this. Most likely, some of the character code that got generated randomly (because the Big5 chinese character cannot be converted) are codes that are not valid UTF-8 code points. Not ever code that you can construct is a valid UTF-8 code point. And the Informix Server checks the input to make sure that there are only valid UTF-8 code points. There's a description of UTF-8 in Wikipedia. This is a good starting point for understanding UTF-8. Regards, Martin -- Martin Fuerderer IBM Informix Development Munich, Germany Information Management Read about the Informix Warehouse Accelerator: http://tinyurl.com/the-iwa-blog IBM Deutschland Research & Development GmbH Chairman of the Supervisory Board: Martina Koederitz Board of Management: Dirk Wittkopp Corporate Seat: Boeblingen, Germany Reg.-Gericht: Amtsgericht Stuttgart, HRB 243294 From: "VINCENT YEUNG" <chingsze@hkcts.com> To: ids@iiug.org, Date: 09/18/2013 14:19 Subject: Cannot insert the utf8 data into database [31443] Sent by: ids-bounces@iiug.org I have created the database of DB_LOCALE = en_us.utf8 on the AIX system. My DB is informix dynamic server 10.00.FC5. My job is to convert the Big5 tables of a old database into a new database of UTF8. I used the c# program to convert the Big5 chinese characters into unicode (UTF8) and then inserted into new database. If the Big5 chinese characters are in standard zone, I can convert into corresponding unicode (UTF8), and then can be inserted into new database. If the Big5 chinese characters not in the standard zone, it cannot be converted to corresponding unicode (it is obvious that those characters are created by myself, they of-cause have no corresponding unicode and it returned a random unicode code). When I inserted those converted data, it showed the error message "Invalid byte in codeset conversion input". I am so strange that because : - i have converted the Big5 chinese characters into unicode (UTF8) even though some not have the corresponding unicode. - in my client side, DB_LOCALE = CLIENT_LOCALE = en.us_utf8. It should be no conversion performed. Why ???? ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hi, True, the original string contains the non-standard big5 character that is not supposed to have been mapped to unicode because the big5 code added by myself. But before inserting into the database, I converted the string into utf8 (the non-standard big5 character is converted to a random unicode). The converted string contains all valid unicode characters. When I inserted the converted string into database, it showed the error message. I am doubt that why the informix performed the conversion since: - - the converted string contains all valid unicode characters (utf8). - DB_LOCALE and CLIENT_LOCALE are all set to en_us.utf8. I think it may involve the environment setting (PC client), or the version of unicode my version of informix supports? Anyway, thanks for your response.
I think the informix dynamic server should have its own acceptable range of unicode (UTF8). If incoming data contains the data out of its confind range, it shows the "error message" and terminate the processing. That is only my assumption...no any supports from the Informix offical web site or materials..
Hello Martin, You are right. I have inserted the invalid utf8 data format into the database. It is my fault, not Informix.
Hi Martin, According to the definition of UTF8, they are all unicode. My problem is not fixed. For some unicodes, I can insert into the Database. But for some, it can not with the error message conversion error. I don't know why ? I don't think the Informix can distinguish which unicode can accept and other not. I will re-read the manuals and wanna fixing the problem. Best Regards, Vincent YEUNG 7/10-2013