UTF-8 best practice?
Posted in 2017
Topics: Data Types & Schema Design, Migration, Import/Export & Data Conversion, Internationalization & Character Sets
Hi Informixers,
all our Informix databases, currently with IDS 12.10FC2 under SLES11, have
been running with the ISO-8859-1 (or -15) codeset, which is quite sufficient
for our purposes since we do not regularly use any foreign characters.
Occasionally such characters (like in Eastern European names) could be
circumscribed if necessary.
We are currently setting up a new server for IDS 12.10.FC8W1 under SLES12. Is
it considered "best practice" to use UTF-8 for databases nowadays? Any pro's
and con's?
If we decide to switch to UTF-8, there are a couple more questions:
1) Should all present CHAR/VARCHAR columns be converted to NCHAR/NVARCHAR?
2) Is there a rule of thumb (or even an accurate calculation) of the extra
column length needed for the added multibyte characters? Our language is
German, so mostly the Umlaut characters like 'ä','ß' etc. are relevant.
3) I assume that we cannot just use dbexport/dbimport for moving the data
between servers, but will have to do codeset conversion from ISO-8859-1 to
UTF-8 on the exported data files before importing them. Are there better tools
for doing this, or pitfalls to consider?
Thanks for your support,
Richard
AFAIR from what Madison said all ER uses UTF8, so using UTF8 across the
board means the skip a layer of translation
Cheers
Paul
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
RICHARD SPITZ
Sent: Tuesday, April 18, 2017 8:39 AM
To: ids@iiug.org
Subject: UTF-8 best practice? [38959]
Hi Informixers,
all our Informix databases, currently with IDS 12.10FC2 under SLES11, have
been running with the ISO-8859-1 (or -15) codeset, which is quite sufficient
for our purposes since we do not regularly use any foreign characters.
Occasionally such characters (like in Eastern European names) could be
circumscribed if necessary.
We are currently setting up a new server for IDS 12.10.FC8W1 under SLES12.
Is
it considered "best practice" to use UTF-8 for databases nowadays? Any pro's
and con's?
If we decide to switch to UTF-8, there are a couple more questions:
1) Should all present CHAR/VARCHAR columns be converted to NCHAR/NVARCHAR?
2) Is there a rule of thumb (or even an accurate calculation) of the extra
column length needed for the added multibyte characters? Our language is
German, so mostly the Umlaut characters like 'ä','ß' etc. are relevant.
3) I assume that we cannot just use dbexport/dbimport for moving the data
between servers, but will have to do codeset conversion from ISO-8859-1 to
UTF-8 on the exported data files before importing them. Are there better
tools
for doing this, or pitfalls to consider?
Thanks for your support,
Richard
****************************************************************************
***
Forum Note: Use "Reply" to post a response in the discussion forum.
Related threads
- Conversion to differeent characters sets
- Problem in changing locale via dbexport/dbimport
- RE: openlink error "Unable to load locale categories"