re: TEXT or BYTE
Posted in 1996
I'd like to try to clear up some confusion about the use of TEXT and locales. See comments enclosed in ">>>>>>" characters within the attached email. DAM Principal Tech Writer SCT Tech Pubs >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> TEXT data type provides the same advantages in all locales; it >> allows you to store large character values in a column. >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > as their names were intended to indicate, a TEXT blob supposedly can be > printed, or at least displayed by an appropriate text editor; whereas > the BYTE data type implies virtually nothing about the contents or format > of the blob data that it stores, except the size cannot exceed 2**31 bytes. > > you should wait for an authoritative answer from someone in the Servers > and Connectivity group, but here are my comments from a "Tools" perspective: > >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> The TEXT and BYTE data types are meant for very large values. The >> database server interprets a TEXT value as a "string" of ASCII >> characters. (It does not place any interpretation on a BYTE value; >> the application program is expected to deal with the valid >> appropriately.) So the chief distinction between TEXT and CHAR (and >> VARCHAR) doesn't have any thing to do with which characters they accept. >> The difference is strictly once of the size of data expected. CHAR (and >> VARCHAR) columns have a physical size limit (I think around 255 >> bytes). TEXT has a much larger size limit and is therefore not stored >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > >* Informix doesn't recommend to enter any "non-printable" > > characters excapt tabs, newlines, newpages into TEXT > > columns; only ASCII characters are supported. > > So what is the use of TEXT in an NLS environment, > > where many non-ASCII characters occur? > > the use of "non-printable" ASCII characters within the character string > values of CHAR, TEXT, and VARCHAR data types is not encouraged by Informix, > because certain specific "non-printable" values may interfere with other > Informix features for processing, displaying, or printing those values; > for example, Informix products will generally not read past an end-of-data > character within a string, so any subsequent characters are ignored. > (See page 2-68 of "NewEra Language Reference" for additional examples.) > >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> If you are using a non-default locale (a locale other than U.S. >> English), then in a TEXT column, the database server can correctly >> interpret ASCII characters and any other characters that the code set >> of the locale defines. Non-printable ASCII characters would still have >> the same functionality in a non-default locale because ASCII is a >> subset of ALL other code sets. No non-ASCII code set should >> "redefine" an ASCII non-printable character, though it might contain >> other non-printable characters unique to the code set. So the >> database server does not interpret ASCII non-printables any differently >> when you use a non-default locale. >> If the database locale defines other non-ASCII characters (8-bit or >> multibyte), then the database server can correctly interpret these >> non-ASCII characters (printable or not) in TEXT columns. In addition, >> a client application can perform code-set conversion (if the code >> set of the client locale differs from the code set of the database >> locale) on a TEXT column (but NOT a BYTE column). >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > but for Informix products that are GLS-enabled, such as the 7.2 database > server, or NewEra 2.10 and 2.11, the recommendation to include only ASCII > characters in string values only applies to U.S.English locales (namely, > en_us.8859-1 for UNIX, or Code Page 1252 for Windows). >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> The restriction to "ASCII characters only" would apply to any >> locale whose code set is ASCII. >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > in an NLS (or GLS) > environment, the definition (and processing rules) of "non-printable" > characters within TEXT values depends on the codeset specified in the > DB_LOCALE value for the database server, and on the CLIENT_LOCALE value > for the front-end application. >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> The ability to correctly interpret characters within TEXT values >> does indeed depend on the code set of the database locale. However, >> how you set the database locale depends on the product you are using: >> o GLS (7.2) products use DB_LOCALE and CLIENT_LOCALE as >> described in Tom's paragraph above. >> >> o NLS products use the LANG and LC_* environment variables to >> establish a locale (and code set). Some earlier versions >> of ESQL/C (and therefore probably of NewEra too), use >> DB_LOCALE and CLIENT_LOCALE to establish code-set conversion >> but not database locale. >> >> o ALS products use DB_LOCALE and CLIENT_LOCALE (but the syntax >> is different). >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > but if, for example, a Japanese end-of-file > character occurs in the middle of a "PRINT filename" statement within a > Japanese (= GLS) locale, then any characters in the file that follow that > end-of-file character probably will NOT be printed. (unfortunately, i do > not have a Japanese environment to verify this theoretical example.) >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> My understanding is that it would be extremely unlikely that a >> Japanese code set would have a "Japanese end-of-file character. >> The Japanese code set would contains the ASCII code set and the >> end-of-file character would still be whatever it is in ASCII >> (consult your locale ASCII table). >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> > > because Informix does not document which non-printable characters in > non-U.S.English codesets may cause unexpected behavior in various NLS or > GLS environments, it is left to the user to assess the risks involved in > storing and processing unprintable characters within string values. > (that is, most characters are probably harmless, but some may be toxic) > > >* Why does Informix support TEXT and BYTE if there > > is no NLS-Support for TEXT? > > I don't see any significant difference in those 2 > > datatypes. >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> >> Difference between TEXT and BYTE is how the database server >> interprets the data. Difference between CHAR/VARCHAR and TEXT >> is that maximum size of the character data these data types >> can contain. >>