Re: Overall unstableness of Informix
Posted in 1996
In article <322E1112.3F8C@transre.com>, "Juan R. Guzman"
<jguzman@transre.com> writes
>Michael Silver wrote:
>>
>> We have had an Informix (7.12.UC2) database running on an HP for over a
>> year. We have been utterly disappointed with it. It crashes about
>> once-a-week. We run a lot of batch work against it using New Era for
>> Motif, which is responsible for many of the crashes. How a programming
>> language can crash the database engine is beyond me. Often times a
>> small batch job will hit the engine so hard, that everyone else in the
>> database times out.
Sounds like the batch job is poorly coded - any queries/updates
should be using indicies. You should be able to get a trace of the
SQL executed and query plans used from NewEra - 4Gl/dbaccess uses
"SET EXPAIN ON". There should be something similar for NewEra??
ALso start a batch job running then log in to the HP box all the
informix enviroment variables set and run onstat -u several times.
Does your batch job show hunderds or thousands of read/write
operations per second. If so it is probably not using an index so
do onstat -g sql <session-id> repeatedly to find the offending piece
of sql.
> We can't use PDQ, because it crashes the database,
>> or causes strange things to happen.
Sounds like either
a) PDQ has not been correctly configured, how you configure it
depends upon the maximum number of simultanoeus PDQ queries will
be running
b) the UNIX kernel has not been correctly configured.
c) the Informix parameter OPTCOMPIND is 2 rather than 0.
>We have a dual processor HP Model
>> K. Our database fits into RAM (we have 512MB). We have dual port fast
>> and wide SCSI, with about 6 GB storage. A relatively small database, on
>> a fast machine, but we get very poor performance. We have had Informix
>> out here at least every 6 months, and nothing gets fixed, which makes
>> the $5000 bill hard to swallow. Even HP doesn't understand he
>> performance
>>
Let me guess - you called for Informix Online Performance Tuning
experts. Online is probably correctly configured but the application
(i.e. batch jobs) / database configuration (database indicies) need
to the performance tuned. This would be done by the people who wrote
the application not Informix Online Tuners or HP people.
Remember Informix is a LARGE company - people at Informix specialise
in either supporting
a) Front end tools i.e. Dbaccess/Isql and either 4GL/NewEra
b) Online configuration and Tuning.
Anybody who programs either 4Gl/NewEra should know about Informix
Database configuration and Tuning (i.e. creation of appropriate
indicies). Sound like the problem is the application/database
indicies and whoever wrote the application is crap.
Sorry to them but when you write an application you should know
how to write it such that it runs quickly.
>> Is this typical? Does everyone have these problems?
Nope I've configured OnLine 7.10.UC1 systems running under HP-UX
9.04 with 28Gb of RAID disk (about 1 Gb of data to start with
growing to 28Gb within 3 years). ONE CPU and 80 users. Word documents
were stored as Blobs within the database and we had no performance/
stability problems. This was in use 24 hours a day 7 days a week
and last I heard it had been running continuously for ~3 months.
>Perhaps we are on the wrong platform?
Nope in fact HP is one of the first platforms that Informix port to
- stick with it. The only thing is that if you run version 10 do not
run 10.01 as a UNIX kernel bug with shared memory hadnling exists.
Instead run the latest patched 10.10 version.
>Is Informix just this bad?
Only if incorrectly configured.
> Should we switch to Oracle or another database?
Nope Oracle is not truely multi-threaded as so gives worse
performance under heavy load. Also most Oracle things e.g.
there multithreaded server (which is not truely multi-threaded)
COSTS EXTRA.
>>We went to a user group meeting and laughed
>> when Informix told us about the datablades. We can't even get the
>> normal version to work. We feel we can get better performance from
>> an
>> NT quad Pentium Pro, running SQL server.
Possibly - 4 CPUs run faster than 2.
>>This is not practical for as
>> large as we expect to grow, but the fact that they can honestly be
>> compared is scary. What to do.
>
E-mail me for more info. - I'm quite willing to help you sort out
your problems.
>These are some pointers I can think of real quick. I'm sure the Informix
>people covered most if not all of them but I'll take a shot just in
>case.
>
>I have On-line 7.11UC1 so somethings I'm about to write might be
>different but I'm sure most things are the same.
>
>1.- Make sure that MULTIPROCESSOR is set to 0. Informix recommends
>setting it to 1 only if you have at least 4 processors.
>
Correct
>2.- Try using 2 cpu vps, it helped in my case. If you are only going to
>use 1 cpu vp (NUMCPUVPs) make sure that SINGLE_CPU_VP is not equal to 0.
>
I would not recommend 2 cpu vps wihtout more info.
Run onstat -g rea when the system is under heavy load.
If there are any threads listed then you do not have enough CPU VP's
configured - these threads are ready to run but are awaiting a turn
but one CPU VP can only run one thread at a time.
>3.- Does HP support KAIO? If it does make sure you're using it.
>
Meed to check with Hp on this one. Your Informix release notes in
$INFORMIXDIR/release/ONLINE_7.1 should give tell you if KAIO is not
supported.
>4.- Maybe you've used to much memory for caching? There comes a point
>where setting too many buffers doesn't give you better performance.
>Actually it might make your machine to start swapping (now that I
>mention it, check to see if it's swapping).
>
Correct - run
1. sar -g 3 any pgouts/s on a regular basis indicate not enough
memory for what is running. Possibly too many Online buffers
configured.
>5.- Checkpoints: I believe that the default of 300 is OK. What I did
>change was the LRU_MAX_DIRTY AND LRU_MIN_DIRTY (7 and 5 respectively)
>and now my checkpoints don't take as long. Do an "onstat -F", you
>shouldn't have any foreground writes (if you, do this is something that
>you really have to get rid of). Of chunk writes and LRU writes we want
>LRU writes.
>
Correct.
>6.- SHMADD shouldn't be a small number (I have 8MB, I had 2 when I first
>set up On-line) because of the overhead it creates to use it.
>
Run onstat -g seg count amount of memory in use under heavy load and
set SHMVIRTSIZE to this value.
>7.- I personally don't use PDQ, tried it, it slowed down my response. I
>think (and I might be wrong on this)that PDQ is for servers that have at
>least 4 processors with a lot of RAM and a huge database. Hopefully
>somebody that has set this up and is happy with it, will put there 2
>c