RE: IDS 10 erratic run times
Posted in 2005
I'm seeing the same problem with my v9.4 update statistics jobs too.
I've had the same eight Update Stat jobs that run in about 3.5 hours, in
parallel, on a 9.21 system for years. On my 9.4 system these same jobs,
running in parallel, take almost 7 hours to run.
I never had the time to look into my problem so I run my jobs single
threaded on my v9.4 database now and they run in less than 3 hours!
Sorry that didn't help you...
Jerry
-----Original Message-----
From: owner-informix-list@iiug.org [mailto:owner-informix-list@iiug.org]
On Behalf Of Neil Truby
Sent: Sunday, September 18, 2005 7:02 PM
To: informix-list@iiug.org
Subject: IDS 10 erratic run times
IDS 10.0FC3 on Solaris 2.9
This is a brand new server which we are preparing to migrate a v9 system
to.
Over the past few weeks I have run multiple dbimports from more-or-less
the same data source - certainly the same schema - but the import time
varies widely.
The fastest I've done it is 9h, the slowest 14. This weekend's took 13h
10m. There's never anyhting else going on on it. In the UPDATE
STATISICS
phase the elapsed time was 3h: during the "fast" run it was just 1h. I
had a look during this period: often onstat -p onstat -m onstat -D would
show periods of perhaps 2 minutes with no activity at all: no bufreads
or physreads, no checkpointing for 20 minutes (the checkpoint interval
is 5 mins), the relevant thread just "sleeping forever".
I haven't changed the disk layout (a 60G database striped across three
disks in a JBOD array with the standard Solaris LVM interlace of 16k,
this stripe mirrored to another 3 disks identically configured attached
to another controller on the same array); I have experimented with
different buffer sizes (perversely one 9h run was with BUFFERS set to
just 100,000 rather than the usual 1,000,000 although there's no
consistent pattern to this now I've cut the LRU MIN and MAX, and I've
had another 9h run with the 2GBytes of BUFFERS since).
It's the fact that all activity just seems to stop for long periods that
puzzles me. If, during such a period, I fire up dbaccess and run a
sysmater query it returns instantly, so I can't see a systemic cause.
The new box is still 1.5 to twice the speed of the old in comparative
testing but this disparity in times is just puzzling me and I don't want
it to come back and bite us after go live.
Any suggestions or inspiration welcome!
thanks
Neil
sending to informix-list