HP/Raid5/Performance Lessons
Posted in 1997
If you are running or wish to run with Raid on HP here are some things
that we recently ran in to.
Machine Configuration:
HP-UX 9.04, 6 cpu's, 1.6 gig mem, 32 gig raid (MTI), 600+ users some shm
and some tcp, Infx 7.11 uc1, around a 30 gig database
Initially our informix instance was running on an HP-UX 9000/T-500 with
4 cpu's 1.6 gig mem - no raid involved. Our cache buffers were in the
90 % range, load on the cpu was running around 1-2, with just some
sluggishness when users connected, checkpoints were around 5 sec.
The datacenter/hardware got moved out to Colorado. We installed MTI's
raid boxes, 2 Luns, 32gig worth of disk space, Raid 5 and 4 cpu's.
I brought the instance up on it...put rootdbs, phys and logical logs and
data out on the Raid box and turned it on. We have mostly a R/O
environment however we still have a lot of insert/update activity as
well. We were maintaining as much as 17 mil rows in some tables.
Unfortunately the tables that were hit the hardest were very fragmented
due to some previous neglect...30-40 or more extents in some cases.
Welll.. when we turned on the box... Informix basically died a slow
death.. each of the 4 CPU's showed 100% utilization, the load was up to
as much as 14.0, users couldn't get in and we ground down to a
screeching halt. Running explains on queries did not show any
particularly high costs, and onchecks did not pull anything out of the
catalogs or doing rowid and data checks.
Onstat -g sch showed the CPU vp's basically spinning their wheels..
Adding additional CPU vp's or AIO vp's made things worse.
We added 2 more CPU's, moved Rootdbs off to disks outside of the Raid,
striped the logical logs on 2 disks outside of Raid and put Physical
logs on their own chunk outside of Raid..
It got worse.. onstat -g ath showed mostly sm_reads and everyone was
sitting in line waiting for semaphores but couldn't get any. It was
worse than most traffic jams in Denver at rush hour.
Since we really didn't need all those millions of rows of data (which
were fragmented), we ended up renaming those tables and recreated them
empty in a new dbspace (with proper sizing) outside of Raid..across 3
disks.
Turned the box back on and now we have our load down to less than 1.0,
our checkpoints are 2-5 seconds with 600 + users, buffers are caching
very nicely in the 95% range..most of the cleaning activity is occuring
at checkpoint onstat -F (chunk writes) and the world is happy again.
Even though the books say to keep CPU vps = # cpu's we've found that 4
CPU'vps are the best, any more than that increase the load (uptime) for
some strange reason that we cannot explain.
We'd like to recreate these tables back on the Raid box and try it again
in the near future once I use some grecian formula to clear up all the
gray hairs I've just acquired..
We will be testing 7.2 with HP 10.10 on Raid 5 over the next several
weeks with a similar CPU configuration so I'll report back on those
findings for you users who are looking at going to HP and Raid..
One last thing - when we went to 6 cpu's from 4, our ontape archives
went from 5 hours to 3 hrs... DLT tapes.
Anyway, some things to think about when working with Raid.
:)
Ed
edbid @ cobe.com