Weird performance problems IDS on AIX 5.3
Posted in 2011
Topics: Performance & Tuning, Installation, Setup & Upgrades, Migration, Import/Export & Data Conversion, Platform-Specific Issues
We are experiencing an issue where the performance of queries on our IDS 10
has slowed down by over 700%. To give everyone a little background, we
migrated hardware from a P4 with an IBM DS8000 to a P7 with an Compellent
using a redundant VIOS setup. We have separate teams that manage each layer of
the solution stack. The Informix team used the IBM TSM backup and restore
method(level 0) to move their IDS 10 and their database files to the new
system. To isolate the issue, I ran ndisk to stress test the storage subsystem
on the AIX layer where the IDS resides. There were no issues and throughput is
above and beyond what we expected.
My question to this group is -
Should we have done a clean install of IDS 10 first on the new hardware and
used the traditional onunload and onload to restore the databases and then
tuned IDS or is it an acceptable practice to simple do a block-level backup
and restore of the IDS install from the old hardware to the new hardware
configuration? Version levels of IDS and the OS were preserved across the
migration path.
Symptoms include -
a. Running out of space on tempspace on heavy batch queries involving sorting
b. Intermittent periods of slow query response times
c. No paging observed
vCPU pool is hardly stressed during peak hours. Figuratively speaking, the new
hardware is a V-12 engine but the ride experience is like only 4 cylinders are
firing when you put your foot on the gas pedal.
We do have PMRs opened. I am only trying to validate the migration path I
thought we should have used when we went from the old hardware to the new
hardware.
Thanks in advance!
In short, wheneever possible I prefer to use migrations to clean up things,
but the fact that you didn't does explain any issues.
Although it's possible I think that some of the symptoms cannot be explained
just by the hardware change.
I would re-verify that nothing else changed... Exhaustion of temporary
dbspace can be caused by wrong query plans.... which can be caused by wrong
statistics...
Sometimes the statistics get old... have you considered that? For example,
if you had tables with statistics calculated when they were very small and
suddenly (maybe due to your tests) they have grown, you could see things
like this (like hash tables made on the wrong table)
If you have a PMR open I'd request help to find out who/which session is
consuming space (there's a little poison here, if anybody from R&D is
reading :) )
Regards.
On Sat, Aug 27, 2011 at 2:48 AM, PRAVEEN ARUMBAKKAM <prav@ymail.com> wrote:
> We are experiencing an issue where the performance of queries on our IDS 10
> has slowed down by over 700%. To give everyone a little background, we
> migrated hardware from a P4 with an IBM DS8000 to a P7 with an Compellent
> using a redundant VIOS setup. We have separate teams that manage each layer
> of
> the solution stack. The Informix team used the IBM TSM backup and restore
> method(level 0) to move their IDS 10 and their database files to the new
> system. To isolate the issue, I ran ndisk to stress test the storage
> subsystem
> on the AIX layer where the IDS resides. There were no issues and throughput
> is
> above and beyond what we expected.
>
> My question to this group is -
>
> Should we have done a clean install of IDS 10 first on the new hardware and
> used the traditional onunload and onload to restore the databases and then
> tuned IDS or is it an acceptable practice to simple do a block-level backup
> and restore of the IDS install from the old hardware to the new hardware
> configuration? Version levels of IDS and the OS were preserved across the
> migration path.
>
> Symptoms include -
>
> a. Running out of space on tempspace on heavy batch queries involving
> sorting
> b. Intermittent periods of slow query response times
> c. No paging observed
>
> vCPU pool is hardly stressed during peak hours. Figuratively speaking, the
> new
> hardware is a V-12 engine but the ride experience is like only 4 cylinders
> are
> firing when you put your foot on the gas pedal.
>
> We do have PMRs opened. I am only trying to validate the migration path I
> thought we should have used when we went from the old hardware to the new
> hardware.
>
> Thanks in advance!
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--0022159749761dea8b04ab7ffe73
Praveen,
Let's start with the Compellent Storage. As a the Lead Database
Administrator that supervisor of the UNIX/SAN administrators and having
changed over to Compellent storage i have learned much. This system by
nature is designed to be a tiered storage sub-system. The question I would
ask your SAN admins is where is the data for your Informix LUNS. They
should provide you some number of how much of your data is in each RAID
level at each Tier (and don't take NO for an answer they can get it, let me
know if you need me to provide the steps for them to follow). This is
important because anything new will be in the top level(s) of storage by
default and as things age and aren't used they will sink to the lowest
tier. As data moves your data will move from Fast disk RAID 10 (SSD, SAS
15K, or Fiber) and work it's way to Slow disk RAID 5 or RAID 6 (SAS 7K, or
SATA).
So that in mind let me explain what will happen as it has happened
with us. We have a large Archive system which is rarely accessed. Since
this system was moved to the fully tiered storage the update stats process
went from running in 8 hours for the whole system to just about 24 hours
(which if we split the DB Spaces we would increase the parallelism) Which is
coming. So this system only has about 1% of if's data in the top tier
storage so very little comes from the fast storage.
Any tool used for testing that writes data and reads it back will find
the performance to be great.
We also have a production database which is fully tiered, we found many
queries still included very old rarely used data. As we have excluded and
phased out the data from reports they have returned to a much better
performance. To maintain the backup (2 hours to 4 hours) and update stats
(3 hours to 10 hours) performance we have made sure that all tables had been
spit into separate DB spaces and increased the number which could be backed
up in parallel (Veritas Administrators).
What i would suggest is that you ask if there is any way you could test
with the whole database in Tier 1 storage. This wouldn't require much work
for your admins, ask for your data to be put in the "High Priority Storage"
profile for this test and check that all data has moved before testing. It
may take hours or days to move all of your data up to this level of storage
(depending on how much). If your performance returns to the expected
performance you have your answer. If not you need to talk to your UNIX
admins.
Another Questions... Where is the Backup Server? On the Compellent
SAN? When are replay's created and is Tiering Run? Ask to see the I/O
Performance report at the time the backup is running. If they have
Enterprise Manager Configured they have the data. Depending on your front
and back end connections your SAN could be maxed. We have also seen this,
with some configuration we made timing better.
As for the UNIX system. You moved from a P4 with dedicated resources to
a P7 with VIO. Not knowing what resources your system had prior and what is
being allocated you maybe resource starved. The questions to ask. What is
the real CPU being assigned to your system, is it dedicated or is it from a
shared Pool. How many Virtual CPU are being assigned if it is Shared? Is
SMT turned on and to what level compared to the number of Core? How much
memory is assigned and is it dedicated. Are your systems using workload
manager and how is it configured. There have been posts on other groups
for AIX and pSeries which have pointed out that a P7 may not provide a large
performance gain compared to past processors for some applications because
of the type of threading used. I don't remember all of the details but
there are hints to make things faster. I would suggest asking the AIX or
Hardware team be on the phone with IBM to learn more about slowness since
migrating to the P7.
As for Informix. If you haven't tuned your onconfig you could have
something extremely wrong for the assigned resources but I would look at the
above first.
Good luck,
Eric B. Rowell
On Fri, Aug 26, 2011 at 9:48 PM, PRAVEEN ARUMBAKKAM <prav@ymail.com> wrote:
> We are experiencing an issue where the performance of queries on our IDS 10
> has slowed down by over 700%. To give everyone a little background, we
> migrated hardware from a P4 with an IBM DS8000 to a P7 with an Compellent
> using a redundant VIOS setup. We have separate teams that manage each layer
> of
> the solution stack. The Informix team used the IBM TSM backup and restore
> method(level 0) to move their IDS 10 and their database files to the new
> system. To isolate the issue, I ran ndisk to stress test the storage
> subsystem
> on the AIX layer where the IDS resides. There were no issues and throughput
> is
> above and beyond what we expected.
>
> My question to this group is -
>
> Should we have done a clean install of IDS 10 first on the new hardware and
> used the traditional onunload and onload to restore the databases and then
> tuned IDS or is it an acceptable practice to simple do a block-level backup
> and restore of the IDS install from the old hardware to the new hardware
> configuration? Version levels of IDS and the OS were preserved across the
> migration path.
>
> Symptoms include -
>
> a. Running out of space on tempspace on heavy batch queries involving
> sorting
> b. Intermittent periods of slow query response times
> c. No paging observed
>
> vCPU pool is hardly stressed during peak hours. Figuratively speaking, the
> new
> hardware is a V-12 engine but the ride experience is like only 4 cylinders
> are
> firing when you put your foot on the gas pedal.
>
> We do have PMRs opened. I am only trying to validate the migration path I
> thought we should have used when we went from the old hardware to the new
> hardware.
>
> Thanks in advance!
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Eric B. Rowell
--0016364eeb6effaa1c04ab9c6f4e
Praveen, Given that magnitude of performance degradation, it almost certianly sounds as if the P7 server you migrated onto has either kernel auditing enabled and/or some level of power saving mode turned on which are HUGE performance killers. Mark
In order to tune AIX on the P7s to run Informix efficiently at my client we had to do the following: - Since we were using COOKED filesystem files in order to take advantage of DIRECT_IO and AIX's CONCURRENT_IO features, we had to remount the chunk filesystems with the NOACCESS flag set to reduce contention for the file inodes (this keeps the chunk's access time from being updated every time Informix writes to a chunk). - Increase the AIX IO queue depth from 8 to 32 or 64 - Disabled the AIX feature known as processor folding. Processor folding disables the SMT threads in the processors if CPU utilization falls below 25% and reenbles the SMT threads in stages if utilization increases. The IBM AIX technician working with us felt that the overhead of constantly ramping the SMT threads down and back up may be hurting performance. With these three changes performance on the P7 machine exceeded the performance experienced on the client's P5 systems. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Mon, Aug 29, 2011 at 9:29 AM, MARK JALKIEWICZ < mark.jalkiewicz@verizon.net> wrote: > Praveen, > > Given that magnitude of performance degradation, it almost certianly sounds > as > if the P7 server you migrated onto has either kernel auditing enabled > and/or > some level of power saving mode turned on which are HUGE performance > killers. > > Mark > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --90e6ba1efe866dfd9a04aba516ed