Speed discrepancy between prod and a lab
Posted in 2012
Jared Heath reported that a lab server (restored from a production image, same hardware, disks, onconfig and an L0 restore) ran IDS 11.70.FC2 batch jobs 25-30% faster than production on HP-UX. Suggestions included checking RAID/disk configuration and throughput, extent layout, index levels and stats, RAID cache battery health, kernel parameters, raw vs. block/cooked devices, and CPU affinity; Madison Pruet noted HDR replication can drag on the primary and suggested more offline recovery threads. Heath verified chunks, index levels, drives and battery as matching. No resolution is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration, Platform-Specific Issues
I've got a performance issue between my lab and production that I'm at a loss to explain at this point. - Exact same hardware (server, cpu and memory identical) - Exact same disk array, disk speeds, raid card & cache - The lab server is a restored image of the production server - The IDS instances are using the same onconfig - The lab instance is a L0 restore of the production one. - The server activity is identical...prod has no server activity beyond the same jobs, and also is user restricted. What I'm seeing is my test setup is out-performing my production by 25-30% on IDS batch jobs. This is 11.70.FC2 on 11.23 HP-UX (rp3440 server). My expectation of the L0 restore is that my production should actually be faster than my lab due to built-up buffers/caching but instead, its far slower than my L0 restored instance. Any thoughts? Does the L0 restore clean out something I'm missing that can account for this speed change. Note that I've got 2 production environments, and the other one (wildly different in usage and what processes on it) does not have this discrepancy in speed...in fact, it runs within 1% of my production run-times. This makes me think there is either something odd with this specific instance, or that I've got a hardware issue on my production system that is going unseen.
Disk? Are you using the same disk configuration (RAID) on prod and lab? j. On Jan 17, 2012, at 10:35 AM, JARED HEATH wrote: > I've got a performance issue between my lab and production that I'm at = a loss=20 > to explain at this point.=20 >=20 > - Exact same hardware (server, cpu and memory identical)=20 > - Exact same disk array, disk speeds, raid card & cache=20 > - The lab server is a restored image of the production server=20 > - The IDS instances are using the same onconfig=20 > - The lab instance is a L0 restore of the production one.=20 > - The server activity is identical...prod has no server activity = beyond the=20 > same jobs, and also is user restricted.=20 >=20 > What I'm seeing is my test setup is out-performing my production by = 25-30% on=20 > IDS batch jobs. This is 11.70.FC2 on 11.23 HP-UX (rp3440 server).=20 >=20 > My expectation of the L0 restore is that my production should actually = be=20 > faster than my lab due to built-up buffers/caching but instead, its = far slower=20 > than my L0 restored instance.=20 >=20 > Any thoughts? Does the L0 restore clean out something I'm missing that = can=20 > account for this speed change.=20 >=20 > Note that I've got 2 production environments, and the other one = (wildly=20 > different in usage and what processes on it) does not have this = discrepancy in=20 > speed...in fact, it runs within 1% of my production run-times. This = makes me=20 > think there is either something odd with this specific instance, or = that I've=20 > got a hardware issue on my production system that is going unseen.=20 >=20 >=20 > = **************************************************************************= *****=20 > Forum Note: Use "Reply" to post a response in the discussion forum.=20= >=20
completely identical. Same # of disks, same speed, same raid enclosure, same raid level, same card and cache size.
I've been over-looking the HDR Remote Secondary that prod has, so I should add that we have one. I don't think it can account for a 25% speed difference. The other production environment has the same type of HDR, and we don't see the speed discrepancy with it.
And it's an L0 restore, so extent layout and stats should be identical, =
but I would check them anyway. oncheck -pe and query the btree levels =
in sysindexes, and/or the USTLOWTS in systables.
And just to eliminate disk from the equation, how fast can you =
read/write 1GB to each of the disks in question? (bonnie =
http://www.textuality.com/bonnie/ is a fun tool to test that).
j.
On Jan 17, 2012, at 10:55 AM, JARED HEATH wrote:
> completely identical.=20
>=20
> Same # of disks, same speed, same raid enclosure, same raid level, =
same card=20
> and cache size.=20
>=20
>=20
> =
**************************************************************************=
*****=20
> Forum Note: Use "Reply" to post a response in the discussion forum.=20=
>=20
I'll check the disks, but the speed checking tool you mention I don't believe I could attempt. The disks that the DB resides on are Raw. I could get a perf test from the lab, but not production. I'm also the system admin though, and the disks are perfectly identical...they are all 15k HP 300gb disks. Unless somebody slipped a 10k drive into the production raid...hmmm...
verified the chunks/idx levels are identical. I also verified all the drives are 15k HP drives, same product ID. Some of them have different firmware levels but firmware shouldn't explain the discrepancy (at least, it better not)
Jared, It is possible that HDR can put a bit of a drag on the primary system. This issue has always been there, but is a bit more pronounced with new= er systems because the difference between CPU speeds and IO speed has grow= n. We recently (11.70xC4) changed the way that the secondary loads pages i= nto memory as it is applying log records and have seen a significant reduct= ion in the average backlog of operations to be performed. The main issue i= s that the secondary apply threads were also having to perform the physic= al IO to load the page into memory. We are now using preload techniques t= o load the page. There are still some things that can be done to help with this. The biggest is to increase the number of offline recovery threads. M.P. = From: "JARED HEATH" <jared.heath@sbcglobal.net> = = To: ids@iiug.org = = Date: 01/17/2012 11:27 AM = = Subject: Re: Speed discrepancy between prod and a lab [25951] = = Sent by: ids-bounces@iiug.org = = verified the chunks/idx levels are identical. I also verified all the drives are 15k HP drives, same product ID. Some= of them have different firmware levels but firmware shouldn't explain the discrepancy (at least, it better not) ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Have you checked the battery on the cache on the disk controller on prod? On 17/01/2012 15:35, JARED HEATH wrote: > I've got a performance issue between my lab and production that I'm at a loss > to explain at this point. > > - Exact same hardware (server, cpu and memory identical) > - Exact same disk array, disk speeds, raid card& cache > - The lab server is a restored image of the production server > - The IDS instances are using the same onconfig > - The lab instance is a L0 restore of the production one. > - The server activity is identical...prod has no server activity beyond the > same jobs, and also is user restricted. > > What I'm seeing is my test setup is out-performing my production by 25-30% on > IDS batch jobs. This is 11.70.FC2 on 11.23 HP-UX (rp3440 server). > > My expectation of the L0 restore is that my production should actually be > faster than my lab due to built-up buffers/caching but instead, its far slower > than my L0 restored instance. > > Any thoughts? Does the L0 restore clean out something I'm missing that can > account for this speed change. > > Note that I've got 2 production environments, and the other one (wildly > different in usage and what processes on it) does not have this discrepancy in > speed...in fact, it runs within 1% of my production run-times. This makes me > think there is either something odd with this specific instance, or that I've > got a hardware issue on my production system that is going unseen. > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. >
- OS/Kernel parameters? - You mention RAW devices. Any chance that one of them is point to block devices instead of character devices? - Any possibility that you have dynamic CPU allocation? Regards On Tue, Jan 17, 2012 at 3:35 PM, JARED HEATH <jared.heath@sbcglobal.net>wrote: > I've got a performance issue between my lab and production that I'm at a > loss > to explain at this point. > > - Exact same hardware (server, cpu and memory identical) > - Exact same disk array, disk speeds, raid card & cache > - The lab server is a restored image of the production server > - The IDS instances are using the same onconfig > - The lab instance is a L0 restore of the production one. > - The server activity is identical...prod has no server activity beyond the > same jobs, and also is user restricted. > > What I'm seeing is my test setup is out-performing my production by 25-30% > on > IDS batch jobs. This is 11.70.FC2 on 11.23 HP-UX (rp3440 server). > > My expectation of the L0 restore is that my production should actually be > faster than my lab due to built-up buffers/caching but instead, its far > slower > than my L0 restored instance. > > Any thoughts? Does the L0 restore clean out something I'm missing that can > account for this speed change. > > Note that I've got 2 production environments, and the other one (wildly > different in usage and what processes on it) does not have this > discrepancy in > speed...in fact, it runs within 1% of my production run-times. This makes > me > think there is either something odd with this specific instance, or that > I've > got a hardware issue on my production system that is going unseen. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Fernando Nunes Portugal http://informix-technology.blogspot.com My email works... but I don't check it frequently... --00248c6a6a42339df004b6bff435
yeah, batteries are showing as "ok". We've seen that one before (sadly)
The OS/Kernel parms should be identical. The lab server is a restored image of
the prod server (just done again yesterday)
I'll cross-reference the devices, but I don't think its possible the way my L0
works it would croak on my lab's full RAW setup if one of the prod devices
wasn't also RAW, right?
This is a single CPU, dual-core server with only one CPU allocated to IDS...so
no CPU stuff should be happening. It doesn't have affinity set, so I guess its
possible it could have some context-switching if oninit swaps from one CPU to
another. I've never tested to see if affinity really helps on these servers.
The raw devices is just a shot in the dark... But L0 takes data from one
"file" and puts it into another "file".
If these files are cooked, direct raw, links to cooked or links to raw
doesn't really make a difference for the backup/restore Informix process.
Regards.
On Tue, Jan 17, 2012 at 9:36 PM, JARED HEATH <jared.heath@sbcglobal.net>wrote:
> The OS/Kernel parms should be identical. The lab server is a restored
> image of
> the prod server (just done again yesterday)
>
> I'll cross-reference the devices, but I don't think its possible the way
> my L0
> works it would croak on my lab's full RAW setup if one of the prod devices
> wasn't also RAW, right?
>
> This is a single CPU, dual-core server with only one CPU allocated to
> IDS...so
> no CPU stuff should be happening. It doesn't have affinity set, so I guess
> its
> possible it could have some context-switching if oninit swaps from one CPU
> to
> another. I've never tested to see if affinity really helps on these
> servers.
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--00235447044cfeecc204b6c042fa