RE: Long Checkpoints with HDR and 11.5.
Posted in 2010
A shop on IDS 11.50.FC3/FC5 (AIX 6.1) saw extreme slowness and very long checkpoints whenever HDR was active; disabling HDR cured it, and upgrading FC3->FC5 didn't. Suggestions: Madison Pruet blamed primary/secondary imbalance (secondary had only 4 CPUVPs vs 6, OFF_RECVRY_THREADS of 10 too low — use a near-prime value ~3x CPUVPs, avoid Online-5 style indexes). Jacques Renaut pointed to buffer-pool flushing on the secondary at checkpoint holding up the primary (APAR IC60754). Resolution reported: IBM said FC5 was only a partial fix, with the full fix in 11.50.xC6 via logical log staging, which must be enabled with LOG_STAGING_DIR and LOG_INDEX_BUILDS. Actual results after upgrading weren't posted.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Installation, Setup & Upgrades, Logging & Checkpoints, Platform-Specific Issues
We are receiving extreme slowness and long checkpoints on our 11.50.FC3
and now 11.50.FC5 versions on AIX 6.1, while running HDR.
Everything comes back to life when we cut HDR, but began again once we
upgraded to FC5.
Has anyone else had this issue and what was the resolution?
IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to
be the answer.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Peter_Logan@spartanstores.com
Sent: Monday, March 15, 2010 11:53 AM
To: ids@iiug.org
Subject: Re: Strange sql behavior [19335]
Yes it does ... and Art's explaination straightened me out ...
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"Fernando Nunes" <domusonline@gmail.com>
To:
ids@iiug.org
Date:
03/15/2010 12:51 PM
Subject:
Re: Strange sql behavior [19333]
Sent by:
ids-bounces@iiug.org
s_h_w_status or s_hw_status?
Does the column exist on table ps_s_payperiod? If it does it is correct
behavior (You have my sympathy if you find this odd...)
regards
On Mon, Mar 15, 2010 at 3:39 PM, Peter_Logan@spartanstores.com <
Peter_Logan@spartanstores.com> wrote:
> Hello,
>
> I'm running ids 11.50.fc4 on Aix 5.3 ... the following query is giving
> strange behavior ... Any advice is appreciated...
>
> SELECT count(*) FROM ps_S_PAYPERIOD_STS WHERE S_HW_STATUS NOT IN
(SELECT
> S_HW_STATUS FROM ps_S_UL_EE_STS_MVW) ;
>
> This returns the count of 0...
>
> SELECT S_HW_STATUS FROM ps_S_UL_EE_STS_MVW ;> This returns a -217, column not found, s_h_w_status ...
>
> The -217 is correct since this column doesn't exist in the view.
>
> The question is why isn't the first sql also returning the error. It's
> like it doesn't care ...
>
> Please advise...
>
> Peter Logan
> Senior Database Administrator
> Phone: 616/878-8309
>
>
>
>
************************************************************************
*******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--0016e6db2b159803c00481d9aded
************************************************************************
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
************************************************************************
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
It sounds like there might be an imbalance between the primary and the
secondary. A couple of questions...
1) How many physical CPUs on the primary/secondary?
2) How many CPUVPS on the primary/secondary?
3) What is OFF_RECVRY_THREADS set to?
4) What is the output of onstat -g cpu on the secondary?
M.P.
=
From: "Knox, Ernest" <Ernest.Knox@searshc.com> =
=
To: ids@iiug.org =
=
Date: 03/15/2010 12:02 PM =
=
Subject: RE: Long Checkpoints with HDR and 11.5. [19336] =
=
Sent by: ids-bounces@iiug.org =
=
We are receiving extreme slowness and long checkpoints on our 11.50.FC3=
and now 11.50.FC5 versions on AIX 6.1, while running HDR.
Everything comes back to life when we cut HDR, but began again once we
upgraded to FC5.
Has anyone else had this issue and what was the resolution?
IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to=
be the answer.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS=
!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Peter_Logan@spartanstores.com
Sent: Monday, March 15, 2010 11:53 AM
To: ids@iiug.org
Subject: Re: Strange sql behavior [19335]
Yes it does ... and Art's explaination straightened me out ...
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"Fernando Nunes" <domusonline@gmail.com>
To:
ids@iiug.org
Date:
03/15/2010 12:51 PM
Subject:
Re: Strange sql behavior [19333]
Sent by:
ids-bounces@iiug.org
s_h_w_status or s_hw_status?
Does the column exist on table ps_s_payperiod? If it does it is correct=
behavior (You have my sympathy if you find this odd...)
regards
On Mon, Mar 15, 2010 at 3:39 PM, Peter_Logan@spartanstores.com <
Peter_Logan@spartanstores.com> wrote:
> Hello,
>
> I'm running ids 11.50.fc4 on Aix 5.3 ... the following query is givin=
g
> strange behavior ... Any advice is appreciated...
>
> SELECT count(*) FROM ps_S_PAYPERIOD_STS WHERE S_HW_STATUS NOT IN
(SELECT
> S_HW_STATUS FROM ps_S_UL_EE_STS_MVW) ;
>
> This returns the count of 0...
>
> SELECT S_HW_STATUS FROM ps_S_UL_EE_STS_MVW ;> This returns a -217, column not found, s_h_w_status ...
>
> The -217 is correct since this column doesn't exist in the view.
>
> The question is why isn't the first sql also returning the error. It'=
s
> like it doesn't care ...
>
> Please advise...
>
> Peter Logan
> Senior Database Administrator
> Phone: 616/878-8309
>
>
>
>
***********************************************************************=
*
*******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
--0016e6db2b159803c00481d9aded
***********************************************************************=
*
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
***********************************************************************=
*
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
***********************************************************************=
********
Forum Note: Use "Reply" to post a response in the discussion forum.
=
M.P. this is what we have:
Answers:
1) physical cpus on both are 8
2) The number of cpuvps on primary is set to 6, on secondary it is set
to 4
3) OFF_RECVRY_THREADS is set to 10
4) Secondary is currently down. cannot give onstat -g cpu info at the
time.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
<mailto:2244650553@messaging.sprintpcs.com>
Page via Skytel: 2244650553@sprint.skytel.com
<mailto:2244650553@sprint.skytel.com>
Informix or MySQL Primary: 9110210@skytel.com
<mailto:9110210@skytel.com>
Informix or MySQL Secondary: 7276872@skytel.com
<mailto:7276872@skytel.com>
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: Villagomez, Mary
Sent: Monday, March 15, 2010 12:29 PM
To: Knox, Ernest
Subject: RE: Long Checkpoints with HDR and 11.5. [19337]
Answers:
1) physical cpus on both are 8
2) The number of cpuvps on primary is set to 6, on secondary it is set
to 4
3) OFF_RECVRY_THREADS is set to 10
4) Secondary is currently down. cannot give onstat -g cpu info at the
time.
Mary Villagomez
IT Database Specialist
Sears Holdings Corporation IT
3333 Beverly Road B2-260A
Hoffman Estates, IL 60179
Phone: 847-286-1768
Pager: 2244650493@sprint.skytel.com
Blackberry: 2244650493@messaging.sprintpcs.com
________________________________
From: Knox, Ernest
Sent: Mon 03/15/2010 12:14 PM
To: Villagomez, Mary
Subject: FW: Long Checkpoints with HDR and 11.5. [19337]
Answer these and I can send it back.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Madison Pruet
Sent: Monday, March 15, 2010 12:08 PM
To: ids@iiug.org
Subject: RE: Long Checkpoints with HDR and 11.5. [19337]
It sounds like there might be an imbalance between the primary and the
secondary. A couple of questions...
1) How many physical CPUs on the primary/secondary?
2) How many CPUVPS on the primary/secondary?
3) What is OFF_RECVRY_THREADS set to?
4) What is the output of onstat -g cpu on the secondary?
M.P.
=
From: "Knox, Ernest" <Ernest.Knox@searshc.com> =
=
To: ids@iiug.org =
=
Date: 03/15/2010 12:02 PM =
=
Subject: RE: Long Checkpoints with HDR and 11.5. [19336] =
=
Sent by: ids-bounces@iiug.org =
=
We are receiving extreme slowness and long checkpoints on our 11.50.FC3=
and now 11.50.FC5 versions on AIX 6.1, while running HDR.
Everything comes back to life when we cut HDR, but began again once we
upgraded to FC5.
Has anyone else had this issue and what was the resolution?
IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to=
be the answer.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS=
!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Peter_Logan@spartanstores.com
Sent: Monday, March 15, 2010 11:53 AM
To: ids@iiug.org
Subject: Re: Strange sql behavior [19335]
Yes it does ... and Art's explaination straightened me out ...
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"Fernando Nunes" <domusonline@gmail.com>
To:
ids@iiug.org
Date:
03/15/2010 12:51 PM
Subject:
Re: Strange sql behavior [19333]
Sent by:
ids-bounces@iiug.org
s_h_w_status or s_hw_status?
Does the column exist on table ps_s_payperiod? If it does it is correct=
behavior (You have my sympathy if you find this odd...)
regards
On Mon, Mar 15, 2010 at 3:39 PM, Peter_Logan@spartanstores.com <
Peter_Logan@spartanstores.com> wrote:
> Hello,
>
> I'm running ids 11.50.fc4 on Aix 5.3 ... the following query is givin=
g
> strange behavior ... Any advice is appreciated...
>
> SELECT count(*) FROM ps_S_PAYPERIOD_STS WHERE S_HW_STATUS NOT IN
(SELECT
> S_HW_STATUS FROM ps_S_UL_EE_STS_MVW) ;
>
> This returns the count of 0...
>
> SELECT S_HW_STATUS FROM ps_S_UL_EE_STS_MVW ;> This returns a -217, column not found, s_h_w_status ...
>
> The -217 is correct since this column doesn't exist in the view.
>
> The question is why isn't the first sql also returning the error. It'=
s
> like it doesn't care ...
>
> Please advise...
>
> Peter Logan
> Senior Database Administrator
> Phone: 616/878-8309
>
>
>
>
***********************************************************************=
*
*******
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
<http://informix-technology.blogspot.com/>
My email works... but I don't check it frequently...
--0016e6db2b159803c00481d9aded
***********************************************************************=
*
*******
Forum Note: Use "Reply" to post a response in the discussion forum.
*************
Ernie Knox wrote: We are receiving extreme slowness and long checkpoints on our 11.50.FC3 and now 11.50.FC5 versions on AIX 6.1, while running HDR. Everything comes back to life when we cut HDR, but began again once we upgraded to FC5. Has anyone else had this issue and what was the resolution? IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to be the answer. Thanks, ******************************************************************* Ernie Knox IT Database Administrator Specialist Sears Holdings 3333 Beverly Rd., B4-266A Hoffman Estates, IL. 60179 Office: (847) 286-5735 Email: Ernest.Knox@searshc.com Blackberry: 2244650553@messaging.sprintpcs.com Page via Skytel: 2244650553@sprint.skytel.com Informix or MySQL Primary: 9110210@skytel.com Informix or MySQL Secondary: 7276872@skytel.com " Yes we can make a Change! " " It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS! " " Lets not forget - GO Pistons and Red Wings! " GSU ******************************************************************* Responce: Have you recently modified your LRU min/max parameters? It not, what do you have them set to on both the primary and secondary? If LRU max dirty is big (as you increased it because of the interval/non-blocking checkpoints), would you now be doing something on the primary that is actually causing your systems to reach that larger number of modified buffers prior to checkpointing? If so it could be the buffer pool flushing on the secondary as part of a checkpoint is done differently then the normal non-blocking buffer pool flushing on primaries. Because of this, if you have very active logical logging systems, you can back up your primary due to various mechanisms used on the seconadry. In FC5, if I remember correctly, we tried to help this out by detecting longer buffer pool flushes on the secondary and automatically reducing LRU min/max on the secondaries, but I don't recall how aggressive we were in reducing it, additionally, you'd have to hit at least 1 long checkpoint to get the LRU settings to take effect. In 11.50.xC6 we introduced logical log staging on secondaries during checkpoints (if enabled via the $ONCONFIG file) which I believe should completely negate the affect of the long checkpoint on the secondary from impacting the primary at all. Of course this is all based on the assumption that the long checkpoints are caused by a large amount of buffer flushing due to modifications to LRU min/max on the primary (which people started to do on 11.10 and 11.50 because of interval checkpoints). If/when you re-enable HDR I'd try to monitor the number of dirty buffers on your primary and secondary prior to checkpointing, especially if you have increased LRU min/max dirty (compared to what you would have been using back in pre 11.x days and blocking checkpoints). Jacques Renaut Informix Advanced Support APD Team
No, Jacques, IDS secondaries do not share buffer pool updates with the primary. Only the logical log buffers are transferred, so only the server's transaction rate would affect the data transfer rate between the servers and that would be about the same regardless of the LRU_MIN/MAX_DIRTY settings. Now changing the size of the logical log buffer on the other hand, would have some effect, but minor, and the discrepancy between the number of CPU VPs on the two servers will have a much bigger effect. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) IIUG Board of Directors (art@iiug.org) See you at the 2010 IIUG Informix Conference April 25-28, 2010 Overland Park (Kansas City), KS www.iiug.org/conf Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Mon, Mar 15, 2010 at 2:03 PM, JACQUES RENAUT <jrenaut@us.ibm.com> wrote: > Ernie Knox wrote: > We are receiving extreme slowness and long checkpoints on our 11.50.FC3 > and now 11.50.FC5 versions on AIX 6.1, while running HDR. > > Everything comes back to life when we cut HDR, but began again once we > upgraded to FC5. > > Has anyone else had this issue and what was the resolution? > > IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to > be the answer. > > Thanks, > ******************************************************************* > Ernie Knox > IT Database Administrator Specialist > Sears Holdings > 3333 Beverly Rd., B4-266A > Hoffman Estates, IL. 60179 > Office: (847) 286-5735 > Email: Ernest.Knox@searshc.com > Blackberry: 2244650553@messaging.sprintpcs.com > Page via Skytel: 2244650553@sprint.skytel.com > Informix or MySQL Primary: 9110210@skytel.com > Informix or MySQL Secondary: 7276872@skytel.com > > " Yes we can make a Change! " > " It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS! > " > " Lets not forget - GO Pistons and Red Wings! " > GSU > ******************************************************************* > > Responce: > > Have you recently modified your LRU min/max parameters? It not, what do you > have them set to on both the primary and secondary? If LRU max dirty is big > (as you increased it because of the interval/non-blocking checkpoints), > would > you now be doing something on the primary that is actually causing your > systems to reach that larger number of modified buffers prior to > checkpointing? If so it could be the buffer pool flushing on the secondary > as > part of a checkpoint is done differently then the normal non-blocking > buffer > pool flushing on primaries. Because of this, if you have very active > logical > logging systems, you can back up your primary due to various mechanisms > used > on the seconadry. In FC5, if I remember correctly, we tried to help this > out > by detecting longer buffer pool flushes on the secondary and automatically > reducing LRU min/max on the secondaries, but I don't recall how aggressive > we > were in reducing it, additionally, you'd have to hit at least 1 long > checkpoint to get the LRU settings to take effect. In 11.50.xC6 we > introduced > logical log staging on secondaries during checkpoints (if enabled via the > $ONCONFIG file) which I believe should completely negate the affect of the > long checkpoint on the secondary from impacting the primary at all. Of > course > this is all based on the assumption that the long checkpoints are caused by > a > large amount of buffer flushing due to modifications to LRU min/max on the > primary (which people started to do on 11.10 and 11.50 because of interval > checkpoints). > > If/when you re-enable HDR I'd try to monitor the number of dirty buffers on > your primary and secondary prior to checkpointing, especially if you have > increased LRU min/max dirty (compared to what you would have been using > back > in pre 11.x days and blocking checkpoints). > > Jacques Renaut > Informix Advanced Support > APD Team > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --0015174c0f3ec920dc0481db0e51
ok - a couple of things.
1) You need to have the same number (or even more) cpuvps on the
secondary.
2) The number of OFF_RECVRY_THREADS should be a near prime number and =
at
least 3 times the number of CPUVPS.
You're already going to make it more difficult for the secondary to kee=
p up
with the primary because it doesn't have the same horsepower. You migh=
t be
able to overcome that a bit by increasing the offline recovery threads.=
Making them a near prime number will tend to balance the usage a bit. =
A
near prime number would be one which is not evenly divisible by
2,5,7,11,13. So 10 is not a good one. Maybe something more like 23 wo=
uld
be better. Remember - the offline recovery threads have to balance the=
updating activity of the primary and the offline recovery threads need =
to
have a balanced load.
We have done some recent work arround checkpoint performance with HDR. =
We
now have the ability to stage the logs in a temporary space which the
checkpoint is underway. We then subsequently will apply those staged l=
ogs.
Another couple of comments.
It is best to not use indexes which are contained within the tablespace=
that the data pages are contained in (i.e. online-5 style indexes).
=
From: "Knox, Ernest" <Ernest.Knox@searshc.com> =
=
To: ids@iiug.org =
=
Date: 03/15/2010 12:43 PM =
=
Subject: RE: Long Checkpoints with HDR and 11.5. [19338] =
=
Sent by: ids-bounces@iiug.org =
=
M.P. this is what we have:
Answers:
1) physical cpus on both are 8
2) The number of cpuvps on primary is set to 6, on secondary it is set
to 4
3) OFF_RECVRY_THREADS is set to 10
4) Secondary is currently down. cannot give onstat -g cpu info at the
time.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
<mailto:2244650553@messaging.sprintpcs.com)>
Page via Skytel: 2244650553@sprint.skytel.com
<mailto:2244650553@sprint.skytel.com>
Informix or MySQL Primary: 9110210@skytel.com
<mailto:9110210@skytel.com>
Informix or MySQL Secondary: 7276872@skytel.com
<mailto:7276872@skytel.com>
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS=
!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: Villagomez, Mary
Sent: Monday, March 15, 2010 12:29 PM
To: Knox, Ernest
Subject: RE: Long Checkpoints with HDR and 11.5. [19337]
Answers:
1) physical cpus on both are 8
2) The number of cpuvps on primary is set to 6, on secondary it is set
to 4
3) OFF_RECVRY_THREADS is set to 10
4) Secondary is currently down. cannot give onstat -g cpu info at the
time.
Mary Villagomez
IT Database Specialist
Sears Holdings Corporation IT
3333 Beverly Road B2-260A
Hoffman Estates, IL 60179
Phone: 847-286-1768
Pager: 2244650493@sprint.skytel.com
Blackberry: 2244650493@messaging.sprintpcs.com
________________________________
From: Knox, Ernest
Sent: Mon 03/15/2010 12:14 PM
To: Villagomez, Mary
Subject: FW: Long Checkpoints with HDR and 11.5. [19337]
Answer these and I can send it back.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS=
!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Madison Pruet
Sent: Monday, March 15, 2010 12:08 PM
To: ids@iiug.org
Subject: RE: Long Checkpoints with HDR and 11.5. [19337]
It sounds like there might be an imbalance between the primary and the
secondary. A couple of questions...
1) How many physical CPUs on the primary/secondary?
2) How many CPUVPS on the primary/secondary?
3) What is OFF_RECVRY_THREADS set to?
4) What is the output of onstat -g cpu on the secondary?
M.P.
=3D
From: "Knox, Ernest" <Ernest.Knox@searshc.com> =3D
=3D
To: ids@iiug.org =3D
=3D
Date: 03/15/2010 12:02 PM =3D
=3D
Subject: RE: Long Checkpoints with HDR and 11.5. [19336] =3D
=3D
Sent by: ids-bounces@iiug.org =3D
=3D
We are receiving extreme slowness and long checkpoints on our 11.50.FC3=
=3D
and now 11.50.FC5 versions on AIX 6.1, while running HDR.
Everything comes back to life when we cut HDR, but began again once we
upgraded to FC5.
Has anyone else had this issue and what was the resolution?
IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to=
=3D
be the answer.
Thanks,
*******************************************************************
Ernie Knox
IT Database Administrator Specialist
Sears Holdings
3333 Beverly Rd., B4-266A
Hoffman Estates, IL. 60179
Office: (847) 286-5735
Email: Ernest.Knox@searshc.com
Blackberry: 2244650553@messaging.sprintpcs.com
Page via Skytel: 2244650553@sprint.skytel.com
Informix or MySQL Primary: 9110210@skytel.com
Informix or MySQL Secondary: 7276872@skytel.com
" Yes we can make a Change! "
" It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS=
=3D
!
"
" Lets not forget - GO Pistons and Red Wings! "
GSU
*******************************************************************
-----Original Message-----
From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of
Peter_Logan@spartanstores.com
Sent: Monday, March 15, 2010 11:53 AM
To: ids@iiug.org
Subject: Re: Strange sql behavior [19335]
Yes it does ... and Art's explaination straightened me out ...
Peter Logan
Senior Database Administrator
Phone: 616/878-8309
From:
"Fernando Nunes" <domusonline@gmail.com>
To:
ids@iiug.org
Date:
03/15/2010 12:51 PM
Subject:
Re: Strange sql behavior [19333]
Sent by:
ids-bounces@iiug.org
s_h_w_status or s_hw_status?
Does the c
This too... ;-) = From: Jacques Renaut/Lenexa/IBM@IBMUS = = To: ids@iiug.org = = Date: 03/15/2010 01:05 PM = = Subject: Re: RE: Long Checkpoints with HDR and 11.5. [19339] = = Sent by: ids-bounces@iiug.org = = Ernie Knox wrote: We are receiving extreme slowness and long checkpoints on our 11.50.FC3= and now 11.50.FC5 versions on AIX 6.1, while running HDR. Everything comes back to life when we cut HDR, but began again once we upgraded to FC5. Has anyone else had this issue and what was the resolution? IBM asked us to upgrade from FC3 to FC5, but that still doesn't seem to= be the answer. Thanks, ******************************************************************* Ernie Knox IT Database Administrator Specialist Sears Holdings 3333 Beverly Rd., B4-266A Hoffman Estates, IL. 60179 Office: (847) 286-5735 Email: Ernest.Knox@searshc.com Blackberry: 2244650553@messaging.sprintpcs.com Page via Skytel: 2244650553@sprint.skytel.com Informix or MySQL Primary: 9110210@skytel.com Informix or MySQL Secondary: 7276872@skytel.com " Yes we can make a Change! " " It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS= ! " " Lets not forget - GO Pistons and Red Wings! " GSU ******************************************************************* Responce: Have you recently modified your LRU min/max parameters? It not, what do= you have them set to on both the primary and secondary? If LRU max dirty is= big (as you increased it because of the interval/non-blocking checkpoints),= would you now be doing something on the primary that is actually causing your= systems to reach that larger number of modified buffers prior to checkpointing? If so it could be the buffer pool flushing on the second= ary as part of a checkpoint is done differently then the normal non-blocking buffer pool flushing on primaries. Because of this, if you have very active logical logging systems, you can back up your primary due to various mechanisms= used on the seconadry. In FC5, if I remember correctly, we tried to help thi= s out by detecting longer buffer pool flushes on the secondary and automatica= lly reducing LRU min/max on the secondaries, but I don't recall how aggress= ive we were in reducing it, additionally, you'd have to hit at least 1 long checkpoint to get the LRU settings to take effect. In 11.50.xC6 we introduced logical log staging on secondaries during checkpoints (if enabled via t= he $ONCONFIG file) which I believe should completely negate the affect of = the long checkpoint on the secondary from impacting the primary at all. Of course this is all based on the assumption that the long checkpoints are cause= d by a large amount of buffer flushing due to modifications to LRU min/max on = the primary (which people started to do on 11.10 and 11.50 because of inter= val checkpoints). If/when you re-enable HDR I'd try to monitor the number of dirty buffer= s on your primary and secondary prior to checkpointing, especially if you ha= ve increased LRU min/max dirty (compared to what you would have been using= back in pre 11.x days and blocking checkpoints). Jacques Renaut Informix Advanced Support APD Team ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
Art Kagel wrote: No, Jacques, IDS secondaries do not share buffer pool updates with the primary. Only the logical log buffers are transferred, so only the server's transaction rate would affect the data transfer rate between the servers and that would be about the same regardless of the LRU_MIN/MAX_DIRTY settings. Now changing the size of the logical log buffer on the other hand, would have some effect, but minor, and the discrepancy between the number of CPU VPs on the two servers will have a much bigger effect. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Reply: No the secondaries do not share buffer pools, however, as log records are applied on the secondary, they mark buffers as dirty, and as part of the checkpoint processing (ie when the secondary gets the logical log record of the checkpoint from the primary) the secondary is required to flush the buffer pool. Sorry if I was unclear in my first post, but trust me, I know what I'm talking about as I entered the defect. The way this buffer pool flush is done on the secondary can indeed hold up the primary. What you would see if user threads waiting on the log buff condition and you would see the dr_prsend thread waiting on the drcb_bqe condition. Now you can see that for other reasons as well, but one reason is cause due to large buffer pool flushing on the secondary. Please refer to APAR IC60754. Jacques Renaut IBM Informix Advanced Support APD Team
Well, the Informix tech replied as follows: Further research revealed that the defect fix in 11.50.xC5 is only a partial fix to the problem. The 11.50.xC6 contains another fix which is supposed to finish the job. We'll let you know how it turns out with FC6. Thanks for all suggestions, ******************************************************************* Ernie Knox IT Database Administrator Specialist Sears Holdings 3333 Beverly Rd., B4-266A Hoffman Estates, IL. 60179 Office: (847) 286-5735 Email: Ernest.Knox@searshc.com Blackberry: 2244650553@messaging.sprintpcs.com Page via Skytel: 2244650553@sprint.skytel.com Informix or MySQL Primary: 9110210@skytel.com Informix or MySQL Secondary: 7276872@skytel.com " Yes we can make a Change! " " It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS! " " Lets not forget - GO Pistons and Red Wings! " GSU ******************************************************************* -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of JACQUES RENAUT Sent: Monday, March 15, 2010 2:08 PM To: ids@iiug.org Subject: Re: RE: Long Checkpoints with HDR and 11.5. [19343] Art Kagel wrote: No, Jacques, IDS secondaries do not share buffer pool updates with the primary. Only the logical log buffers are transferred, so only the server's transaction rate would affect the data transfer rate between the servers and that would be about the same regardless of the LRU_MIN/MAX_DIRTY settings. Now changing the size of the logical log buffer on the other hand, would have some effect, but minor, and the discrepancy between the number of CPU VPs on the two servers will have a much bigger effect. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Reply: No the secondaries do not share buffer pools, however, as log records are applied on the secondary, they mark buffers as dirty, and as part of the checkpoint processing (ie when the secondary gets the logical log record of the checkpoint from the primary) the secondary is required to flush the buffer pool. Sorry if I was unclear in my first post, but trust me, I know what I'm talking about as I entered the defect. The way this buffer pool flush is done on the secondary can indeed hold up the primary. What you would see if user threads waiting on the log buff condition and you would see the dr_prsend thread waiting on the drcb_bqe condition. Now you can see that for other reasons as well, but one reason is cause due to large buffer pool flushing on the secondary. Please refer to APAR IC60754. Jacques Renaut IBM Informix Advanced Support APD Team ************************************************************************ ******* Forum Note: Use "Reply" to post a response in the discussion forum.
Ernie Knox wrote: Well, the Informix tech replied as follows: Further research revealed that the defect fix in 11.50.xC5 is only a partial fix to the problem. The 11.50.xC6 contains another fix which is supposed to finish the job. We'll let you know how it turns out with FC6. Thanks for all suggestions Reply: Yes, the "another fix which is supposed to finish the job" is the logical log staging that I mentioned in my original post. Just a reminder, this "fix" is not enabled by default. You need to "turn it on" by configuring LOG_STAGING_DIR and LOG_INDEX_BUILDS in your onconfig file. You should be able to refer to the Informix Administrator's Guide, Using High-Availabilty Data Replication Chapter, the section on "Preventing Blocking Checkpoints on HDR Servers". Just make sure you are looking at the 11.50.xC6 Doc. Jacques Renaut IBM Informix Advanced Support APD Team
Just to be a bit more precise... The mini-feature to stage the log files is an outgrowth of the delayed apply feature for RSS that was implemented last summer. Also -- the ot= her suggestions that have been raised are still valid because they will mak= e it easier for the secondary to balance with the primary. M.P. = From: Jacques Renaut/Lenexa/IBM@IBMUS = = To: ids@iiug.org = = Date: 03/16/2010 04:21 PM = = Subject: Re: RE: RE: Long Checkpoints with HDR and 11.5. [19346] = = Sent by: ids-bounces@iiug.org = = Ernie Knox wrote: Well, the Informix tech replied as follows: Further research revealed that the defect fix in 11.50.xC5 is only a partial fix to the problem. The 11.50.xC6 contains another fix which is= supposed to finish the job. We'll let you know how it turns out with FC6. Thanks for all suggestions Reply: Yes, the "another fix which is supposed to finish the job" is the logic= al log staging that I mentioned in my original post. Just a reminder, this "fi= x" is not enabled by default. You need to "turn it on" by configuring LOG_STAGING_DIR and LOG_INDEX_BUILDS in your onconfig file. You should = be able to refer to the Informix Administrator's Guide, Using High-Availabilty = Data Replication Chapter, the section on "Preventing Blocking Checkpoints on= HDR Servers". Just make sure you are looking at the 11.50.xC6 Doc. Jacques Renaut IBM Informix Advanced Support APD Team ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =
We will look into this. I missed this among the many other replies. Sorry about that. Thanks, ******************************************************************* Ernie Knox IT Database Administrator Specialist Sears Holdings 3333 Beverly Rd., B4-266A Hoffman Estates, IL. 60179 Office: (847) 286-5735 Email: Ernest.Knox@searshc.com Blackberry: 2244650553@messaging.sprintpcs.com Page via Skytel: 2244650553@sprint.skytel.com Informix or MySQL Primary: 9110210@skytel.com Informix or MySQL Secondary: 7276872@skytel.com " Yes we can make a Change! " " It's always a great day to watch Sports - GO LIONS, TIGERS, and BEARS! " " Lets not forget - GO Pistons and Red Wings! " GSU ******************************************************************* -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of JACQUES RENAUT Sent: Tuesday, March 16, 2010 4:21 PM To: ids@iiug.org Subject: Re: RE: RE: Long Checkpoints with HDR and 11.5. [19346] Ernie Knox wrote: Well, the Informix tech replied as follows: Further research revealed that the defect fix in 11.50.xC5 is only a partial fix to the problem. The 11.50.xC6 contains another fix which is supposed to finish the job. We'll let you know how it turns out with FC6. Thanks for all suggestions Reply: Yes, the "another fix which is supposed to finish the job" is the logical log staging that I mentioned in my original post. Just a reminder, this "fix" is not enabled by default. You need to "turn it on" by configuring LOG_STAGING_DIR and LOG_INDEX_BUILDS in your onconfig file. You should be able to refer to the Informix Administrator's Guide, Using High-Availabilty Data Replication Chapter, the section on "Preventing Blocking Checkpoints on HDR Servers". Just make sure you are looking at the 11.50.xC6 Doc. Jacques Renaut IBM Informix Advanced Support APD Team ************************************************************************ ******* Forum Note: Use "Reply" to post a response in the discussion forum.