HA Checkpoint Trigger
Posted in 2015
User experienced blocking checkpoints triggered by "HA" on Informix 11.70.FC7W2 with primary and read-only RSS servers. Root cause was physical log running out of space during checkpoint processing. Solutions offered: increase physical log size using onparams (bouncing required), turn off AUTO_CKPTS and set checkpoint interval, or use newer physical log management features.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Performance & Tuning, Triggers, Constraints & Referential Integrity, Logging & Checkpoints, Clustering, Grid & MACH11
Hi, We are running 11.70.FC7W2 with 1 primary and 1 read-only RSS server. We just noticed 2 blocking checkpoints and the trigger says "HA". Anyone have any experience with this? I saw the documentation gave 3 examples of what can cause this and we did not add a new RSS server or promote any secondary servers. The phys log looked ok on the RSS server. However I did see the following message on the primary online.log: 13:33:03 Performance Advisory: The physical log is running out of room during checkpoint processing. 13:33:03 Results: Transactions are being blocked until the checkpoint is complete. 13:33:03 Action: Increase the physical log size. Any ideas? In the past, on a rare occasion, we would see PLOG as the trigger during heavy write activity. --Dave From the Manual: High availability. For example: - A new RSS or SDS node is added to a High Availability cluster - A secondary server is promoted to a primary server - The physical log file is low on a secondary server --f46d04138c7dbdecaa05117fb1ab
Dave -
Have never encountered HA as a trigger for CKPT, but it sounds like it's
simply the plog is too small. We raise a "chkpt request" when the plog hits
75% full, but will keep writing pages to the plog. (We must let critical
sections finish before we start flushing pages - but the ckpt flag is up now).
The plog will wrap/overflow ("physical log overflow has occurred" is the msg
in the msg log) or write to disk if PLOG_OVERFLOW_PATH (I think that's it) is
set in the onconfig. The engine is simply blocking transactions to avoid plog
overflow. Overflow by itself isn't bad - but if the engines dies you'll get
stuck in fast recovery coming back up. If PLOG_OVERFLOW_PATH is defined, pages
that were written there will start being returned to the plog, but this can be
VERY slow depending on # of pages written to disk. That's why it's
recommending increasing it's size. Since there is 1 plog, you can't increase
it "on the fly" - the engine has to bounce. AND - if you've been around a long
time, the old approach was to modify PHYSIZE and bounce. As of either 10 or
11, that won't work anymore, even though the msg log will report a changed
size (saw this just last month). You must use onparams to resize/move it (at
least up through 11.7).
Any changes in 12+ in the area anyone?
Thanks -
Mark
www.markscranton.com
mark@markscranton.com
The Mark Scranton Group
"All Informix ... all the time."
I saw this just a few weeks ago at a client site. I followed advice from Art Kagel to turn off AUTO_CKPTS and set a checkpoint interval, and it seems to have worked. Mike -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Informix DBA Sent: Tuesday, March 17, 2015 11:57 AM To: ids@iiug.org Subject: HA Checkpoint Trigger [34833] Hi, We are running 11.70.FC7W2 with 1 primary and 1 read-only RSS server. We just noticed 2 blocking checkpoints and the trigger says "HA". Anyone have any experience with this? I saw the documentation gave 3 examples of what can cause this and we did not add a new RSS server or promote any secondary servers. The phys log looked ok on the RSS server. However I did see the following message on the primary online.log: 13:33:03 Performance Advisory: The physical log is running out of room during checkpoint processing. 13:33:03 Results: Transactions are being blocked until the checkpoint is complete. 13:33:03 Action: Increase the physical log size. Any ideas? In the past, on a rare occasion, we would see PLOG as the trigger during heavy write activity. --Dave >From the Manual: High availability. For example: - A new RSS or SDS node is added to a High Availability cluster - A secondary server is promoted to a primary server - The physical log file is low on a secondary server --f46d04138c7dbdecaa05117fb1ab **************************************************************************** *** Forum Note: Use "Reply" to post a response in the discussion forum.
The trigger of a checkpoint is not always related to if the servers block
or not. The
blocking is generally due to a lack of resources. In your case it is
because the
physical log was short of resources.
There has been some great administrative improvements with the physical
log. To highlight
a few
New dedicated plog space
Expand the physical log
Move the physical log online
1. The Plog Space
This is a new type of space which will allow only the physical log
in it. No other object is allowed to be stored in a plog space.
By default the plog space grow dynamical along with the
physical log. This saves you having to move the plog and
take twice the space.
2. Move the physical log to a new location while the system is running.
This is the
reason why the undocumented way of changing the configuration file and
bouncing
the server no longer works. You do not need to bounce the server, but
rather just
issue the command and the physical log will be moved during your next
checkpoint.
John F. Miller III
STSM, Lead Architect
miller3@us.ibm.com
503-747-1366
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 03/17/2015 01:35:40 PM:
> From: "MARK SCRANTON" <mark@markscranton.com>
> To: ids@iiug.org
> Date: 03/17/2015 01:36 PM
> Subject: Re: HA Checkpoint Trigger [34836]
> Sent by: ids-bounces@iiug.org
>
> Dave -
>
> Have never encountered HA as a trigger for CKPT, but it sounds like it's
> simply the plog is too small. We raise a "chkpt request" when the plog
hits
> 75% full, but will keep writing pages to the plog. (We must let critical
> sections finish before we start flushing pages - but the ckpt flag
> is up now).
> The plog will wrap/overflow ("physical log overflow has occurred" is the
msg
> in the msg log) or write to disk if PLOG_OVERFLOW_PATH (I think that's
it) is
> set in the onconfig. The engine is simply blocking transactions to avoid
plog
> overflow. Overflow by itself isn't bad - but if the engines dies you'll
get
> stuck in fast recovery coming back up. If PLOG_OVERFLOW_PATH is
> defined, pages
> that were written there will start being returned to the plog, but
> this can be
> VERY slow depending on # of pages written to disk. That's why it's
> recommending increasing it's size. Since there is 1 plog, you can't
increase
> it "on the fly" - the engine has to bounce. AND - if you've been
> around a long
> time, the old approach was to modify PHYSIZE and bounce. As of either 10
or
> 11, that won't work anymore, even though the msg log will report a
changed
> size (saw this just last month). You must use onparams to resize/move it
(at
> least up through 11.7).
>
> Any changes in 12+ in the area anyone?
>
> Thanks -
> Mark
> www.markscranton.com
> mark@markscranton.com
> The Mark Scranton Group
> "All Informix ... all the time."
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Thanks for the info John. What flavor of IDS brought this change? Apparently I missed it along the way or I'm stuck on an older version most of the time. Thanks - Mark Mark Scranton The Mark Scranton Group www.markscranton.com
The physical logging performed on the primary is not necessarily going = to be the same as it is on the secondary. A common reason for this is the= creation of an index. Consider that the primary may be doing a checkpoint as the index is bei= ng created during the final phase of the index creation (as the sorted ind= ex pages are being written back to the chunk). Maybe 70% of the index pag= es writes came prior to the checkpoint and were flushed as part of that checkpoint. 30% would be written after the checkpoint and would remain= as dirty pages to be flushed as part of the next checkpoint. So on the primary, the checkpoint flushed 70% of the pages. The newly created index is transferred to the secondary after the index= is fully created on the primary. However, on the secondary 100% of the pa= ges will be written after the last checkpoint. This will tend to cause the= physical log to become full on the secondary even though it was not ful= l on the primary. When the PLOG on the secondary gets dangerously close to full, it is forced to request an emergency checkpoint request. Normally the PLOG on the secondary is used the same as on the primary. = But with things requiring significant PLOG activity and which are actually performed first on the primary and then on the secondary, the PLOG usag= e will drift between the primary and secondary. The primary will tend to= have a fuller PLOG than the secondary and cause a checkpoint, but at th= at time the secondary PLOG is not near as full. Later as the work gets mo= ved to the secondary, the secondary PLOG will become full while the primary= is not so much. Generally this is found in index creation. M.Pruet From: "Informix DBA" <in4mixdba@gmail.com> To: ids@iiug.org Date: 03/17/2015 12:59 PM Subject: HA Checkpoint Trigger [34833] Sent by: ids-bounces@iiug.org Hi, We are running 11.70.FC7W2 with 1 primary and 1 read-only RSS server. W= e just noticed 2 blocking checkpoints and the trigger says "HA". Anyone h= ave any experience with this? I saw the documentation gave 3 examples of what can cause this and we d= id not add a new RSS server or promote any secondary servers. The phys log= looked ok on the RSS server. However I did see the following message on= the primary online.log: 13:33:03 Performance Advisory: The physical log is running out of room during checkpoint processing. 13:33:03 Results: Transactions are being blocked until the checkpoint i= s complete. 13:33:03 Action: Increase the physical log size. Any ideas? In the past, on a rare occasion, we would see PLOG as the trigger during heavy write activity. --Dave >From the Manual: High availability. For example: - A new RSS or SDS node is added to a High Availability cluster - A secondary server is promoted to a primary server - The physical log file is low on a secondary server --f46d04138c7dbdecaa05117fb1ab ***********************************************************************= ******** Forum Note: Use "Reply" to post a response in the discussion forum. =