Re: Extremly long checkpoints: How to find the reason and solve the problem (7.31 FD3)
Posted in 2004
Topics: High Availability & Replication, Server Administration, Logging & Checkpoints, Cloud, Docker & Containers
Are you Replicating to other server ?
Some time ago I got very long checkpoints in an
Informix Server with High Availability Data Replication (HDR)
on the primary server.
I found out that the reason was an extremely complex query
executed in the secondary server (Stand By Server) for generating
a report (DSS Reports in OLTP System , not a good idea !! ).
Chekpoints for HDR Systems are always Synchronous .
It doesn't matter if you configure the system to be Asynchronous (I don't
remember
the name of the parameter in the Onconfig File), the only thing that really
gets Asynchronous are the transactions (No 2-Fase Commit Protocol),
the onconfig's parameter should be named TwoFaseCommit instead of
the name that I don't remember.
Primary Server Always wait an acknowledge message of the other servers
for finishing its own checkpoint.
If you are not replicating ignore this message, I just wanted to express
my frustrating experience with HDR.
Enterprise Replication (ER) would solve the problem.
Regards
----- Original Message -----
From: "John Hardin" <johnh@aproposretail.com>
To: <informix-list@iiug.org>
Sent: Tuesday, May 18, 2004 10:12 AM
Subject: Re: Extremly long checkpoints: How to find the reason and solve the
problem (7.31 FD3)
> On Tue, 11 May 2004 01:27:16 -0700, J?rg Spilker wrote:
>
> > we have some serious problems with extremly long checkpoints.
> > Interactive users can't really work.
>
> > 09:30:28 Checkpoint Completed: duration was 1465 seconds.>
> Youch!
>
> What is the output from "onstat -FR"?
>
> --
> John Hardin KA7OHZ <johnh@aproposretail.com>
> Internal Systems Administrator voice: (425) 672-1304
> Apropos Retail Management Systems, Inc. fax: (425) 672-0192
> -----------------------------------------------------------------------
> ...the Fates notice those who buy chainsaws...
> -- www.darwinawards.com
>
sending to informix-list
Francisco Roldan wrote: > Are you Replicating to other server ? > Some time ago I got very long checkpoints in an > Informix Server with High Availability Data Replication (HDR) > on the primary server. > > I found out that the reason was an extremely complex query > executed in the secondary server (Stand By Server) for generating > a report (DSS Reports in OLTP System , not a good idea !! ). > > Chekpoints for HDR Systems are always Synchronous . > It doesn't matter if you configure the system to be Asynchronous (I don't > remember > the name of the parameter in the Onconfig File), the only thing that really > gets Asynchronous are the transactions (No 2-Fase Commit Protocol), > the onconfig's parameter should be named TwoFaseCommit instead of > the name that I don't remember. > Primary Server Always wait an acknowledge message of the other servers > for finishing its own checkpoint. > > If you are not replicating ignore this message, I just wanted to express > my frustrating experience with HDR. > Enterprise Replication (ER) would solve the problem. > > Regards > > snip ... DRINTERVAL -1 (Synchronous) or DRINTERVAL > 0 Asynchronous There appears to be some activity which the checkpoint is dependent on (not the checkpoint itself) which is synchronous; you can see this when the checkpoint completes on the primary but hasn't started / completed on the secondary when DRINTERVAL > 0. Something to do with flushing the physical log buffer on the secondary, and threads in critical section. What was the "extremely complex query executed in the secondary server (Stand By Server) for generating a report (DSS Reports in OLTP System , not a good idea !! )." (I thought that was a nice way to split out DSS from the OLTP primary by putting DSS on the secondary). Did you log a Tech Support case?? Were there a lot of writes involved on the secondary to temp tables?