Checkpoint duration too long sometimes--no fixed pattern
Posted in 2003
Topics: Storage & Space Management, Logging & Checkpoints, Platform-Specific Issues
Environment:
Informix Dynamic Server 7.31 UD2 R1
HP-UX 11.0
As you can see below
10:44:50 Checkpoint Completed: duration was 0
seconds.
10:59:59 Checkpoint Completed: duration was 0
seconds.
11:15:08 Checkpoint Completed: duration was 0
seconds.
11:30:17 Checkpoint Completed: duration was 0
seconds.
11:45:26 Checkpoint Completed: duration was 0
seconds.
12:00:53 Checkpoint Completed: duration was 19
seconds.
12:16:02 Checkpoint Completed: duration was 0
seconds.
12:31:34 Checkpoint Completed: duration was 23
seconds.
12:46:44 Checkpoint Completed: duration was 0
seconds.
13:01:54 Checkpoint Completed: duration was 1
seconds.
13:17:05 Checkpoint Completed: duration was 2
seconds.
13:32:15 Checkpoint Completed: duration was 0
seconds.
13:47:24 Checkpoint Completed: duration was 0
seconds.
14:02:33 Checkpoint Completed: duration was 1
seconds.
14:17:43 Checkpoint Completed: duration was 1
seconds.
14:32:54 Checkpoint Completed: duration was 0
seconds.
14:48:03 Checkpoint Completed: duration was 0
seconds.
15:03:13 Checkpoint Completed: duration was 2
seconds.
15:18:23 Checkpoint Completed: duration was 0
seconds.
15:35:08 Checkpoint Completed: duration was 97
seconds.
we have at times checkpoints of duration more than a
minute. This is totally unacceptable in production
environment.
We are unable to diagnose a pattern for this. It could
happen at any time throughout the day. I don't have
any clues where to look for and what could be the
possible reason for this behaviour.
Here are the related checkpoint parameters which we
have set in our environment for your reference.
MULTIPROCESSOR 1 # 0 forsingle-processor, 1 for multi-processor
NUMCPUVPS 5 # Number of user (cpu)vps
SINGLE_CPU_VP 0
BUFFERS 500000
CLEANERS 128
CKPTINTVL 900
LRUS 128
LRU_MAX_DIRTY 2
LRU_MIN_DIRTY 1
Also interesting most of the times we have 0 LRU
writes in the system and only writes happen are the
chunkwrites at the time of checkpoint.
Fg Writes LRU Writes Chunk Writes
0 2 14620
Are we hitting some kind of bug? Or doing something
wrong.
Any help on this is highly appreciated.
Do let me know if you need more information.
Thanks and regards,
Vineet
__________________________________
Do you Yahoo!?
The New Yahoo! Search - Faster. Easier. Bingo.
http://search.yahoo.com
Vineet,
May be you're hitting a bug. I bet that this is happening due to some
adhoc report or query(ies). Longtime back at my work I place, I had a
similar situation. Found out that some of the reports were run, with
sqls creating temp tables with log. I went thru all the sqls and
changed temp table creation to be "with no log". That subsided the
problem. During high checkpoint time, you could trace the sql with
"onstat -g sql" and "onstat -g ses <sql pid>" would give you defaulting
report/application. And I am sure you know that it could be better to
reduce checkpoint interval period to 5/6 minutes.
In my system, I monitor errorlog file for checkpoints for more than 6
seconds. If it does occur then I page myself so that I can quickly see
what is going on.
hope this helps,
Vivek
> -----Original Message-----
> From: vin.us [mailto:vin_us@yahoo.com]
> Sent: Wednesday, April 30, 2003 1:52 PM
> To: ids; forum.subscriber; vin.us
> Subject: Checkpoint duration too long sometimes--no fixed
> pattern [1054]
>
>
>
> Environment:
>
>
> Informix Dynamic Server 7.31 UD2 R1
>
> HP-UX 11.0
>
> As you can see below
>
> 10:44:50 Checkpoint Completed: duration was 0
> seconds.
> 10:59:59 Checkpoint Completed: duration was 0
> seconds.
> 11:15:08 Checkpoint Completed: duration was 0
> seconds.
> 11:30:17 Checkpoint Completed: duration was 0
> seconds.
> 11:45:26 Checkpoint Completed: duration was 0
> seconds.
> 12:00:53 Checkpoint Completed: duration was 19
> seconds.
> 12:16:02 Checkpoint Completed: duration was 0
> seconds.
> 12:31:34 Checkpoint Completed: duration was 23
> seconds.
> 12:46:44 Checkpoint Completed: duration was 0
> seconds.
> 13:01:54 Checkpoint Completed: duration was 1
> seconds.
> 13:17:05 Checkpoint Completed: duration was 2
> seconds.
> 13:32:15 Checkpoint Completed: duration was 0
> seconds.
> 13:47:24 Checkpoint Completed: duration was 0
> seconds.
> 14:02:33 Checkpoint Completed: duration was 1
> seconds.
> 14:17:43 Checkpoint Completed: duration was 1
> seconds.
> 14:32:54 Checkpoint Completed: duration was 0
> seconds.
> 14:48:03 Checkpoint Completed: duration was 0
> seconds.
> 15:03:13 Checkpoint Completed: duration was 2
> seconds.
> 15:18:23 Checkpoint Completed: duration was 0
> seconds.
> 15:35:08 Checkpoint Completed: duration was 97
> seconds.
>
> we have at times checkpoints of duration more than a
> minute. This is totally unacceptable in production
> environment.
>
> We are unable to diagnose a pattern for this. It could
> happen at any time throughout the day. I don't have
> any clues where to look for and what could be the
> possible reason for this behaviour.
> Here are the related checkpoint parameters which we
> have set in our environment for your reference.
>
> MULTIPROCESSOR 1 # 0 for> single-processor, 1 for multi-processor
> NUMCPUVPS 5 # Number of user (cpu)> vps
> SINGLE_CPU_VP 0>
>
> BUFFERS 500000
>
> CLEANERS 128
>
> CKPTINTVL 900
>
> LRUS 128
>
> LRU_MAX_DIRTY 2
>
> LRU_MIN_DIRTY 1>
>
> Also interesting most of the times we have 0 LRU
> writes in the system and only writes happen are the
> chunkwrites at the time of checkpoint.
>
> Fg Writes LRU Writes Chunk Writes
> 0 2 14620
>
> Are we hitting some kind of bug? Or doing something
> wrong.
>
> Any help on this is highly appreciated.
>
> Do let me know if you need more information.
>
> Thanks and regards,
>
> Vineet
>
>
>
>
>
> __________________________________
> Do you Yahoo!?
> The New Yahoo! Search - Faster. Easier. Bingo.
> http://search.yahoo.com
>
>
Hi
Sonia,
Thanks for the prompt reply.
What I really don't understand is at the time of
checkpoint only the dirty pages are flushed to the
disk or everything in the buffers are flushed to make
it consistent.
Actually we do have some runaway queries in the system
which keep initiated time to time. Those are select
queries involving multiple table joins and read lot of
data. I have to kill those queries as the load on the
system goes high some time.
so are those queries in any way contributes to the
long checkpoint duration.
Also I haven't checked the inserts,updates happeing on
the system during that time but I think it is uniform
throughout the day.I mean there are no massive
updates, inserts during that time though I can look
for it.
Thanks,
Vineet
--- Sonia Guerra <sguerra@merkafon.com> wrote:
> Hi Vineet,
> As you know checkpoint duration is based into the
> amount of dirty data to be
> flushed to disk from memory, wich means if more
> dirty pages were modified
> between checkpoints, the time spend to flush them
> will be bigger, so I
> recommend you first to look for update statements
> like INSERTS, UPDATES,
> DELETES and UPDATE STATISTICS using onstat -g sql
> command.
>
> I experienced something like this before and I start
> looking for that,
> finally I found developers was running massive
> updates and deletes at this
> time, without my permission if I could say that, so
> the reschedule their
> proccesses and the problem was solved.
>
> You could do the same getting the engine activity
> through a log file as
> result of onstat -g sql, onstat -u, etc. commands
> output.
>
> If you need detailed information or clues to
> monitor, please let me know.
>
> Best Regards,
> Sonia
>
>
> ----- Original Message -----
> From: "Vineet Mehr...." <vin_us@yahoo.com>
> To: <ids@iiug.org>
> Sent: Wednesday, April 30, 2003 3:51 PM
> Subject: Checkpoint duration too long sometimes--no
> fixed pattern [1054]
>
>
> > Environment:
> >
> >
> > Informix Dynamic Server 7.31 UD2 R1
> >
> > HP-UX 11.0
> >
> > As you can see below
> >
> > 10:44:50 Checkpoint Completed: duration was 0
> > seconds.
> > 10:59:59 Checkpoint Completed: duration was 0
> > seconds.
> > 11:15:08 Checkpoint Completed: duration was 0
> > seconds.
> > 11:30:17 Checkpoint Completed: duration was 0
> > seconds.
> > 11:45:26 Checkpoint Completed: duration was 0
> > seconds.
> > 12:00:53 Checkpoint Completed: duration was 19
> > seconds.
> > 12:16:02 Checkpoint Completed: duration was 0
> > seconds.
> > 12:31:34 Checkpoint Completed: duration was 23
> > seconds.
> > 12:46:44 Checkpoint Completed: duration was 0
> > seconds.
> > 13:01:54 Checkpoint Completed: duration was 1
> > seconds.
> > 13:17:05 Checkpoint Completed: duration was 2
> > seconds.
> > 13:32:15 Checkpoint Completed: duration was 0
> > seconds.
> > 13:47:24 Checkpoint Completed: duration was 0
> > seconds.
> > 14:02:33 Checkpoint Completed: duration was 1
> > seconds.
> > 14:17:43 Checkpoint Completed: duration was 1
> > seconds.
> > 14:32:54 Checkpoint Completed: duration was 0
> > seconds.
> > 14:48:03 Checkpoint Completed: duration was 0
> > seconds.
> > 15:03:13 Checkpoint Completed: duration was 2
> > seconds.
> > 15:18:23 Checkpoint Completed: duration was 0
> > seconds.
> > 15:35:08 Checkpoint Completed: duration was 97
> > seconds.
> >
> > we have at times checkpoints of duration more than
> a
> > minute. This is totally unacceptable in production
> > environment.
> >
> > We are unable to diagnose a pattern for this. It
> could
> > happen at any time throughout the day. I don't
> have
> > any clues where to look for and what could be the
> > possible reason for this behaviour.
> > Here are the related checkpoint parameters which
> we
> > have set in our environment for your reference.
> >
> > MULTIPROCESSOR 1 # 0 for> > single-processor, 1 for multi-processor
> > NUMCPUVPS 5 # Number of user
> (cpu)> > vps
> > SINGLE_CPU_VP 0> >
> >
> > BUFFERS 500000
> >
> > CLEANERS 128
> >
> > CKPTINTVL 900
> >
> > LRUS 128
> >
> > LRU_MAX_DIRTY 2
> >
> > LRU_MIN_DIRTY 1> >
> >
> > Also interesting most of the times we have 0 LRU
> > writes in the system and only writes happen are
> the
> > chunkwrites at the time of checkpoint.
> >
> > Fg Writes LRU Writes Chunk Writes
> > 0 2 14620
> >
> > Are we hitting some kind of bug? Or doing
> something
> > wrong.
> >
> > Any help on this is highly appreciated.
> >
> > Do let me know if you need more information.
> >
> > Thanks and regards,
> >
> > Vineet
> >
> >
> >
> >
> >
> > __________________________________
> > Do you Yahoo!?
> > The New Yahoo! Search - Faster. Easier. Bingo.
> > http://search.yahoo.com
> >
> >
> >
>
>
__________________________________
Do you Yahoo!?
The New Yahoo! Search - Faster. Easier. Bingo.
http://search.yahoo.com
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g