Re: page cleaners time out
Posted in 1992
inf_bb@hermes1.sps.mot.com (informix users bb) writes:
>
>we are running online ver 5.00.UC2 with i-star and are experiencing a
>problem with the page cleaners timing out. we are running on a sun
>690 with the database accelerator kernal enhancement.
>
>a tbstat -u gives the following:
>
>(hermes1)/home_hermes1/baskett>tbstat -u
>
>RSAM Version 5.00.UC2 -- On-Line -- Up 08:17:49 -- 31856 Kbytes
>
>Users
>address flags pid user tty wait tout locks nreads nwrites
>10001ba8 ------D 12325 root ttyp5 0 0 0 1384 247
>10001c18 ------D 0 root ttyp5 0 0 0 0 0
>10001c88 ------F 12349 root 0 0 0 0 0
>10001cf8 ------F 0 root 0 0 0 0 0
>10001d68 ------F 0 root 0 0 0 0 0
>10001dd8 ------F 0 root 0 0 0 0 0
>10001e48 ------F 12353 root 0 0 0 0 0
>10001eb8 ------F 12354 root 0 0 0 0 0
>10001f28 ------F 12355 root 0 0 0 0 0
>10001f98 ------F 0 root 0 0 0 0 0
>10002008 ------F 12357 root 0 0 0 0 0
>10002078 ------F 12358 root 0 0 0 0 0
>100023f8 ------- 13354 germano - 0 0 1 0 0
>10002468 ------- 657 wingz ttypd 0 0 3 73 0
>100024d8 ------- 10770 zill ttyq9 0 0 2 25 0
> 15 active, 100 total
>
>Transactions
>address flags user locks log begin isolation retrys coordinator
>10004768 A---- 10001ba8 0 0 NOTRANS 0
>1000493c A---- 10001c18 0 0 NOTRANS 0
>100057dc A---- 100024d8 2 0 NOTRANS 0
>100059b0 A---- 10002468 3 0 NOTRANS 0
>10005f2c A---- 100023f8 1 0 NOTRANS 0
> 5 active, 100 total
>
>looking at the online log:
>
>
>Mon Nov 23 00:14:33 1992
>
>00:14:33 ERROR: page cleaner # 8 has timed out
>00:15:02 Checkpoint Completed
>00:31:34 ERROR: page cleaner # 2 has timed out
>00:31:50 Checkpoint Completed
>00:48:21 ERROR: page cleaner # 3 has timed out
>00:49:01 Checkpoint Completed
>01:04:47 Checkpoint Completed
>01:20:09 ERROR: page cleaner # 4 has timed out
>01:20:25 Checkpoint Completed
>01:37:47 Checkpoint Completed>
>does anyone know what is causing this? seen this before? any fixes?
For reasons that I cannot explain (I don't really understand the thinking
myself), page cleaners are given 2 minutes to do any single write()
operation. When tbinit detects that they have not "checked back" for over
2 minutes, it considers them to be "having problems" and terminates them.
This is what you are seeing.
As to the cause: you page cleaners are timing out during checkpoints. From
this, I conclude that your checkpoints are doing a lot of work (i.e. a lot
of dirty buffers are being flushed). Since chunk writes (writes performed
by page cleaners during checkpoints) are done by sorting data by by page
number and writing the sorted pages from a large buffer in a single i/o,
it is possible that the page cleaners are doing a lot of i/o per
write(), and depending on the speed of your disk, 2 minutes may not be an
unreasonable amount of time to take. Unfortunately, this is not a tuneable
thing.
What can you do? Your best bet is to try to reduce the amount of writing
going on during checkpoints. Run tbstat -F and compare the number of
chunk writes to the sum of idle and lru writes. If chunk writes predominate,
then my guess is correct. Go into tbconfig and tune your LRU_MAX_DIRTY and
LRU_MIN_DIRTY parameters, decreasing them (the default values from 4.10 on
are 60 and 50, respectively; maintain the difference - 10 - but decrease
the absolute numbers, say to 40 and 30). This will result in more idle
and lru writing by the page cleaners, making their job during checkpoints
a bit easier. This should prevent them from timing out.
Let us know if this helps!
Dave
--
Disclaimer: These opinions are not those of Informix Software, Inc.
**************************************************************************
"I look back with some satisfaction on what an idiot I was when I was 25,
but when I do that, I'm assuming I'm no longer an idiot." - Andy Rooney