Re: page cleaner #1 timed out
Posted in 1996
> From informix-list-owner@rmy.emory.edu Tue Apr 30 03:41:00 1996 > From: Guy Germonpre <germonpre_guy/dis@bekaert.com> > Subject: page cleaner #1 timed out > Date: 30 Apr 1996 07:43:59 GMT guy, > Online 5.03.UD2 on Compaq proliant 4500. From time to time we get the > error > "ERROR: page cleaner # 1 has timed out" > > What does it mean ? How can we solve it ? it means that the page cleaner process was waiting to do a write to a device, and it exceeded the max wait time that it has set internally. so it sat waiting to write to the disk, and never got the chance. it then assumes that something is wrong with the device, or writing to the device, and times out. if all of your devices are ok, this usually means that there is a lot of contention for the device, and there was a pretty good size queue waiting to access the device, and the page cleaner process was way down in the queue. there could be a lot of reasons for this, but the first would have to do with the number of disks and controllers you have, and how your dbspaces are spread out over those disks. if you have only one or a few disks, and/or all of your hot dbspaces are on one disk or even one controller, then this could lead to a lot of contention for that disk or controller, and lead to situations like this. the manaual states that having a page cleaner for every physical device is a good starting point, but that is also in conjunction with the number of lru queues that you have. there is a little bit of fine tuning here; the number of page cleaners, the number of lru queues, and the min and max percentage that the queues are allowed to fill before they are flushed to disk. look at the tbstat -F command. it depends on your situation, but generally you should at least have few or no idle writes. you then need to balance the lru writes to the chunk writes. lru writes are less efficient, and are done when the page cleaners flush the lru queues to disk. chunk writes are more efficient, and are done during a checkpoint. lru writes don't impact the performance of the system as much, whereas all operations are suspended during a checkpoint, so to the users, the system will appear to freeze up during the checkpoint, whether it lasts for 1 second or 1409 seconds, (i had a checkpoint this long when i first brought up my 7.1 system!). it might be ok for the system to freeze every 10 or 20 minutes for 5 seconds during a checkpoint, or that might be unacceptable to your users. from this you need to balance the number of lru queues, and how full they become before they are flushed. the lower the min and max tbconfig parameters are set, the more often they will be flushed, (more lru writes), and the less work there will be to do at checkpoint. this could also be part of your problem. if you have a couple of page cleaners all going to the same device, one could easily time out during a checkpoint, on the other hand, if you have the same situation, one could time out during an lru write. it is a bit tricky, but you should see the results from playing around with these parameters. the best bet to start out with is users/2 or 8, whichever is larger, for your lrus, one page cleaner for every physical device, and i would set the min and max parameters a little low to start out with, maybe 30 and 20 to start. take it from there. this has become a lot longer winded than i had planned, i hope it helps. of course, once again, nothing replaces rtfm. mickm