RE: large chkpoints with update program running
Posted in 2001
(DSS talking)
You should also point out that checkpoint writes are more efficient than LRU
writes. While it may give the impression that the machine is slow, if you
have a single update process (and no users you're trying to run around),
then you want checkpoint writes. Of course if you have to run users around
I know a good taxi service...
cheers
j.
> -----Original Message-----
> From: obnoxio@hotmail.com [mailto:obnoxio@hotmail.com]
> Sent: Thursday, February 15, 2001 7:51 AM
> To: informix-list@iiug.org
> Subject: Re: large chkpoints with update program running
>
>
> In the year of Our Lord Thu, 15 Feb 2001 11:47:32 -0000, "Tam
> McLaughlin"
> <tamm@scotlegal.com> spake, saying:
>
> >The system is going v v slow and it is most likely due to an update
> >progrm than has been running for over 2 hours. The following
> is the results
> >from onstat -u | grep paladm, every 10 secs which shows it
> is doing a lot of
> >writes. This has increased the checkpoint interval. Is this
> normal? Would
> >you expect the checkpoint interval to go from 0 or 1 secs
> all the way up?
> >I understand that it depends on how much data is being
> updates etc etc.
> >Just looking for clues or anything I can change/tune to
> speed things up a
> >bit.
> >
> >We are on IDS 7.30, 3 processors and Unixware 7.0.1
>
> >onstat -m> >
> >10:07:39 Checkpoint Completed: duration was 3 seconds.
> >10:17:46 Dynamically added 1 cpu VP
> >10:18:09 Checkpoint Completed: duration was 6 seconds.
> >10:22:56 Checkpoint Completed: duration was 13 seconds.
> >10:25:16 Logical Log 5136 Complete.
> >10:25:20 Logical Log 5136 - Backup Started
> >10:25:27 Logical Log 5136 - Backup Completed
> >10:26:25 Checkpoint Completed: duration was 15 seconds.
> >10:29:38 Checkpoint Completed: duration was 20 seconds.
> >10:32:46 Checkpoint Completed: duration was 16 seconds.
> >10:35:43 Checkpoint Completed: duration was 12 seconds.
> >10:40:24 Logical Log 5137 - Backup Started
> >10:40:24 Logical Log 5137 Complete.
> >10:40:33 Logical Log 5137 - Backup Completed
> >10:45:54 Checkpoint Completed: duration was 7 seconds.> >
> >
> >onstat -p> >
> >
> >Informix Dynamic Server Version 7.30.UC10X2 -- On-Line -- Up 9 days
> >20:17:39 -- 437432 Kbytes
> >
> >Profile
> >dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> >150524382 14181166 3479905643 95.67 3995912 3283893
> 92954833 95.70
> >
> >isamtot open start read write rewrite delete commit
> >rollbk
> >2137261542 35594822 133115757 1587923637 42933432 1135739
> 223641 331363
> >25597
> >
> >gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> >0 0 0 0 0 0 0
> >
> >ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> >6 0 0 241667.47 150249.09 1143 2838
> >
> >bufwaits lokwaits lockreqs deadlks dltouts ckpwaits
> compress seqscans
> >9731674 79 2044529143 4 0 604
> 215290 2455026
> >
> >ixda-RA idx-RA da-RA RA-pgsused lchwaits
> >65678222 72603 49220912 114938522 561723
> >
> >onstat -c (the parameters that are relevant)> >
> >Informix Dynamic Server Version 7.30.UC10X2 -- On-Line -- Up 9 days
> >20:18:16 -- 437432 Kbytes
> >
> >PHYSDBS physspace
> >PHYSFILE 10000>
> Way too small. You're checkpointing more quickly than every
> 10 minutes because
> your PHYSLOG is too small. Based on what I can see in the
> message log, you need
> at least 50000.
>
> >LOGFILES 30
> >LOGSIZE 5000
> >TBLSPACE_STATS 0
> >SERVERNUM 0
> >DBSERVERNAME pluto
> >DBSERVERALIASES pluto_net
> >NETTYPE ipcshm,2,64,CPU
> >NETTYPE tlitcp,2,8,NET>
> This isn't germane, but these are pretty daft settings. Try:
> NETTYPE ipcshm,2,200,CPU
> NETTYPE tlitcp,2,200,NET
>
> >DEADLOCK_TIMEOUT 60
> >RESIDENT 1
> >MULTIPROCESSOR 1
> >NUMCPUVPS 2
> >SINGLE_CPU_VP 0
> >LOCKS 20000>
> I generally have more locks than buffers, this seems odd. I'd
> go for at least
> one lock per buffer.
>
> >BUFFERS 150000
> >NUMAIOVPS 40>
> I take it you don't have KAIO? If you do, use it, it makes a
> big difference.
>
> >PHYSBUFF 124
> >LOGBUFF 22> >LOGSMAX 60
> >CLEANERS 20
> >SHMVIRTSIZE 120000
> >SHMADD 8000
> >SHMTOTAL 0
> >CKPTINTVL 600
> >LRUS 93>
> Eh? 93?
>
> >LRU_MAX_DIRTY 4
> >LRU_MIN_DIRTY 2
> >LTXHWM 40
> >LTXEHWM 45
> >RA_PAGES 48
> >RA_THRESHOLD 42
> >DBSPACETEMP tempspace:tempspace2
> >OPTCOMPIND 0>
> There are basically two approaches you can take:
>
> 1. reduce individual checkpoint time
> * DECREASE LRU_MIN/MAX_DIRTY (0/1), PHYSFILE and CKPTINTVL
>
> 2. reduce total checkpoint time
> * INCREASE LRU_MIN/MAX_DIRTY (80/90), PHYSFILE and CKPTINTVL
>
> Which way you go depends on you. Maybe you can also up LRUS
> and CLEANERS to 127
> and monitor onstat -u to see how many cleaners are used --
> maybe 20 isn't
> enough.
>