Re[4]: Long checkpoints (LRU_xxx_DIRTY/LRUS)--NETTYPE va
Posted in 1999
Thanks already for your recommendations ! Further details follow :
PHYSICAL LOG OVERFLOW :
------------------------
We never have checkpoints caused by the physical log being 75% full as
our physical log is so big (250M). Physical log overflows can happen
and can be very dangerous when the checkpoint generated at 75% of the
physical log is not completed yet by the time the physical log is 100%
filled (this can happen as threads in a critical section can continue
at the moment the 75% threshold is reached, so before the checkpoint
starts). I was last week on the "Internal Architecture and Advanced
Administration" course and this briefly described which unpleasant
things can happen if you encounter such a physical overflow situation.
SPLITUP ROOTDBS/PHYSDBS :
--------------------------
Although the rootdbs dbspace does not seem to be heavily accessed
(physdbs is), I will split these two dbspaces over two disks.
NETTYPE SETTINGS :
------------------ There is a lot of confusion about the best NETTYPE settings. Manuals
and SAP recommendations are not so clear. The statements regarding
NETTYPE settings in the article from Art Kagel are also different.
I currently run both protocols on CPUVPs as we have a lot of CPU power
on our SUN E6000, but looking at the SYSCPU% in onstat -g glo, I
noticed that 6 CPUVPs (I have 6 poll threads, 1 for shm connections
and 5 for tli) have considerably higher SYSCPU values. Maybe it is
better to have fewer tli poll threads with more connections each ?
CKPTINTVL : 1200 => 300 :
------------------------- I am really afraid of setting CKPTINTVL to 5 minutes instead of 20 as
the checkpoint duration is apparently not really influenced a lot by
the number of dirty pages. Right now, I have generated 3 checkpoints
within 2 minutes and each of these still took more than 17 seconds.
PHYSBUFF=4096 / LOGBUFF=128 :
-----------------------------
PHYSBUFF is indeed high. The manuals and course notes indeed say that
when pages/io are considerably less than 75% of buffsize, you should
decrease PHYSBUFF/LOGBUFF, but I don't know why. I know we waste some
RAM because of this, but is there maybe another reason for this rule ?
Anyway there is no reason to keep them so high; so I'll decrease them.
AFF_NPROCS/AFF_SPROCS both 0 :
------------------------------
I've never tried to use affinity although it is supported on our SUN
platform. Apparently, affinity does not stop other processes to use
these processors to which I have bound the CPU VPs (oninits). Do you
think the benifits are bigger than the drawbacks ? I could give it a
try ! Do you suggest to bind all 16 CPUVPs or only a subset of them ?
RA_PAGES=32 / RA_THRESHOLD=24 :
-------------------------------
SAP R/3 does not use sequential scans at all (although in some cases,
sequential scans would be much faster than always using an index).
The reason is that practically all indexes within SAP databases start
with the same field for which there is always an equality test in the
where-clause of selects. The annoying thing is that for most SAP
customers there is only one value for that field in all tables. So
field nunique in sysindexes is for almost all indexes 1. [I know it
is better to have many distinct values for the first field of indexes]
We do have a lot of index-data read-aheads however. Below you find
the read-ahead values in the onstat -p output (after 4 days uptime) :
ixda-RA idx-RA da-RA RA-pgsused
319609326 913755 1802453 320612727
As you can see above, most read-aheads are of type index-data (ixda)
read-aheads and the number of read-ahead pages used is also very high.
SHVIRTSIZE=1310720 :
--------------------
This is indeed a high value but even with this value Informix creates
additional virtual memory segments on a regular basis. I just noticed
that it already attached 3 additional ones since startup last sunday.
[Our database server has 5,5 GB of RAM.] The problem is that SAP
uses some kind of virtual processors too (called workprocesses) which
handle the incoming work from different users. These workprocesses
are started at startup of the application instances (on sunday) and
immediately create corresponding Informix sessions (1-1 maping). As
these SAP workprocesses keep on running until SAP is stopped (next
sunday), the Informix sessions remain the same the whole week (even
though the users log off from SAP application). I even run onmode -F
to free unused memory pages and segments every 4 hours (hoping that
this will avoid the creation of additional virtual memory segments).
[The system is being used by about 1000 users in the U.S. and Europe]
RAW DEVICES + KAIO ENABLED
--------------------------
We are indeed using raw devices with kaio. On other platforms, there
seem to be kernel parameters to use kaio (more) efficiently. I never
found any for Solaris. Does anyone know of KAIO tuning for Solaris ?
BUFFWAITS TOO HIGH :
---------------------
This seems to be another frequently asked question. My buffwaits
value also seems to be quite high. Another topic for improvement ?
Thank you in advance,
Mario Opsomer
Below you find the full onstat -p output after 4 days uptime :
INFORMIX-OnLine Version 7.24.UC3 On-Line Up 4 days 05:12:17 2979040 Kb
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
408723361 1099920114 313877840 0.00 37614362 101573940 145330116 74.12
isamtot open start read write rewrite delete commit
867326465 17840586 284739153 3525796186 15164840 5812304 8918041
1459960
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 6 1007067.78 253082.22 239 478
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
82364257 6684 3274634654 0 0 2258 2305319 277238
ixda-RA idx-RA da-RA RA-pgsused lchwaits
319609326 913755 1802453 320612727 25005902
______________________________ Reply Separator _________________________________
Subject: Re: Re[2]: Long checkpoints (LRU_xxx_DIRTY/LRUS)--NETTYPE va
Author: "Obnoxio The Clown" <obnoxio@hotmail.com> at internet
Date: 25/3/99 11:15
> Our physical log is 250 MB, which is quite big (but I want to
avoid
> physical log overflows by all means). The value for LOGBUFF is
128.
Hm. I've never heard of a _physical_ log overflow. I take it then that
you never have the physical log kick off a checkpoint? Perhaps you can
shrink that a bit. Or a lot. :-)
> PHYSBUFF is 4096. We are using SAP, so we must use unbuffered
logging.
Also a bit high in most sites