Re: Digital UNIX V4.0F + 7.31.FD1 Performance
Posted in 2007
A Digital UNIX 4.0F site running IDS 7.31.FD1 (BaaN IV) saw performance degrade steadily over days until the server was bounced. NOAGE was already set, and onstat showed nothing obviously wrong. Tilman identified the known 7.3x LRU buffer-priority issue: onstat -P showed ~87% btree pages and onstat -R many MED_HIGH/HIGH buffers, so index pages were squeezing out data pages. Setting the undocumented environment variable LRUAGE=1 (exported before starting oninit, not in ONCONFIG, as Frank noted) rebalanced the cache (data ~57%) and stopped the degradation. A later unrelated crash (out of virtual shared memory) was attributed to memory/kernel shared-memory limits or a runaway session, with an upgrade recommended.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration
Do you have NOAGE turned on in your ONCONFIG file? Many UNIX releases are very
aggressive about aging long-running processes. Since IDS's oninit processes
NEVER go down they are classified as long-running and have their run-priority
and timeslices adjusted accordingly unless you have NOAGE set to '1' in your
ONCONFIG file.
Art S. Kagel
----- Original Message -----
From: Robert Langhammer <ids@iiug.org>
At: 6/26 10:46:02
Hello!
I have a real strange problem.
We use this system since many years and with no problems.
Since last time (i think around 2-3 mounth) it is need to boot the server on
weekend because the performance of the system go realy down!
And in the last 3 weeks it is much more, restarts need every 3 days.
The system is running after the restart with best performance and no problems
but after time the system performance goes down and down.
We do all what is need in the Tuning guide, no onstat parameter shows strange
constellations that shows us this and this parameter is bad or wrong.
also many and needes kernel parameters standing on maximum.
funny is that with the years this problem comes more and more and at moment i
dont know what to do.
we plan on august a archiving to other database, but in this time our company
need realy a faster machine we have at moment with this performancesituation
realy big problem with our BaaN IV ;-) system.
Please help.
Greeting, Robert
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
yesyes, i have
NOAGE 1 # Process aging
D(*^ that's the only thing I can think of that bouncing IDS would tend to fix.
Art S. Kagel
----- Original Message -----
From: Robert Langhammer <ids@iiug.org>
At: 6/26 11:02:12
yesyes, i have
NOAGE 1 # Process aging
*******************************************************************************
Forum Note: Use "Reply" to post a response in the discussion forum.
perhaps that need....
onstat -p when system going slow and slow
Informix Dynamic Server Version 7.31.FD1 -- On-Line -- Up 20:38:13 -- 1810432
Kbytes
Profile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
31740042 54676124 4514343286 99.30 1025297 2061142 7319399 85.99
isamtot open start read write rewrite delete commit rollbk
4365661745 10430340 100229604 3763751492 1068912 334275 569623 389143 119
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 65100.68 6000.37 123 246
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
3422826 80 1305292289 0 0 329 125060 220366
ixda-RA idx-RA da-RA RA-pgsused lchwaits
8234629 5386553 5342889 18698297 210046
Check you ps list and make sure there is not a hung program (including a
dbaccess or isql process) that is hogging your system resources.
On 6/26/07 11:23 AM, "ROBERT LANGHAMMER" <r.langhammer@roco.cc> wrote:
> perhaps that need....
>
> onstat -p when system going slow and slow>
> Informix Dynamic Server Version 7.31.FD1 -- On-Line -- Up 20:38:13 -- 1810432
> Kbytes
>
> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 31740042 54676124 4514343286 99.30 1025297 2061142 7319399 85.99
>
> isamtot open start read write rewrite delete commit rollbk
> 4365661745 10430340 100229604 3763751492 1068912 334275 569623 389143 119
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 65100.68 6000.37 123 246
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 3422826 80 1305292289 0 0 329 125060 220366
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 8234629 5386553 5342889 18698297 210046
>
>
>
******************************************************************************>
*
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
when i use: ps -eo pcpu,pid,lstart,cputime,logname,comm|sort|more
Output is like:
0,0 17475 Mi 27 Jun 08:07:51 2007 0:00.79 informix cpuinfo
0,0 17515 Mi 27 Jun 08:12:42 2007 0:00.02 telnetd
0,0 17516 Mi 27 Jun 08:12:42 2007 0:00.03 informix ksh
0,0 17703 Mi 27 Jun 08:35:14 2007 0:00.03 langh ps
0,0 17704 Mi 27 Jun 08:35:14 2007 0:00.01 langh sort
0,0 17706 Mi 27 Jun 08:35:04 2007 0:00.05 sjohann onstat
%CPU PID STARTED TIME LOGNAME COMMAND
22,0 10281 Di 26 Jun 13:31:05 2007 03:58:52 informix oninit
38,7 10293 Di 26 Jun 13:31:02 2007 03:48:38 informix oninit
48,0 10294 Di 26 Jun 13:31:05 2007 06:27:10 informix oninit
I see our system use maximum CPU resources.
okay we have tables that have many data inside 2.135.000 and one with more
7.000.000
Funny is only that this tables get every weekend when some cleaningjobs
running ned update statistics low/medium/high and after restart very very fast.
but alloversystemperformance and special in this tables goes now special for
this big tables down and down.....
after restart, without any update statistics again very fast.....
:-(
There is an issue in 7.3x BUFFER management which may cause performance
problems.
In 7.3x Informix introduced a new concpet of priorisation of buffers according
to the page type on top of the old 'LRU' concept. Within in this concept Index
branch pages have a higher proiority than data pages. This may leads to a
situation where nearly all buffers are occupied by INDEX pages, and very
little are left over for data pages - leading to a very small effective size
of the BUFFER. Now when users attempt to read new data, they constantly swap
out each others data pages. This situation usually builds up over time and is
more likely in dtabases with many, large, broad indexes .
You may use 'onstat -P' (capital P) to get an inidcation whether that's your
probelm, if the summry shows a very high percentage of btree pages, you
probably hit this problem.
Another hint is, when 'onstat -R' reports high number of MED_HIGH and HIGH
Buffers.
You may restore the old behaviour by setting the undocumented env variable
NOLRUPRIO to 1 before starting IDS, however, I recall there used to be
problems with that as well which got fixed only in later IDS version 7.31.FD7
- PTS 161710 .
Alternatively you may set the equally undocumented variable LRUAGE, which
causes the priority of all buffers to get downgraded periodically and also
avoid that issue. Set either of the two variables to 1 (but not both at the
same time) before you start oninit.
Please let us know whether that helped.
Apart from that 7.31.FD1 is really old and has some serious bugs - I suggest
to contact IBM Informix support and get a newver version.
(not sure what the latest 7.31.FDx is for True64)
Rgds
Tilman
Percentages:
Data 12.89
Btree 86.78
Other 0.33
it seems to me that is this problem what you specified.....
also:
informix1:/home/informix/langh/topdb->onstat -R
Informix Dynamic Server Version 7.31.FD1 -- On-Line -- Up 21:35:58 -- 1810432
Kbytes
8 buffer LRU queue pairs priority levels
# f/m pair total % of length LOW MED_LOW MED_HIGH HIGH
0 F 36068 98.5% 35526 0 11705 22406 1415
1 m 1.5% 542 0 539 3 0
2 f 38827 99.2% 38528 0 15603 21456 1469
3 m 0.8% 299 0 298 1 0
4 f 37752 99.1% 37426 0 14441 21571 1414
5 m 0.9% 326 0 322 4 0
6 f 37752 99.3% 37478 0 14130 21838 1510
7 m 0.7% 274 0 272 2 0
8 f 37019 99.1% 36691 0 13536 21711 1444
9 m 0.9% 328 1 327 0 0
10 f 37402 99.1% 37083 0 13771 21884 1428
11 m 0.9% 319 0 318 1 0
12 f 37054 99.2% 36750 0 13625 21749 1376
13 m 0.8% 304 0 302 2 0
14 f 37298 99.2% 36991 0 13777 21691 1524
15 m 0.8% 307 0 305 2 0
2699 dirty, 299172 queued, 300000 total, 524288 hash buckets, 2048 buffer size
start clean at 5% (of pair total) dirty, or 1875 buffs dirty, stop at 2%
0 priority downgrades, 0 priority upgrades
So i try first to change this settings what you write and hope it works after.
We have an old Baan IV and it is not easy put new Databaseversions.
It needs often new OS Version, new portingsets and after upgrade databe.
Till some month our ERP-System run good and there is one SAYING: "never change
a running system"
now it is time to think about.....
I know since yesterday from IBM Passport Advantage Site one document where
standing all Softwareproducts and with what system there are working.
In this document standing 7_31_FD7 is working and compiled with True 64 unix
V4.0D we have 4.0F it should be running......
So i think if i take this step and hope so solve this problem.
First i must check if our Networkermodule (V5.5 also older version) is working
with them together.......
Yes, sounds as if this could indeed be your problem - I suggest to set LRUAGE
(unless you are able to upgrade to 7.31.FD7 or higher)
LRUAGE is set per default in all SAP systems running on IDS 7.31.xD .
I don't recall issues with it in 7.31.xD1 and higher.
If performance still degrades afterwards, we'd have to look further .
Regarding Legatao -to my knowledge - the module is not tight to the fixpack
version, but rather the release.
i.e. as long as you stay on 7.31.FDx you should be ok. At least there
shouldn't be any differences in onbar that affect the Legato module.
But of course I don't know taht for sure, so better check with Legato EMC.
Please let me know in case you get a negative answer from them.
Regards
Tilman
Thank you, today on evening we make this changes in $ONCONFIG, restart the machine and hope it works! Greetings from Austria
ROBERT LANGHAMMER wrote: > Thank you, today on evening we make this changes in $ONCONFIG, restart the > machine and hope it works! > > Greetings from Austria > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > > LRUAGE has to be set as an environment variable before starting the server. It's not an variable to set in $ONCONFIG. If you interactively log in execute the following two lines: LRUAGE=1 export LRUAGE or put this into informix's .profile. If you have a startup script executed on server reboot put these lines also in there.
so, ein erstes Feedback!
Werte bei onstat -P ergeben folgendes:
Percentages:
Data 56.87
Btree 40.33
Other 2.80
onstat -R
8 buffer LRU queue pairs priority levels
# f/m pair total % of length LOW MED_LOW MED_HIGH HIGH
0 F 35745 97.9% 35000 0 34996 4 0
1 m 2.1% 745 0 743 2 0
2 f 43189 98.4% 42485 0 42466 19 0
3 m 1.6% 704 0 702 2 0
4 f 36950 97.9% 36185 0 36177 8 0
5 m 2.1% 765 0 765 0 0
6 f 36689 97.9% 35908 0 35898 10 0
7 m 2.1% 781 0 777 4 0
8 f 36618 97.9% 35834 0 35825 10 0
9 m 2.1% 784 0 781 3 0
10 f 36476 97.8% 35688 0 35675 14 0
11 m 2.2% 788 0 783 5 0
12 f 36475 97.9% 35724 0 35722 2 0
13 m 2.1% 751 0 750 1 0
14 f 36623 97.8% 35828 1 35808 19 0
15 m 2.2% 795 0 793 2 0
6113 dirty, 298765 queued, 300000 total, 524288 hash buckets, 2048 buffer size
start clean at 5% (of pair total) dirty, or 1875 buffs dirty, stop at 2%
15778194 priority downgrades, 0 priority upgrades
Was mir vorallem auch noch auffällt das jetzt die CPUs wirklich arbeiten und
nicht so wie in den letzten Wochen sich bei 1/3 Auslastung dahin langweilen.
Derzeit kann ich noch keine "Verlangsamung" der Maschine feststellen, werde
dies aber intensiv die nächten Tage beobachten 8eventuell keinen Restart am
Wochenende) um dann hier wieder ein Feedback geben.
So, erstes Ergebnis ist da! Datenbank läuft jetzt ohne der oben im Thread angesprochenen Probleme. Das heißt sie läuft an den folgenden Tagen genauso schnell (oder langsam) wie kurz nach dem Neustart. Heute ist aber ein Fehler aufgetreten und dann ist uns die DB weggerauscht, hier der Auszug aus dem online.log: 13:09:46 Assert Failed: No Exception Handler 13:09:46 Informix Dynamic Server Version 7.31.FD1 13:09:46 Who: Session(25410, bogenspd@<UNKNOWN-HOST>, 11675, 1562137552) Thread(27323, sqlexec, 25810ca30, 3) File: mtex.c Line: 446 13:09:46 Results: Exception Caught. Type: MT_EX_BE, Context: mt_ex_setup: no mem 13:09:46 Action: Please notify Informix Technical Support. 13:09:46 shmat: [ENOMEM][12]: out of available data space, check system MAXMEM 13:09:46 out of virtual shared memory 13:09:47 shmat: [ENOMEM][12]: out of available data space, check system MAXMEM 13:09:47 out of virtual shared memory 13:10:33 Logical Log 340542 Complete. 13:10:52 See Also: /tmp/af.6ea37ff9 13:10:52 mtex.c, line 446, thread 27323, proc id 798, No Exception Handler. 13:10:52 PANIC: Attempting to bring system down 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 Zu erwähnen ist, der Schalter MAXMEM ist auf den maximal höchten Wert eingestellt den das OS zuläßt! M.f.G.
This looks as if your IDS instance ran out of memeory . Maybe some user session ran out of control and grabbed all the memory or there was exceptional load ? Do you have any idea whether some special jobs where running. Look into the af-file it might giuve you a hint. Contact IBM Informix support to help you indentifiying the root cause. Could also be some memory leak in IDS I do think upgrading the instance to a later IDS version is a good idea. 7.31.FD10 is the latest for True64 if I am correct, but you better check with IBM support. However, unless you identified the root cause for the memory outage you won't know whether upgrade does fix your problem. HTH Tilman P.S: Again, please post in english in this forum
Looks like your OS isn't configured with enough shared memory per process or number of shared memory segments per process. You need to adjust kernel settings upwards. Art S. Kagel ----- Original Message ----- From: Robert Langhammer <ids@iiug.org> At: 7/04 7:29:09 So, erstes Ergebnis ist da! Datenbank läuft jetzt ohne der oben im Thread angesprochenen Probleme. Das hei t sie läuft an den folgenden Tagen genauso schnell (oder langsam) wie kurz nach dem Neustart. Heute ist aber ein Fehler aufgetreten und dann ist uns die DB weggerauscht, hier der Auszug aus dem online.log: 13:09:46 Assert Failed: No Exception Handler 13:09:46 Informix Dynamic Server Version 7.31.FD1 13:09:46 Who: Session(25410, bogenspd@<UNKNOWN-HOST>, 11675, 1562137552) Thread(27323, sqlexec, 25810ca30, 3) File: mtex.c Line: 446 13:09:46 Results: Exception Caught. Type: MT_EX_BE, Context: mt_ex_setup: no mem 13:09:46 Action: Please notify Informix Technical Support. 13:09:46 shmat: [ENOMEM][12]: out of available data space, check system MAXMEM 13:09:46 out of virtual shared memory 13:09:47 shmat: [ENOMEM][12]: out of available data space, check system MAXMEM 13:09:47 out of virtual shared memory 13:10:33 Logical Log 340542 Complete. 13:10:52 See Also: /tmp/af.6ea37ff9 13:10:52 mtex.c, line 446, thread 27323, proc id 798, No Exception Handler. 13:10:52 PANIC: Attempting to bring system down 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 13:10:52 semctl: errno = 22 Zu erwähnen ist, der Schalter MAXMEM ist auf den maximal höchten Wert eingestellt den das OS zulä t! M.f.G. ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.