Re: How can I tell what's eating my global memory pool?
Posted in 2003
squung@aol.com (Everett Mills) wrote in message news:<b53eea08.0308120717.47d798f3@posting.google.com>...
Addenda: I just read this in another forum, sounds like the same
problem...
From: Alexey Sonkin [mailto:alexeis@grandvirtual.com]
Sent: Thursday, August 14, 2003 12:42 PM
To: ids@iiug.org
Subject: Severe memory leak in global pool: mgm_query [1700]
Hi, everybody,
After taking a look at the size of memory pools (onstat -g mem) on our
machine, I was shocked : it appeared, that the size of global pool is
138 MB!!!
Informix Dynamic Server 2000 Version 9.21.UC4 -- On-Line (Prim) --
Up 22
days 14:11:09 -- 1790448 Kbytes
Pool Summary:
name class addr totalsize freesize #allocfrag #freefrag
........
global V 8f816020 134889472 1006776 518347 6502
........
This global pool can't be reduced by 'onmode -F'
(its 'free' size is too small)
I've made an attempt to analyze this pool with
'onstat -g afr global'. The resulting file was 16 MB in size mostly
contained records like:
Allocations for pool name global:
addr size memid
....
8f8199d8 240 mgm_query
8f819ac8 240 mgm_query
....
This is total statistics of that file:
Total_bytes #fragmets fragment_name
-------------------------------------------
40 1 messages
48 1 fragman
56 1 initseg
56 1 shmblklist
160 4 Log
280 1 filetable
280 1 opentable
280 1 hashfiletab
456 1 mgm
568 1 gentcb
1496 1 SAPI
1832 3 overhead
2592 1 ostcb
3648 8 rsam
4528 11 sql
6392 107 misc
7432 163 osenv
19256 310 mcbmsg
41024 2 btcleaner
73528 69 scb
274944 2 aio
559144 139 log
3147248 13 datarep
5822416 3060 net
122795024 509837 mgm_query
------------------------------------------
That is, this memory leak is caused by MGM_QUERY fragments:
We have 509837 fragments 240 bytes each, 122 MB in total !!!!
I have a suspicion, that 'mgm' here stays for 'memory grant manager',
that is, memory leak is caused by PDQ (parallel data query) queries!
I've attempted to make a search in the PDQ (product defect query)
tool, and was unable to find anything similar.
By making further experiments with PDQ queries and running simple
script in parallel (onstat -g afr global | grep mgm_query | wc -l),
I've found, that each PDQ query allocates from 4 to 1000 mgm_query
fragments and doesn't release them to the system upon completion!!!
Hey, Informix tech-support, can You file a bug in Your 'Atlas', or
whatever You use now at IBM?
------------------------------------------
Alexey Sonkin
Senior Database Administrator
Very similar, except mine's filled up with 43000 + net entries, and
this server is not being access thought IP except at few (less than a
dozen) times per day.
> Guys-
> We have a pair of (supposed to be) identical IDS servers running
> in two of our plants. They have similar software, identical hardware,
> etc. They are both running IDS 7.31.FD4 on HP-UX 11.0. The less busy
> one, however has a global memory pool that grows until it chews up all
> of the memory allocated to informix.
>
> Here's what the memory usage on the sick server looks like:
> Informix Dynamic Server Version 7.31.FD4 -- On-Line (Prim) -- Up
> 13 days 17:
> 57:44 -- 317504 Kbytes
>
> Segment Summary:
> id key addr size ovhd class
> blkused bl
> kfree
> 4100 1381451777 c00000000023c000 156409856 36048 R*
> 19084 9
>
> 1543 1381451780 c000000009766000 167772160 3200 V*
> 7526 12
> 954
> 7178 1381451783 c000000013766000 942080 656 M
> 108 7
>
> Total: - - 325124096 - -
> 26718 12
> 970
>
> Here's the well one:
>
> Informix Dynamic Server Version 7.31.FD4 -- On-Line (Prim) -- Up
> 77 days 16:
> 18:11 -- 317504 Kbytes
>
> Segment Summary:
> id key addr size ovhd class
> blkused bl
> kfree
> 6660 1381451777 c00000000023c000 156409856 36048 R*
> 19084 9
>
> 7175 1381451780 c000000009766000 167772160 3200 V*
> 4059 16
> 421
> 1546 1381451783 c000000013766000 942080 656 M
> 108 7
>
> Total: - - 325124096 - -
> 23251 16
> 437
>
> (* segment locked in memory)
>
> Note that the less busy server has in use 7526 blocks of virtual
> segment memory vs. 4059 in the good server (snapshot at peak usage).
> Tomorrow the bad one will probably jump to about 8500, while the other
> continues at 3500-4000. I recently discovered that the global memory
> pool is where this is going. Here's an example (onstat -g mem | grep
> global over a few days):
>
> Sick Server:
> Pool Summary:
> name class addr totalsize freesize #allocfrag
> #freefrag
> 8/8/03
> global V c00000000976e028 35987456 4511120 28294
> 13396
>
> 8/11/03
> global V c00000000976e028 44130304 5764160 35748
> 17361
>
> 8/12/03
> global V c00000000976e028 46628864 6142168 38397
> 18653
>
> Good Server:
> Pool Summary:
> name class addr totalsize freesize #allocfrag
> #freefrag
> 8/8/03
> global V c00000000976e028 11026432 689096 3492 558
>
> 8/11/03
> global V c00000000976e028 11034624 746872 3304 548
>
> 8/12/03
> global V c00000000976e028 11132928 688456 3711 498
>
> Note how the sick server's global pool grows daily. The question is:
> How can I track down what's eating that up. The Docs only references
> to it are vague at best.
>
> --EEM