IDS hangs; IBM diagnoses resident memory problem
Posted in 2009
An IDS 10.0FC8X6 instance on AIX 5.3 (p570, 32GB, ~20GB of shared memory with RESIDENT=-1) hung with "Blocked by Checkpoint", forcing a reboot and 52 minutes of downtime. IBM's AIX team blamed exceeding the 80% pinned/resident memory limit and suggested more RAM, a smaller IDS footprint, or non-resident operation. List members disputed that: some argued RESIDENT should rarely be used, others defended it, while several pointed at the online log's "KAIO: out of OS resources (errno 12)" and the HDR ping timeout, suggesting the real cause was an inability to flush to disk / KAIO resource exhaustion (asking about IFMX_AIXKAIO_NUM_REQ) or a known AIX getuserattr memory leak (APAR IY91677, already patched). No definitive root cause or resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Server Administration, Logging & Checkpoints, Networking & sqlhosts Configuration, Platform-Specific Issues, Internationalization & Character Sets, Versions, Editions & End-of-Life
IDS 10.0FC8X6 on AIX 5.3
Our Informix instance simply packed up at a low-activity that is, it just
hanged, with a "Bloked by Checkpoint" message.
We took an onstat -a and AIX dump, the latter re-booting the server, after
which everything came up fine. But we had a 52-minute downtime on a 24x7
system.
PMRs were then opened with IBM Informix and AIX support.
The AIX team reckon that the server was or tried to use more than the 80%
AIX limit on "pinned" (their word for resident I think) memory and recommend
either:
- buying mroe memory
- reducing Informix's footprint
- making Informix run non-resident
I don't think more memory is necessary.
What we did think we'd do is change RESIDENT from -1 to 1, so that the
virtual portion is non-resident. That would have avoided this problem it
seems, but this doesn't fit with current IBM orthodoxy I think.
The IBM Informix technican, uderstadndably cautious, was prepared to say
only "I don't believe from a IDS point of view you will see any impact on
performance, unless you start to see swapping"
I need to make a recommendation to the customer who is a little annoyed that
"IBM" (as he perceives the two different teams who looked at the problem)
can't idenify the issue so how can we avoid it happening again?
Any thoughts on the RESIDENT setting (or anything else)?
thanks
Neil
A few facts:
Server is an IBM p570 with 32g memory.
Current Informix memory usage is:
IBM Informix Dynamic Server Version 10.00.FC8X6 -- On-Line (Prim) -- Up 6
days 21:11:49 -- 20788800 Kbytes
$ onstat -c | egrep -i "buff|shmvirt"SHMVIRTSIZE 10274660 # initial virtual shared memory segment size
BUFFERPOOL
size=4K,buffers=2500000,lrus=512,lru_min_dirty=0.10,lru_max_dirty=0.20
Server had been up for 342 days and Informix for 189.
Online log:
Thu Jan 8 00:00:36 2009
00:00:36 1800 buffers dirty
00:00:36 oldest lsn loguniq 97192, logpos 0x1211018
00:00:36 1742 dirty pages are to be flushed
00:00:37 dskflush() took 0 seconds
00:00:37 wait4critex() took 0 seconds
00:00:37 63 buffers dirty
00:00:37 oldest lsn loguniq 97192, logpos 0x1211018
00:00:37 safe_dskflush() took 0 seconds
00:00:37 Checkpoint Completed: duration was 0 seconds.
00:00:37 Checkpoint loguniq 97192, logpos 0x1f35018, timestamp: 0x17dca0ed
00:00:37 Maximum server connections 923
00:00:37 Buffer manager: starting coarse downgrades.
00:00:37 Buffer manager: finished coarse downgrades.
Starting threshold adjustments.
00:00:37 Buffer manager: finished threshold adjustments.
00:15:37 1033 buffers dirty
00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
00:15:37 837 dirty pages are to be flushed
00:15:37 dskflush() took 0 seconds
00:15:37 wait4critex() took 0 seconds
00:15:37 196 buffers dirty
00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
00:15:37 safe_dskflush() took 0 seconds
00:15:37 Checkpoint Completed: duration was 0 seconds.
00:15:37 Checkpoint loguniq 97192, logpos 0x2cdf018, timestamp: 0x17e18f4e
00:15:37 Maximum server connections 923
00:15:37 Buffer manager: starting coarse downgrades.
00:15:37 Buffer manager: finished coarse downgrades.
Starting threshold adjustments.
00:15:37 Buffer manager: finished threshold adjustments.
00:23:33 Logical Log 97192 Complete, timestamp: 0x17e4135a.
00:23:33 Logical Log 97192 - Backup Started
00:23:34 KAIO: out of OS resources, errno = 12, pid = 53534
00:23:40 Logical Log 97192 - Backup Completed
00:25:36 DR: ping timeout
00:25:36 DR: Receive error
00:25:36 ASF Echo-Thread Server: asfcode = -25582: oserr = 4: errstr = :Network connection is broken.
System error = 4.
00:30:37 368 buffers dirty
00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
00:30:37 347 dirty pages are to be flushed
00:30:37 dskflush() took 0 seconds
00:30:37 wait4critex() took 0 seconds
00:30:37 21 buffers dirty
00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
00:30:37 safe_dskflush() took 0 seconds
01:28:27 IBM Informix Dynamic Server Started.
01:28:27 WARNING: If you intend to use J/Foundation or GLS for Unicodefeature(GLU) with this Server instance, please make sure that your SHMBASE
value specifies in onconfig is 0x700000010000000 o
r above. Otherwise you will have problems while attaching or dynamimically
adding virtual shared memory segments. Please refer to Server machine notes
for more information.
01:28:47 Segment locked: addr=700000000000000, size=10765910016
01:28:47 Requested shared memory segment size rounded from 10274660KB to10274672KB
01:29:00 IBM Informix Dynamic Server Started.
> From: neil.truby@ardenta.com > Subject: IDS hangs; IBM diagnoses resident memory problem > Date: Wed, 14 Jan 2009 22:44:05 +0000 > To: informix-list@iiug.org > > > Any thoughts on the RESIDENT setting (or anything else)? > > thanks > Neil > No shit. Boom! Thank you Neil for living through the one scenario which why I have always said don't use the RESIDENT flag. Period. (Actually sorry you had to face this issue.) This was the dumbest setting and truthfully serves no real purpose for 90+% of the instances of IDS running today. It makes less sense today when you have Unix kernels which are 'self tuning' prior to hitting run level 5 which is before Informix is started. In simple terms, when you add a 'memory resident application' you're actually taking away memory on a permanent basis that the server thought it has to work with. Machines don't swap as much today as they did back in the early 90s. It used to be that you had a 2 to 1 swap to active memory ratio. Now you have what? 2 GB of swap? Oh well, I've got a doggie crying to go out in 10 degree F weather not including the nice wind chill. Later. -G _________________________________________________________________ Windows Live™: Keep your life in sync. http://windowslive.com/howitworks?ocid=TXT_TAGLM_WL_t1_allup_howitworks_012009
Hello Ian, i do not agree with you on that, have seen situations where setting resident to 1 or -1 or -2 made performance from not existing to really good; specially on systems with a lot of filesystem io not all systems can limit the amount of cache which will be allocated for filesystem cache. This can kill database performance. Superboer. On 15 jan, 00:27, Ian Michael Gumby <im_gu...@hotmail.com> wrote: > > From: neil.tr...@ardenta.com > > Subject: IDS hangs; IBM diagnoses resident memory problem > > Date: Wed, 14 Jan 2009 22:44:05 +0000 > > To: informix-l...@iiug.org > > > Any thoughts on the RESIDENT setting (or anything else)? > > > thanks > > Neil > > No shit. > Boom! > Thank you Neil for living through the one scenario which why I have always said don't use the RESIDENT flag. Period. > (Actually sorry you had to face this issue.) > > This was the dumbest setting and truthfully serves no real purpose for 90+% of the instances of IDS running today. > > It makes less sense today when you have Unix kernels which are 'self tuning' prior to hitting run level 5 which is before Informix is started. > In simple terms, when you add a 'memory resident application' you're actually taking away memory on a permanent basis that the server thought it has to work with. > > Machines don't swap as much today as they did back in the early 90s. It used to be that you had a 2 to 1 swap to active memory ratio. Now you have what? 2 GB of swap? > > Oh well, I've got a doggie crying to go out in 10 degree F weather not including the nice wind chill. > Later. > > -G > > _________________________________________________________________ > Windows Live™: Keep your life in sync.http://windowslive.com/howitworks?ocid=TXT_TAGLM_WL_t1_allup_howitwor...
for the problem, you could get some stack traces from running threads... someone must be able to tell what is going on... maybe a bad disk or controller???? i would ask IBM/informix what diags they need. -->>The AIX team reckon that the server was or tried to use more than the 80% ->>AIX limit on "pinned" (their word for resident I think) memory and recommend hate to say this; but this sounds like bullshit to me; you could do a test on a test box and see what happens if you use 80% or more of memory resident. if it does hang you can still ask why to informix and AIX. this sounds afwull; a lame excuse for a bug. Superboer.
2009/1/15 <superboer7@t-online.de>: > for the problem, you could get some stack traces from running > threads... > someone must be able to tell what is going on... > maybe a bad disk or controller???? > i would ask IBM/informix what diags they need. > > -->>The AIX team reckon that the server was or tried to use more than > the 80% > ->>AIX limit on "pinned" (their word for resident I think) memory and > recommend > hate to say this; but this sounds like bullshit to me; you could do a > test > on a test box and see what happens if you use 80% or more of memory > resident. > if it does hang you can still ask why to informix and AIX. this sounds > afwull; > a lame excuse for a bug. > > Superboer. > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list > Neil Forget memory/resident. Your problem is squarely down to an inability (at that particular time) to be able to write to disk quickly enough. Hit the same problem on similar hardware about 12 months ago. Identified that as a single query doing a fairly large count, sum, group by, order by query (for a year rather than a month) and thrashing the KAIO subsystem on the temp dbspaces. IDS Support identified disk write issues and thought it was a disk down. P Series and Disk (two separate hardware calls :-(( ) couldn't find any problems so we restarted IDS (needed to crash it as it was stuck) and all was well and has been well since. Keith
"Keith Simmons" <smiley73@googlemail.com> wrote in message news:mailman.142.1232013875.1831.informix-list@iiug.org... > Neil > > Forget memory/resident. Your problem is squarely down to an inability > (at that particular time) to be able to write to disk quickly enough. > Hit the same problem on similar hardware about 12 months ago. > Identified that as a single query doing a fairly large count, sum, > group by, order by query (for a year rather than a month) and > thrashing the KAIO subsystem on the temp dbspaces. IDS Support > identified disk write issues and thought it was a disk down. P Series > and Disk (two separate hardware calls :-(( ) couldn't find any > problems so we restarted IDS (needed to crash it as it was stuck) and > all was well and has been well since. Do you have PMR numbers by any chance?
Never had an issue with " RESIDENT " but then again currently we are not
using it. Too many things sharing the server.
Are you sure you aren't running into a memory leak like the descriped in the
following:
IBM APAR: IY91677: ADD MEMORY CACHING FOR LIBS GETUSERATTR/PUTUSERATTR CALLS
Is this Patch Installed? instfix -ivk IY91677
What is your Level on AIX 5.3? oslevel -s
5300-06-03-0732
If the patch isn't installed and the second set of numbers is below 06 you
may want to watch the msc process to see grows over time...
Quick script to put in cron or run from time to time.
>cat capture_msc.ksh
HOSTNAME=`hostname`
export INFORMIXDIR=/usr/informix
export INFORMIXSERVER=%badserver%
export PATH=$PATH:$HOME/bin:.:$INFORMIXDIR/bin:export
LIBPATH=$INFORMIXDIR/lib:$INFORMIXDIR/lib/esql:$INFORMIXDIR/lib/tools:/usr/lib
MSC_PID=`onstat -g sch |grep msc |sed "s/ */~/g"|cut -f3 -d'~' |head -1`
ps_elf_data=`ps -elf |grep $MSC_PID |grep -v grep`
runtime=`date '+%Y%m%d%H%M%S'`
echo "$runtime $ps_elf_data" >> /u/informix/log/msc_info_dms.out
Hope this helps.
Eric B. Rowell
On Wed, Jan 14, 2009 at 5:44 PM, Neil Truby <neil.truby@ardenta.com> wrote:
> IDS 10.0FC8X6 on AIX 5.3
>
> Our Informix instance simply packed up at a low-activity that is, it just
> hanged, with a "Bloked by Checkpoint" message.
>
> We took an onstat -a and AIX dump, the latter re-booting the server, after
> which everything came up fine. But we had a 52-minute downtime on a 24x7
> system.
>
> PMRs were then opened with IBM Informix and AIX support.
>
> The AIX team reckon that the server was or tried to use more than the 80%
> AIX limit on "pinned" (their word for resident I think) memory and
> recommend
> either:
> - buying mroe memory
> - reducing Informix's footprint
> - making Informix run non-resident
>
> I don't think more memory is necessary.
> What we did think we'd do is change RESIDENT from -1 to 1, so that the
> virtual portion is non-resident. That would have avoided this problem it
> seems, but this doesn't fit with current IBM orthodoxy I think.
> The IBM Informix technican, uderstadndably cautious, was prepared to say
> only "I don't believe from a IDS point of view you will see any impact on
> performance, unless you start to see swapping"
> I need to make a recommendation to the customer who is a little annoyed
> that
> "IBM" (as he perceives the two different teams who looked at the problem)
> can't idenify the issue so how can we avoid it happening again?
>
> Any thoughts on the RESIDENT setting (or anything else)?
>
> thanks
> Neil
>
> A few facts:
> Server is an IBM p570 with 32g memory.
> Current Informix memory usage is:
> IBM Informix Dynamic Server Version 10.00.FC8X6 -- On-Line (Prim) -- Up 6
> days 21:11:49 -- 20788800 Kbytes>
> $ onstat -c | egrep -i "buff|shmvirt"> SHMVIRTSIZE 10274660 # initial virtual shared memory segment> size
> BUFFERPOOL
> size=4K,buffers=2500000,lrus=512,lru_min_dirty=0.10,lru_max_dirty=0.20>
> Server had been up for 342 days and Informix for 189.
>
> Online log:
> Thu Jan 8 00:00:36 2009
>
> 00:00:36 1800 buffers dirty
> 00:00:36 oldest lsn loguniq 97192, logpos 0x1211018
> 00:00:36 1742 dirty pages are to be flushed
>
> 00:00:37 dskflush() took 0 seconds
> 00:00:37 wait4critex() took 0 seconds
> 00:00:37 63 buffers dirty
> 00:00:37 oldest lsn loguniq 97192, logpos 0x1211018
> 00:00:37 safe_dskflush() took 0 seconds
> 00:00:37 Checkpoint Completed: duration was 0 seconds.
> 00:00:37 Checkpoint loguniq 97192, logpos 0x1f35018, timestamp: 0x17dca0ed
>
> 00:00:37 Maximum server connections 923
> 00:00:37 Buffer manager: starting coarse downgrades.
> 00:00:37 Buffer manager: finished coarse downgrades.
> Starting threshold adjustments.
> 00:00:37 Buffer manager: finished threshold adjustments.
> 00:15:37 1033 buffers dirty
> 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> 00:15:37 837 dirty pages are to be flushed
>
> 00:15:37 dskflush() took 0 seconds
> 00:15:37 wait4critex() took 0 seconds
> 00:15:37 196 buffers dirty
> 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> 00:15:37 safe_dskflush() took 0 seconds
> 00:15:37 Checkpoint Completed: duration was 0 seconds.
> 00:15:37 Checkpoint loguniq 97192, logpos 0x2cdf018, timestamp: 0x17e18f4e
>
> 00:15:37 Maximum server connections 923
> 00:15:37 Buffer manager: starting coarse downgrades.
> 00:15:37 Buffer manager: finished coarse downgrades.
> Starting threshold adjustments.
> 00:15:37 Buffer manager: finished threshold adjustments.
> 00:23:33 Logical Log 97192 Complete, timestamp: 0x17e4135a.
> 00:23:33 Logical Log 97192 - Backup Started
> 00:23:34 KAIO: out of OS resources, errno = 12, pid = 53534
> 00:23:40 Logical Log 97192 - Backup Completed
> 00:25:36 DR: ping timeout
> 00:25:36 DR: Receive error
> 00:25:36 ASF Echo-Thread Server: asfcode = -25582: oserr = 4: errstr = :> Network connection is broken.
> System error = 4.
> 00:30:37 368 buffers dirty
> 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> 00:30:37 347 dirty pages are to be flushed
>
> 00:30:37 dskflush() took 0 seconds
> 00:30:37 wait4critex() took 0 seconds
> 00:30:37 21 buffers dirty
> 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> 00:30:37 safe_dskflush() took 0 seconds
> 01:28:27 IBM Informix Dynamic Server Started.
> 01:28:27 WARNING: If you intend to use J/Foundation or GLS for Unicode> feature(GLU) with this Server instance, please make sure that your SHMBASE
> value specifies in onconfig is 0x700000010000000 o
> r above. Otherwise you will have problems while attaching or dynamimically
> adding virtual shared memory segments. Please refer to Server machine notes
> for more information.
>
> 01:28:47 Segment locked: addr=700000000000000, size=10765910016
> 01:28:47 Requested shared memory segment size rounded from 10274660KB to> 10274672KB
> 01:29:00 IBM Informix Dynamic Server Started.>
>
>
> _______________________________________________
> Informix-list mailing list
> Informix-list@iiug.org
> http://www.iiug.org/mailman/listinfo/informix-list
>
--
Eric B. Rowell
"Eric Rowell" <erowell@gmail.com> wrote in message news:mailman.143.1232031123.1831.informix-list@iiug.org... Never had an issue with " RESIDENT " but then again currently we are not using it. Too many things sharing the server. Are you sure you aren't running into a memory leak like the descriped in the following: IBM APAR: IY91677: ADD MEMORY CACHING FOR LIBS GETUSERATTR/PUTUSERATTR CALLS Is this Patch Installed? instfix -ivk IY91677 What is your Level on AIX 5.3? oslevel -s 5300-06-03-0732 # oslevel -s 5300-06-05-0806 We had that bug when we first went live, but on IBM's advice patched to obviate it I think. rgds Neil
Hi Neil,
I'd go along with what Keith already said. It seems that IDS was not
able to flush dirty buffers during the checkpoint not only fast enough
(like in a long checkpoint), but to write at all. The next online.log
line points in that direction:
00:23:34 KAIO: out of OS resources, errno = 12, pid = 53534
And then immediately at the next checkpoint it blocked.
Do you have env. var. IFMX_AIXKAIO_NUM_REQ exported in your
environment? If yes, to which value?
We also had memory leak Eric mentioned, but it only resulted in new
connection/authentication attempts being refused, nothing else.
Cheers
Davorin
Neil Truby wrote:
> IDS 10.0FC8X6 on AIX 5.3
>
> Our Informix instance simply packed up at a low-activity that is, it
> just hanged, with a "Bloked by Checkpoint" message.
>
> We took an onstat -a and AIX dump, the latter re-booting the server,
> after which everything came up fine. But we had a 52-minute downtime on
> a 24x7 system.
>
> PMRs were then opened with IBM Informix and AIX support.
>
> The AIX team reckon that the server was or tried to use more than the
> 80% AIX limit on "pinned" (their word for resident I think) memory and
> recommend either:
> - buying mroe memory
> - reducing Informix's footprint
> - making Informix run non-resident
>
> I don't think more memory is necessary.
> What we did think we'd do is change RESIDENT from -1 to 1, so that the
> virtual portion is non-resident. That would have avoided this problem
> it seems, but this doesn't fit with current IBM orthodoxy I think.
> The IBM Informix technican, uderstadndably cautious, was prepared to say
> only "I don't believe from a IDS point of view you will see any impact
> on performance, unless you start to see swapping"
> I need to make a recommendation to the customer who is a little annoyed
> that "IBM" (as he perceives the two different teams who looked at the
> problem) can't idenify the issue so how can we avoid it happening again?
>
> Any thoughts on the RESIDENT setting (or anything else)?
>
> thanks
> Neil
>
> A few facts:
> Server is an IBM p570 with 32g memory.
> Current Informix memory usage is:
> IBM Informix Dynamic Server Version 10.00.FC8X6 -- On-Line (Prim) --
> Up 6 days 21:11:49 -- 20788800 Kbytes>
> $ onstat -c | egrep -i "buff|shmvirt"> SHMVIRTSIZE 10274660 # initial virtual shared memory segment> size
> BUFFERPOOL
> size=4K,buffers=2500000,lrus=512,lru_min_dirty=0.10,lru_max_dirty=0.20>
> Server had been up for 342 days and Informix for 189.
>
> Online log:
> Thu Jan 8 00:00:36 2009
>
> 00:00:36 1800 buffers dirty
> 00:00:36 oldest lsn loguniq 97192, logpos 0x1211018
> 00:00:36 1742 dirty pages are to be flushed
>
> 00:00:37 dskflush() took 0 seconds
> 00:00:37 wait4critex() took 0 seconds
> 00:00:37 63 buffers dirty
> 00:00:37 oldest lsn loguniq 97192, logpos 0x1211018
> 00:00:37 safe_dskflush() took 0 seconds
> 00:00:37 Checkpoint Completed: duration was 0 seconds.
> 00:00:37 Checkpoint loguniq 97192, logpos 0x1f35018, timestamp: 0x17dca0ed
>
> 00:00:37 Maximum server connections 923
> 00:00:37 Buffer manager: starting coarse downgrades.
> 00:00:37 Buffer manager: finished coarse downgrades.
> Starting threshold adjustments.
> 00:00:37 Buffer manager: finished threshold adjustments.
> 00:15:37 1033 buffers dirty
> 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> 00:15:37 837 dirty pages are to be flushed
>
> 00:15:37 dskflush() took 0 seconds
> 00:15:37 wait4critex() took 0 seconds
> 00:15:37 196 buffers dirty
> 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> 00:15:37 safe_dskflush() took 0 seconds
> 00:15:37 Checkpoint Completed: duration was 0 seconds.
> 00:15:37 Checkpoint loguniq 97192, logpos 0x2cdf018, timestamp: 0x17e18f4e
>
> 00:15:37 Maximum server connections 923
> 00:15:37 Buffer manager: starting coarse downgrades.
> 00:15:37 Buffer manager: finished coarse downgrades.
> Starting threshold adjustments.
> 00:15:37 Buffer manager: finished threshold adjustments.
> 00:23:33 Logical Log 97192 Complete, timestamp: 0x17e4135a.
> 00:23:33 Logical Log 97192 - Backup Started
> 00:23:34 KAIO: out of OS resources, errno = 12, pid = 53534
> 00:23:40 Logical Log 97192 - Backup Completed
> 00:25:36 DR: ping timeout
> 00:25:36 DR: Receive error
> 00:25:36 ASF Echo-Thread Server: asfcode = -25582: oserr = 4: errstr => : Network connection is broken.
> System error = 4.
> 00:30:37 368 buffers dirty
> 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> 00:30:37 347 dirty pages are to be flushed
>
> 00:30:37 dskflush() took 0 seconds
> 00:30:37 wait4critex() took 0 seconds
> 00:30:37 21 buffers dirty
> 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> 00:30:37 safe_dskflush() took 0 seconds
> 01:28:27 IBM Informix Dynamic Server Started.
> 01:28:27 WARNING: If you intend to use J/Foundation or GLS for Unicode> feature(GLU) with this Server instance, please make sure that your
> SHMBASE value specifies in onconfig is 0x700000010000000 o
> r above. Otherwise you will have problems while attaching or
> dynamimically adding virtual shared memory segments. Please refer to
> Server machine notes for more information.
>
> 01:28:47 Segment locked: addr=700000000000000, size=10765910016
> 01:28:47 Requested shared memory segment size rounded from 10274660KB> to 10274672KB
> 01:29:00 IBM Informix Dynamic Server Started.>
>
>
I may be missing the point completely but why are you focusing on RESIDENT?
Because of the KAIO error? Or because you have system logs that point to memory
issues?
Your HDR has raised an error... It should be bullet proof, but I've seen
situations where this can leave the primary waiting for the checkpoint
acknowledge, so effectively "blocked by checkpoint".
Any idea of what happened to the network? Do the system logs point to lack of
memory? Did it also affect the network connections?
What do you had on the secondary online log?
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
"Fernando Nunes" <domusonline@gmail.com> wrote in message
news:gkolep$rj5$1@news.motzarella.org...
> I may be missing the point completely but why are you focusing on
> RESIDENT? Because of the KAIO error? Or because you have system logs that
> point to memory issues?
Because IBM AIX support tells us that the problemm is related to an atempt
to exceed more than 80% "pinned" memory.
> Your HDR has raised an error... It should be bullet proof, but I've seen
> situations where this can leave the primary waiting for the checkpoint
> acknowledge, so effectively "blocked by checkpoint".
> Any idea of what happened to the network?
Nothing, I don't think, since I was able to ssh to the box remotely without
any problem.
> Do the system logs point to lack of memory?
According to AIX Support, yes.
> Did it also affect the network connections?
Not that I noticed.
> What do you had on the secondary online log?
00:15:37 Maximum server connections 15
00:23:33 Logical Log 97192 Complete, timestamp: 0x17e49e7f.
00:25:54 DR: ping timeout
00:25:54 DR: Receive error
00:25:54 ASF Echo-Thread Server: asfcode = -25582: oserr = 119: errstr = :Network connection is broken.
System error = 119.
00:25:56 DR: Turned off on secondary server
00:25:56 Process exited with return code 127: /bin/sh /bin/sh -c
/opt/informix/10.0/etc/alarmprogram_1.sh 315 "Data Replication failure." "DR: Turned off on secondary
01:29:21 DR: Secondary server connected
01:29:21 DR: Secondary server needs failure recovery
01:29:37 DR: Failure recovery from disk in progress ...
01:29:37 Logical Log 97193 Complete, timestamp: 0x17f01c61.
01:29:37 452 buffers dirty
01:29:37 oldest lsn loguniq 40605, logpos 0x2a58018
01:29:37 319 dirty pages are to be flushed
01:29:38 dskflush() took 0 seconds
01:29:38 wait4critex() took 0 seconds
01:29:38 136 buffers dirty
On Jan 15, 2:09 am, superbo...@t-online.de wrote: > Hello Ian, > > i do not agree with you on that, have seen situations where setting > resident to 1 or -1 or -2 > made performance from not existing to really good; > specially on systems with a lot of filesystem io > not all systems can limit the amount of cache which will be allocated > for filesystem cache. > This can kill database performance. > > Superboer. Well you're always welcome to disagree. Here's the issue in a nutshell. Why do you make the memory pages resident and not let them be swapped? Yes I know its for faster access. However here's the thing. You make IDS memory resident, you've not forced other things to swap in and out in an effort to compensate. If you were really using IDS under a consistent and 'heavy enough' load, IDS would not swap out. It would/should remain memory resident unless another process was getting starved and it had to be swapped out. So setting the flag really doesn't really do much if your database server is really being used. As we watch kernels become self tuning, when they power up, the kernel sees what resources it has available and it will use a formula to help tune the kernel to give you an 'optimum' use of resources. This is going to occur before the machine is even at run level 1. You are probably not starting your engine until run level 5. So if you make IDS memory resident, you've now tied up resources that the machine would have expected to be available. This is why you shouldn't just set memory resident 'automatically'. I think that the larger problem is that those individuals who are teaching performance tuning classes never really took an OS course or had a lot of experience as Sysadmins and built machines from the ground up. What they are teaching is that it may help improve performance, and it probably doesn't hurt anything so set it to true. Yes I've seen people say to students. The issue isn't just with IDS. It can happen with other databases. You're not 'starving' IDS, but you can be starving an OS process that IDS relies on and if the OS gets starved, bad things can happen. Does that make sense? As I've said, the odds are that in 90+% of the instances of IDS running, you do not need to consider turning on memory residence.
Eric,
Sorry to top post...
You mention a memory leak.
Ok so what happens when you have a memory leak and you have RESIDENT
Memory turned on?
With respect to Keith's post, you can reboot and the problem can go
away for a long time. It doesn't mean that the problem is solved or
unrelated to using the resident memory flag.
On Jan 15, 8:51 am, Eric Rowell <erow...@gmail.com> wrote:
> Never had an issue with " RESIDENT " but then again currently we are not
> using it. Too many things sharing the server.
>
> Are you sure you aren't running into a memory leak like the descriped in the
> following:
>
> IBM APAR: IY91677: ADD MEMORY CACHING FOR LIBS GETUSERATTR/PUTUSERATTR CALLS
>
> Is this Patch Installed? instfix -ivk IY91677
>
> What is your Level on AIX 5.3? oslevel -s
> 5300-06-03-0732
>
> If the patch isn't installed and the second set of numbers is below 06 you
> may want to watch the msc process to see grows over time...
>
> Quick script to put in cron or run from time to time.
>
> >cat capture_msc.ksh
>
> HOSTNAME=`hostname`
>
> export INFORMIXDIR=/usr/informix
> export INFORMIXSERVER=%badserver%
> export PATH=$PATH:$HOME/bin:.:$INFORMIXDIR/bin:> export
> LIBPATH=$INFORMIXDIR/lib:$INFORMIXDIR/lib/esql:$INFORMIXDIR/lib/tools:/usr/lib
>
> MSC_PID=`onstat -g sch |grep msc |sed "s/ */~/g"|cut -f3 -d'~' |head -1`
> ps_elf_data=`ps -elf |grep $MSC_PID |grep -v grep`
> runtime=`date '+%Y%m%d%H%M%S'`
> echo "$runtime $ps_elf_data" >> /u/informix/log/msc_info_dms.out
> Hope this helps.
>
> Eric B. Rowell
>
>
>
> On Wed, Jan 14, 2009 at 5:44 PM, Neil Truby <neil.tr...@ardenta.com> wrote:
> > IDS 10.0FC8X6 on AIX 5.3
>
> > Our Informix instance simply packed up at a low-activity that is, it just
> > hanged, with a "Bloked by Checkpoint" message.
>
> > We took an onstat -a and AIX dump, the latter re-booting the server, after
> > which everything came up fine. But we had a 52-minute downtime on a 24x7
> > system.
>
> > PMRs were then opened with IBM Informix and AIX support.
>
> > The AIX team reckon that the server was or tried to use more than the 80%
> > AIX limit on "pinned" (their word for resident I think) memory and
> > recommend
> > either:
> > - buying mroe memory
> > - reducing Informix's footprint
> > - making Informix run non-resident
>
> > I don't think more memory is necessary.
> > What we did think we'd do is change RESIDENT from -1 to 1, so that the
> > virtual portion is non-resident. That would have avoided this problem it
> > seems, but this doesn't fit with current IBM orthodoxy I think.
> > The IBM Informix technican, uderstadndably cautious, was prepared to say
> > only "I don't believe from a IDS point of view you will see any impact on
> > performance, unless you start to see swapping"
> > I need to make a recommendation to the customer who is a little annoyed
> > that
> > "IBM" (as he perceives the two different teams who looked at the problem)
> > can't idenify the issue so how can we avoid it happening again?
>
> > Any thoughts on the RESIDENT setting (or anything else)?
>
> > thanks
> > Neil
>
> > A few facts:
> > Server is an IBM p570 with 32g memory.
> > Current Informix memory usage is:
> > IBM Informix Dynamic Server Version 10.00.FC8X6 -- On-Line (Prim) -- Up 6
> > days 21:11:49 -- 20788800 Kbytes>
> > $ onstat -c | egrep -i "buff|shmvirt"> > SHMVIRTSIZE 10274660 # initial virtual shared memory segment> > size
> > BUFFERPOOL
> > size=4K,buffers=2500000,lrus=512,lru_min_dirty=0.10,lru_max_dirty=0.20>
> > Server had been up for 342 days and Informix for 189.
>
> > Online log:
> > Thu Jan 8 00:00:36 2009
>
> > 00:00:36 1800 buffers dirty
> > 00:00:36 oldest lsn loguniq 97192, logpos 0x1211018
> > 00:00:36 1742 dirty pages are to be flushed
>
> > 00:00:37 dskflush() took 0 seconds
> > 00:00:37 wait4critex() took 0 seconds
> > 00:00:37 63 buffers dirty
> > 00:00:37 oldest lsn loguniq 97192, logpos 0x1211018
> > 00:00:37 safe_dskflush() took 0 seconds
> > 00:00:37 Checkpoint Completed: duration was 0 seconds.
> > 00:00:37 Checkpoint loguniq 97192, logpos 0x1f35018, timestamp: 0x17dca0ed
>
> > 00:00:37 Maximum server connections 923
> > 00:00:37 Buffer manager: starting coarse downgrades.
> > 00:00:37 Buffer manager: finished coarse downgrades.> > Starting threshold adjustments.
> > 00:00:37 Buffer manager: finished threshold adjustments.
> > 00:15:37 1033 buffers dirty
> > 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> > 00:15:37 837 dirty pages are to be flushed
>
> > 00:15:37 dskflush() took 0 seconds
> > 00:15:37 wait4critex() took 0 seconds
> > 00:15:37 196 buffers dirty
> > 00:15:37 oldest lsn loguniq 97192, logpos 0x1f35018
> > 00:15:37 safe_dskflush() took 0 seconds
> > 00:15:37 Checkpoint Completed: duration was 0 seconds.
> > 00:15:37 Checkpoint loguniq 97192, logpos 0x2cdf018, timestamp: 0x17e18f4e
>
> > 00:15:37 Maximum server connections 923
> > 00:15:37 Buffer manager: starting coarse downgrades.
> > 00:15:37 Buffer manager: finished coarse downgrades.> > Starting threshold adjustments.
> > 00:15:37 Buffer manager: finished threshold adjustments.
> > 00:23:33 Logical Log 97192 Complete, timestamp: 0x17e4135a.
> > 00:23:33 Logical Log 97192 - Backup Started
> > 00:23:34 KAIO: out of OS resources, errno = 12, pid = 53534
> > 00:23:40 Logical Log 97192 - Backup Completed
> > 00:25:36 DR: ping timeout
> > 00:25:36 DR: Receive error
> > 00:25:36 ASF Echo-Thread Server: asfcode = -25582: oserr = 4: errstr = :> > Network connection is broken.
> > System error = 4.
> > 00:30:37 368 buffers dirty
> > 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> > 00:30:37 347 dirty pages are to be flushed
>
> > 00:30:37 dskflush() took 0 seconds
> > 00:30:37 wait4critex() took 0 seconds
> > 00:30:37 21 buffers dirty
> > 00:30:37 oldest lsn loguniq 97192, logpos 0x2cdf018
> > 00:30:37 safe_dskflush() took 0 seconds
> > 01:28:27 IBM Informix Dynamic Server Started.
> > 01:28:27 WARNING: If you intend to use J/Foundation or GLS for Unicode> > feature(GLU) with this Server instance, please make sure that your SHMBASE
> > value specifies in onconfig is 0x700000010000000 o
> > r above. Otherwise you will have problems while attaching or dynamimically
> > adding virtual shared memory segments. Please refer to Server machine notes
> > for more information.
>
> > 01:28:47 Segment locked: addr=700000000000000, size=10765910016
> > 01:28:47 Requested shared memory segment size rounded from 10274660KB to> > 10274672KB
> > 01:29:00 IBM Informix Dynamic Server Started.>
> > _______________________________________________
> > Informix-list mailing list
> > Informix-l...@iiug.org
> >http://www.iiug.org/mailman/listinfo/informix-list
>
> --
> Eric B. Rowell
> From: domusonline@gmail.com > Subject: Re: IDS hangs; IBM diagnoses resident memory problem > Date: Fri, 16 Jan 2009 00:48:05 +0000 > To: informix-list@iiug.org [SNIP] > > I may be missing the point completely but why are you focusing on RESIDENT? > Because of the KAIO error? Or because you have system logs that point to memory > issues? > > Your HDR has raised an error... It should be bullet proof, but I've seen > situations where this can leave the primary waiting for the checkpoint > acknowledge, so effectively "blocked by checkpoint". > > Any idea of what happened to the network? Do the system logs point to lack of > memory? Did it also affect the network connections? > What do you had on the secondary online log? > > > -- > Fernando Nunes I think the point was that Neil focused on the Resident Memory issue because IBM told him to. It could be something in IDS, or it could be something in AIX where he was blocked because of a lack of resources. You know, something like "memory" ? The hard thing is that Neil could reboot and not be able to repeat the issue again. The problem could be a compounding of several issues. The point I'm trying to make is that the concept of resident memory being set to true is not necessarily a good thing and unless you need it, don't use it. But hey! What do I know? Its not like I've ever tuned a kernel or even Tuned a Fish before.... ;-) (You have to be over 45 to understand that joke. ;-) -g _________________________________________________________________ Windows Live™: Keep your life in sync. http://windowslive.com/explore?ocid=TXT_TAGLM_WL_t1_allup_explore_012009
Neil Truby wrote:
>
> "Fernando Nunes" <domusonline@gmail.com> wrote in message
> news:gkolep$rj5$1@news.motzarella.org...
>
>> I may be missing the point completely but why are you focusing on
>> RESIDENT? Because of the KAIO error? Or because you have system logs
>> that point to memory issues?
>
> Because IBM AIX support tells us that the problemm is related to an
> atempt to exceed more than 80% "pinned" memory.
>
>> Your HDR has raised an error... It should be bullet proof, but I've
>> seen situations where this can leave the primary waiting for the
>> checkpoint acknowledge, so effectively "blocked by checkpoint".
>> Any idea of what happened to the network?
>
> Nothing, I don't think, since I was able to ssh to the box remotely
> without any problem.
>
>> Do the system logs point to lack of memory?
> According to AIX Support, yes.
>
>> Did it also affect the network connections?
> Not that I noticed.
>
>> What do you had on the secondary online log?
> 00:15:37 Maximum server connections 15
> 00:23:33 Logical Log 97192 Complete, timestamp: 0x17e49e7f.
> 00:25:54 DR: ping timeout
> 00:25:54 DR: Receive error
> 00:25:54 ASF Echo-Thread Server: asfcode = -25582: oserr = 119: errstr> = : Network connection is broken.
> System error = 119.
> 00:25:56 DR: Turned off on secondary server
> 00:25:56 Process exited with return code 127: /bin/sh /bin/sh -c
> /opt/informix/10.0/etc/alarmprogram_1.sh 3> 15 "Data Replication failure." "DR: Turned off on secondary
After this point the secondary has not written anything to the log.
Notice that the primary did not assume the replication was down.
If the primary does not receive the checkpoint acknowledge, it will block. It
_should_ timeout, but for that the HDR threads should be ok. Maybe they weren't.
I'd suggest you ask IDS support if they noticed this. Maybe they can check the
thread status...
From the secondary point of view the primary died. When it came up, all the
normal messages...
> 01:29:21 DR: Secondary server connected
> 01:29:21 DR: Secondary server needs failure recovery
>
> 01:29:37 DR: Failure recovery from disk in progress ...
> 01:29:37 Logical Log 97193 Complete, timestamp: 0x17f01c61.
> 01:29:37 452 buffers dirty
> 01:29:37 oldest lsn loguniq 40605, logpos 0x2a58018
> 01:29:37 319 dirty pages are to be flushed
>
> 01:29:38 dskflush() took 0 seconds
> 01:29:38 wait4critex() took 0 seconds
> 01:29:38 136 buffers dirty>
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
Neil Truby wrote:
>
> "Fernando Nunes" <domusonline@gmail.com> wrote in message
> news:gkolep$rj5$1@news.motzarella.org...
>
>> I may be missing the point completely but why are you focusing on
>> RESIDENT? Because of the KAIO error? Or because you have system logs
>> that point to memory issues?
>
> Because IBM AIX support tells us that the problemm is related to an
> atempt to exceed more than 80% "pinned" memory.
>
>> Your HDR has raised an error... It should be bullet proof, but I've
>> seen situations where this can leave the primary waiting for the
>> checkpoint acknowledge, so effectively "blocked by checkpoint".
>> Any idea of what happened to the network?
>
> Nothing, I don't think, since I was able to ssh to the box remotely
> without any problem.
>
>> Do the system logs point to lack of memory?
> According to AIX Support, yes.
>
>> Did it also affect the network connections?
> Not that I noticed.
>
>> What do you had on the secondary online log?
> 00:15:37 Maximum server connections 15
> 00:23:33 Logical Log 97192 Complete, timestamp: 0x17e49e7f.
> 00:25:54 DR: ping timeout
> 00:25:54 DR: Receive error
> 00:25:54 ASF Echo-Thread Server: asfcode = -25582: oserr = 119: errstr> = : Network connection is broken.
> System error = 119.
> 00:25:56 DR: Turned off on secondary server
> 00:25:56 Process exited with return code 127: /bin/sh /bin/sh -c
> /opt/informix/10.0/etc/alarmprogram_1.sh 3> 15 "Data Replication failure." "DR: Turned off on secondary
After this point the secondary has not written anything to the log.
Notice that the primary did not assume the replication was down.
If the primary does not receive the checkpoint acknowledge, it will block. It
_should_ timeout, but for that the HDR threads should be ok. Maybe they weren't.
I'd suggest you ask IDS support if they noticed this. Maybe they can check the
thread status...
From the secondary point of view the primary died. When it came up, all the
normal messages...
> 01:29:21 DR: Secondary server connected
> 01:29:21 DR: Secondary server needs failure recovery
>
> 01:29:37 DR: Failure recovery from disk in progress ...
> 01:29:37 Logical Log 97193 Complete, timestamp: 0x17f01c61.
> 01:29:37 452 buffers dirty
> 01:29:37 oldest lsn loguniq 40605, logpos 0x2a58018
> 01:29:37 319 dirty pages are to be flushed
>
> 01:29:38 dskflush() took 0 seconds
> 01:29:38 wait4critex() took 0 seconds
> 01:29:38 136 buffers dirty>
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
> Well you're always welcome to disagree. > > Here's the issue in a nutshell. > Why do you make the memory pages resident and not let them be swapped? To make damn sure that the filesystemcache does not get its hands on the memory. > The issue isn't just with IDS. It can happen with other databases. true. > You're not 'starving' IDS, but you can be starving an OS process that > IDS relies on and if the OS gets starved, bad things can happen. That smells like one has not enough memory in the machine to begin with a system like that should be retuned or memory should be added. Superboer.
> From: superboer7@t-online.de > Subject: Re: IDS hangs; IBM diagnoses resident memory problem > Date: Mon, 19 Jan 2009 01:30:01 -0800 > To: informix-list@iiug.org > > > Well you're always welcome to disagree. > > > > Here's the issue in a nutshell. > > Why do you make the memory pages resident and not let them be swapped? > > To make damn sure that the filesystemcache does not get its hands on > the memory. > How often do you see machines swap these days? > > The issue isn't just with IDS. It can happen with other databases. > true. > > You're not 'starving' IDS, but you can be starving an OS process that > > IDS relies on and if the OS gets starved, bad things can happen. > > That smells like one has not enough memory in the machine to begin > with > a system like that should be retuned or memory should be added. > > And how much is enough memory these days? 32GB? 64GB? 128GB? The point is that yes, its cheaper to always add more memory or upgrade a piece of hardware. Sometimes its not possible. And more to the point, if you add more memory, why do you need to make your IDS memory resident? The answer is that you don't. The OS won't swap unless it has to, and it won't swap out pages that have recently been active. So if you're using your IDS, then you IDS won't swap because the memory is always going to be more active than a different application. This also goes to the point that you'd be better off with a good system admin, network admin and IT Architect who has planned for growth and instead of trying to put everything on one server, they've split the load up. (Ok, so you really have two database servers because they are using ER/HDR) Part of the issue is that there could be a memory leak and the memory requirements grow to a point that you force swap. Part of the problem is that the kernel tuning parameters don't take into account that there will be applications started after run level 5 and the OS buffers and expected memory may not be there. Again, it goes back to the point. Unless you know you need and can prove you need Memory Residence, don't set the flag on. _________________________________________________________________ Windows Live™ Hotmail®: Chat. Store. Share. Do more with mail. http://windowslive.com/explore?ocid=TXT_TAGLM_WL_t1_hm_justgotbetter_explore_012009