problem with IDS 11.70FC4 and RedHat 5.6
Posted in 2013
A user on IDS 11.70.FC4 / RHEL 5.6 reported a recurring hang: after an ASF 'Network connection is broken' (-25582) message the instance stopped answering requests and had to be bounced manually; startup also logged 'Insufficient free huge pages'. Respondents asked for onstat -g seg output and an AF stack trace (none was produced), suggested tuning VP_MEMORY_CACHE_KB and enlarging/fixing shared memory settings, and Art Kagel explained how to allocate huge pages via /proc/sys/vm/nr_hugepages or vm.nr_hugepages in sysctl.conf. IBM support advised upgrading from FC4 and opening a PMR. No resolution of the hang is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Performance & Tuning, Storage & Space Management, Stored Procedures & SPL, Security, Permissions & Auditing, Logging & Checkpoints, Networking & sqlhosts Configuration, Versions, Editions & End-of-Life
I have IDS 11.70 FC4 intalled in RedHat 5.6 and we are having a recurrent
problem with one of our instances
This morning we got the following error that make us restart the instance, the
online log entrances follow:
Mon Feb 18 07:59:11 2013
07:59:11 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
Network connection is broken.
07:59:37 Checkpoint Completed: duration was 1 seconds.
07:59:37 Mon Feb 18 - loguniq 190, logpos 0x4817108, timestamp: 0xbd95965a
Interval: 34917
07:59:37 Maximum server connections 621
07:59:37 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
Plog used 7490, Llog used 16131
07:59:38 IBM Informix Dynamic Server Stopped.
07:59:43 IBM Informix Dynamic Server Started.
07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
segment.
Requested: 2629509120 bytes. Available: 0 bytes.
The default memory page size will be used.
07:59:43 Segment locked: addr=0x44000000, size=2629509120
07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
segment.
Requested: 2147483648 bytes. Available: 0 bytes.
The default memory page size will be used.
07:59:43 Segment locked: addr=0xe0bb2000, size=2147483648
Mon Feb 18 07:59:45 2013
07:59:45 Event alarms enabled. ALARMPROG =
'/opt/IBM/informix/etc/alarmprogram.sh'
07:59:45 Booting Language <c> from module <>
07:59:45 Loading Module <CNULL>
07:59:45 Booting Language <builtin> from module <>
07:59:45 Loading Module <BUILTINNULL>
07:59:51 DR: DRAUTO is 0 (Off)
07:59:51 DR: ENCRYPT_HDR is 0 (HDR encryption Disabled)
07:59:51 Event notification facility epoll enabled.
07:59:51 Could not disable priority aging: errno = 4
07:59:51 Could not disable priority aging: errno = 4
07:59:51 Could not disable priority aging: errno = 4
07:59:51 Could not disable priority aging: errno = 4
07:59:51 Could not disable priority aging: errno = 4
07:59:51 Could not disable priority aging: errno = 4
07:59:51 IBM Informix Dynamic Server Version 11.70.FC4 Software Serial Number
AAA#B000000
07:59:51 Performance Advisory: The physical log size is smaller than the
recommended size for a
server configured with RTO_SERVER_RESTART.
07:59:51 Results: Fast recovery performance might not be optimal.
07:59:51 Action: For best fast recovery performance when RTO_SERVER_RESTART is
enabled,
increase the physical log size to at least 2306867 KB. For servers
configured with a large buffer pool, this might not be necessary.
07:59:51 IBM Informix Dynamic Server Initialized -- Shared Memory Initialized.
07:59:51 Started 1 B-tree scanners.
07:59:51 B-tree scanner threshold set at 5000.
07:59:51 B-tree scanner range scan size set to -1.
07:59:51 B-tree scanner ALICE mode set to 6.
07:59:51 B-tree scanner index compression level set to med.
07:59:51 Physical Recovery Started at Page (1:33083).
07:59:52 Physical Recovery Complete: 255 Pages Examined, 255 Pages Restored.
07:59:52 Logical Recovery Started.
07:59:52 10 recovery worker threads will be started.
07:59:53 Logical Recovery has reached the transaction cleanup phase.
07:59:53 Logical Recovery Complete.
0 Committed, 1 Rolled Back, 0 Open, 0 Bad Locks
07:59:54 Dataskip is now OFF for all dbspaces
07:59:54 requested number of KAIO events (8192) exceeds limit (6144). using
6144.
07:59:55 Checkpoint Completed: duration was 1 seconds.
07:59:55 Mon Feb 18 - loguniq 190, logpos 0x48192c8, timestamp: 0xbd9594aa
Interval: 34918
07:59:55 Maximum server connections 0
07:59:55 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
Plog used 272, Llog used 1
07:59:55 On-Line Mode
07:59:56 SCHAPI: Started dbScheduler thread.
07:59:57 Booting Language <spl> from module <>
07:59:57 Loading Module <SPLNULL>
07:59:57 Auto Registration is synced
07:59:57 SCHAPI: Started 2 dbWorker threads.
07:59:59 Defragmenter cleaner thread now running
07:59:59 Defragmenter cleaner thread cleaned:0 partitions
08:00:56 Loading Module <$INFORMIXDIR/extend/ifxmngr/ifxmngr.bld>
08:00:56 The C Language Module </opt/IBM/informix/extend/ifxmngr/ifxmngr.bld>
loaded
The configuration of the instance is as follow:
RESIDENT -1
SHMBASE 0x44000000L
SHMVIRTSIZE 2097152
SHMADD 32768
EXTSHMADD 32768
SHMTOTAL 6291456
MULTIPROCESSOR 1
VPCLASS cpu,num=4,noage
VPCLASS soc,num=1,noage
VP_MEMORY_CACHE_KB 0
SINGLE_CPU_VP 0
BUFFERPOOL
size=2K,buffers=1048576,lrus=32,lru_min_dirty=50.00,lru_max_dirty=60.00
I can't find the reason for this problem, can any of you point me in the right
direction?
I'll be very greatfull.
Best Regards.
Javier Hernández
Hi,
What is the output of onstat -g seg ?
Its good to tune VP_MEMORY_CACHE_KB 10240 also.
rgds
schillache
Hello.
What´s the stack trace generated in AF file?
You should be hitting some bug, I saw it once, in 11.70, but can´t find it
without your stack trace.
Please post your stack and we could help you more, ok?
(You should upgrade as soon as possible. xC4 has several troubles. If you
cannot, you should maximize your SHM settings, so that your segments doesn´t
grow/reduce too often - or even never, full your SHM until you reach your
SHMTOTAL).
(I remember that´s what I did, in my case).
Regards.
Alexandre Marini
IBM Informix Certified Professional v10 / v11.50 / v11.70
IBM Information Management Informix Technical Professional
IBM Infosphere DataStage Technical Professional
Informix Senior DBA - Orizon Brasil
BRIUG website administrator
Informix independent consultant
> To: ids@iiug.org
> From: javier_hdzt@yahoo.com.mx
> Subject: problem with IDS 11.70FC4 and RedHat 5.6 [29524]
> Date: Mon, 18 Feb 2013 17:34:20 -0500
>
> I have IDS 11.70 FC4 intalled in RedHat 5.6 and we are having a recurrent
> problem with one of our instances
>
> This morning we got the following error that make us restart the instance,
the
> online log entrances follow:
>
> Mon Feb 18 07:59:11 2013
>
> 07:59:11 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
> Network connection is broken.
>
> 07:59:37 Checkpoint Completed: duration was 1 seconds.
> 07:59:37 Mon Feb 18 - loguniq 190, logpos 0x4817108, timestamp: 0xbd95965a
> Interval: 34917
>
> 07:59:37 Maximum server connections 621
> 07:59:37 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
> Plog used 7490, Llog used 16131
>
> 07:59:38 IBM Informix Dynamic Server Stopped.
>
> 07:59:43 IBM Informix Dynamic Server Started.
> 07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
> segment.
>
> Requested: 2629509120 bytes. Available: 0 bytes.
>
> The default memory page size will be used.
> 07:59:43 Segment locked: addr=0x44000000, size=2629509120
> 07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
> segment.
>
> Requested: 2147483648 bytes. Available: 0 bytes.
>
> The default memory page size will be used.
> 07:59:43 Segment locked: addr=0xe0bb2000, size=2147483648
>
> Mon Feb 18 07:59:45 2013
>
> 07:59:45 Event alarms enabled. ALARMPROG =
> '/opt/IBM/informix/etc/alarmprogram.sh'
> 07:59:45 Booting Language <c> from module <>
> 07:59:45 Loading Module <CNULL>
> 07:59:45 Booting Language <builtin> from module <>
> 07:59:45 Loading Module <BUILTINNULL>
> 07:59:51 DR: DRAUTO is 0 (Off)
> 07:59:51 DR: ENCRYPT_HDR is 0 (HDR encryption Disabled)
> 07:59:51 Event notification facility epoll enabled.
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 IBM Informix Dynamic Server Version 11.70.FC4 Software Serial Number
> AAA#B000000
> 07:59:51 Performance Advisory: The physical log size is smaller than the
> recommended size for a
> server configured with RTO_SERVER_RESTART.
> 07:59:51 Results: Fast recovery performance might not be optimal.
> 07:59:51 Action: For best fast recovery performance when RTO_SERVER_RESTART
is
> enabled,
> increase the physical log size to at least 2306867 KB. For servers
> configured with a large buffer pool, this might not be necessary.
> 07:59:51 IBM Informix Dynamic Server Initialized -- Shared Memory
Initialized.
>
> 07:59:51 Started 1 B-tree scanners.
> 07:59:51 B-tree scanner threshold set at 5000.
> 07:59:51 B-tree scanner range scan size set to -1.
> 07:59:51 B-tree scanner ALICE mode set to 6.
> 07:59:51 B-tree scanner index compression level set to med.
> 07:59:51 Physical Recovery Started at Page (1:33083).
> 07:59:52 Physical Recovery Complete: 255 Pages Examined, 255 Pages Restored.
> 07:59:52 Logical Recovery Started.
> 07:59:52 10 recovery worker threads will be started.
> 07:59:53 Logical Recovery has reached the transaction cleanup phase.
> 07:59:53 Logical Recovery Complete.
>
> 0 Committed, 1 Rolled Back, 0 Open, 0 Bad Locks
>
> 07:59:54 Dataskip is now OFF for all dbspaces
> 07:59:54 requested number of KAIO events (8192) exceeds limit (6144). using
> 6144.
> 07:59:55 Checkpoint Completed: duration was 1 seconds.
> 07:59:55 Mon Feb 18 - loguniq 190, logpos 0x48192c8, timestamp: 0xbd9594aa
> Interval: 34918
>
> 07:59:55 Maximum server connections 0
> 07:59:55 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
> Plog used 272, Llog used 1
>
> 07:59:55 On-Line Mode
> 07:59:56 SCHAPI: Started dbScheduler thread.
> 07:59:57 Booting Language <spl> from module <>
> 07:59:57 Loading Module <SPLNULL>
> 07:59:57 Auto Registration is synced
> 07:59:57 SCHAPI: Started 2 dbWorker threads.
> 07:59:59 Defragmenter cleaner thread now running
> 07:59:59 Defragmenter cleaner thread cleaned:0 partitions
> 08:00:56 Loading Module <$INFORMIXDIR/extend/ifxmngr/ifxmngr.bld>
> 08:00:56 The C Language Module </opt/IBM/informix/extend/ifxmngr/ifxmngr.bld>
> loaded
>
> The configuration of the instance is as follow:
>
> RESIDENT -1
> SHMBASE 0x44000000L
> SHMVIRTSIZE 2097152
> SHMADD 32768
> EXTSHMADD 32768
> SHMTOTAL 6291456
>
> MULTIPROCESSOR 1
> VPCLASS cpu,num=4,noage
> VPCLASS soc,num=1,noage
> VP_MEMORY_CACHE_KB 0
> SINGLE_CPU_VP 0
>
> BUFFERPOOL
> size=2K,buffers=1048576,lrus=32,lru_min_dirty=50.00,lru_max_dirty=60.00>
> I can't find the reason for this problem, can any of you point me in the
right
> direction?
>
> I'll be very greatfull.
>
> Best Regards.
>
> Javier Hernández
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Original post:
I have IDS 11.70 FC4 intalled in RedHat 5.6 and we are having a recurrent
problem with one of our instances
This morning we got the following error that make us restart the instance, the
online log entrances follow:
Mon Feb 18 07:59:11 2013
07:59:11 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
Network connection is broken.
07:59:37 Checkpoint Completed: duration was 1 seconds.
07:59:37 Mon Feb 18 - loguniq 190, logpos 0x4817108, timestamp: 0xbd95965a
Interval: 34917
07:59:37 Maximum server connections 621
07:59:37 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
Plog used 7490, Llog used 16131
07:59:38 IBM Informix Dynamic Server Stopped.
07:59:43 IBM Informix Dynamic Server Started.
<Bunch of stuff removed>
I can't find the reason for this problem, can any of you point me in the right
direction?
I'll be very greatfull.
Best Regards.
Javier Hernández
Response:
From the MSGPATH output you posted, it's not clear what the problem is. The
server stopped message appears to be from somebody shutting the server down
(onmode command). The only error before that is the 1 ASF error, which doesn't
look like the sort of thing that require a restart. Can you be more specific
on what you are trying to fix?
Jacques Renaut
IBM Informix Advanced Support
APD Team
OK, if you are talking about the Huge pages problem, here's what you do:
cat /proc/meminfo|fgrep -i huge
That will tell you what size in KB huge pages are defined on your system
(usually 2048K). On my system I get:
# cat /proc/meminfo|fgrep -i hugeAnonHugePages: 0 kB
HugePages_Total: 24
HugePages_Free: 24
HugePages_Rsvd: 0
HugePages_Surp: 0
Hugepagesize: 2048 kB
So I have 24 2MB huge pages defined. You can modify the number of pages
(assuming you have enough memory) by writing the desired number to
/proc/sys/vm/nr_hugepages
Your server wants 2GB of huge pages, so:
echo 1024 > /proc/sys/vm/nr_hugepages
Should do the trick. You can make the change permanent by adding the
following lines to the /etc/sysctl.conf file:
vm.nr_hugepages = 1024
vm.nr_hugepages_mempolicy = 1024
If too much memory is already committed, you may not get all of the 1GB you
are requesting so it is best to try that with the Informix server(s)
offline or to change the sysctl.conf file and reboot the machine.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Mon, Feb 18, 2013 at 5:34 PM, JAVIER HERNáNDEZ
<javier_hdzt@yahoo.com.mx>wrote:
> I have IDS 11.70 FC4 intalled in RedHat 5.6 and we are having a recurrent
> problem with one of our instances
>
> This morning we got the following error that make us restart the instance,
> the
> online log entrances follow:
>
> Mon Feb 18 07:59:11 2013
>
> 07:59:11 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
> Network connection is broken.
>
> 07:59:37 Checkpoint Completed: duration was 1 seconds.
> 07:59:37 Mon Feb 18 - loguniq 190, logpos 0x4817108, timestamp: 0xbd95965a
> Interval: 34917
>
> 07:59:37 Maximum server connections 621
> 07:59:37 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked
> 0,
> Plog used 7490, Llog used 16131
>
> 07:59:38 IBM Informix Dynamic Server Stopped.
>
> 07:59:43 IBM Informix Dynamic Server Started.
> 07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
> segment.
>
> Requested: 2629509120 bytes. Available: 0 bytes.
>
> The default memory page size will be used.
> 07:59:43 Segment locked: addr=0x44000000, size=2629509120
> 07:59:43 Insufficient free huge pages in /proc/meminfo for shared memory
> segment.
>
> Requested: 2147483648 bytes. Available: 0 bytes.
>
> The default memory page size will be used.
> 07:59:43 Segment locked: addr=0xe0bb2000, size=2147483648
>
> Mon Feb 18 07:59:45 2013
>
> 07:59:45 Event alarms enabled. ALARMPROG =
> '/opt/IBM/informix/etc/alarmprogram.sh'
> 07:59:45 Booting Language <c> from module <>
> 07:59:45 Loading Module <CNULL>
> 07:59:45 Booting Language <builtin> from module <>
> 07:59:45 Loading Module <BUILTINNULL>
> 07:59:51 DR: DRAUTO is 0 (Off)
> 07:59:51 DR: ENCRYPT_HDR is 0 (HDR encryption Disabled)
> 07:59:51 Event notification facility epoll enabled.
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 Could not disable priority aging: errno = 4
> 07:59:51 IBM Informix Dynamic Server Version 11.70.FC4 Software Serial
> Number
> AAA#B000000
> 07:59:51 Performance Advisory: The physical log size is smaller than the
> recommended size for a
> server configured with RTO_SERVER_RESTART.
> 07:59:51 Results: Fast recovery performance might not be optimal.
> 07:59:51 Action: For best fast recovery performance when
> RTO_SERVER_RESTART is
> enabled,> increase the physical log size to at least 2306867 KB. For servers
> configured with a large buffer pool, this might not be necessary.
> 07:59:51 IBM Informix Dynamic Server Initialized -- Shared Memory
> Initialized.
>
> 07:59:51 Started 1 B-tree scanners.
> 07:59:51 B-tree scanner threshold set at 5000.
> 07:59:51 B-tree scanner range scan size set to -1.
> 07:59:51 B-tree scanner ALICE mode set to 6.
> 07:59:51 B-tree scanner index compression level set to med.
> 07:59:51 Physical Recovery Started at Page (1:33083).
> 07:59:52 Physical Recovery Complete: 255 Pages Examined, 255 Pages
> Restored.
> 07:59:52 Logical Recovery Started.
> 07:59:52 10 recovery worker threads will be started.
> 07:59:53 Logical Recovery has reached the transaction cleanup phase.
> 07:59:53 Logical Recovery Complete.
>
> 0 Committed, 1 Rolled Back, 0 Open, 0 Bad Locks
>
> 07:59:54 Dataskip is now OFF for all dbspaces
> 07:59:54 requested number of KAIO events (8192) exceeds limit (6144). using
> 6144.
> 07:59:55 Checkpoint Completed: duration was 1 seconds.
> 07:59:55 Mon Feb 18 - loguniq 190, logpos 0x48192c8, timestamp: 0xbd9594aa
> Interval: 34918
>
> 07:59:55 Maximum server connections 0
> 07:59:55 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked
> 0,
> Plog used 272, Llog used 1
>
> 07:59:55 On-Line Mode
> 07:59:56 SCHAPI: Started dbScheduler thread.
> 07:59:57 Booting Language <spl> from module <>
> 07:59:57 Loading Module <SPLNULL>
> 07:59:57 Auto Registration is synced
> 07:59:57 SCHAPI: Started 2 dbWorker threads.
> 07:59:59 Defragmenter cleaner thread now running
> 07:59:59 Defragmenter cleaner thread cleaned:0 partitions
> 08:00:56 Loading Module <$INFORMIXDIR/extend/ifxmngr/ifxmngr.bld>
> 08:00:56 The C Language Module
> </opt/IBM/informix/extend/ifxmngr/ifxmngr.bld>
> loaded
>
> The configuration of the instance is as follow:
>
> RESIDENT -1
> SHMBASE 0x44000000L
> SHMVIRTSIZE 2097152
> SHMADD 32768
> EXTSHMADD 32768
> SHMTOTAL 6291456
>
> MULTIPROCESSOR 1
> VPCLASS cpu,num=4,noage
> VPCLASS soc,num=1,noage
> VP_MEMORY_CACHE_KB 0
> SINGLE_CPU_VP 0
>
> BUFFERPOOL
> size=2K,buffers=1048576,lrus=32,lru_min_dirty=50.00,lru_max_dirty=60.00>
> I can't find the reason for this problem, can any of you point me in the
> right
> direction?
>
> I'll be very greatfull.
>
> Best Regards.
>
> Javier Hernández
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--bcaec54b48d069beaf04d614da30
This is the output of onstat -g seg as of 19/02/2013 9:22 CST
IBM Informix Dynamic Server Version 11.70.FC4 -- On-Line -- Up 16:08:10 --5850380 Kbytes
Segment Summary:
id key addr size ovhd class blkused blkfree
9437191 52564801 44000000 3843305472 45472304 R* 938305 2
9469960 52564802 129143000 2147483648 25167624 V* 111927 412361
Total: - - 5990789120 - - 1050232 412363
(* segment locked in memory)
No reserve memory is allocated
The proble is that after the ASF error apeared the server stoped answering requests and needed to be bounced in order to function again, it's not the first time that it happens to us and we are trying to get it fixed for real this time.
Regretfully the server didn't generated a stack trace, it only stoped answering requests and we bounced it manually in order to fix the problem. We'll plan on doing the upgrade ASAP as you recomend.
Original post: The proble is that after the ASF error apeared the server stoped answering requests and needed to be bounced in order to function again, it's not the first time that it happens to us and we are trying to get it fixed for real this time. Response: What you are describing doesn't sound like a simple issue, so I would recommend opening a PMR with support. Jacques Renaut IBM Informix Advanced Support APD Team
Related threads
- System Or Internal Error InterruptedIOException
- 25582 error on high volume of short live trans
- ASF Echo-Thread Server: asfcode = -25582 oserr = 4