Storage Area Network Fails for Unix/Netware
Posted in 2007
A user running IDS 9.40.UC8 on RHEL 3 with chunks on an HP SAN (QLogic QLA2300 HBA, SecurePath failover) reported that whenever a SAN controller failed, the whole Linux box froze and only a power cycle recovered it; Windows hosts failed over fine. The online.log showed an I/O error on the logical-log chunk, 'Critical media failure' and an engine PANIC. Suggestions: determine whether Linux or just IDS is hung (try killing oninit), test I/O with dd/cp while flipping HBA channels to isolate a kernel/driver issue, check ONCONFIG chunk-offline behaviour, and consider RAID10 rather than RAID5. No resolution is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Platform-Specific Issues
We want to run our Informix 9.40 database on a SAN, and we've tried this using an HBA fiber-optic (dual channel) card, with HP Secure-Path failover software. This has worked, to a point, except that each time a SAN controller fails for any reason, the Linux-based Informix server has simply locked up! The only way to get Informix going again is to suddenly shut down the power to the Linux server... even Ctrl-Alt-Delete won't work, as will no other Ctrl-X command work either (Ctrl-C, Ctrl-D, Ctrl-Z... no keyboard strokes will return control to the user. Attaching a Windows system to the SAN is never a problem. If the controller fails, and then comes back, the Windows system reattaches as if nothing had happened, because it continues to poll the controller while it's down. Netware fails, perhaps because it ceases to poll the controller. However, according to HP, our Linux (Enterprise 3) should be able to failover and reattach... but it doesn't! Of course, we can't tell whether 1) the HBA card is failing to flip channels, to pick up the other LUN, or 2) the Linux O.S. has a problem making this failover happen, or 3) perhaps it's the Informix engine that's incapable of a small burp of this nature. Has anyone else attempted a SAN database connection like this? O.S.: Linux Enterprise 3, 2.4.21-40.ELsmp #1 SMP Hardware: HP 380 G4, 4GB RAM Informix: IDS Version 9.40.UC8 HBA card: Qlogic QLA2300 Failover: SecurePath 3.0CFullUpdate-4.0.SP2
CHARLES KIRBY wrote: > even Ctrl-Alt-Delete won't work > Thanks. I needed a laugh. -- This message has been scanned for viruses and dangerous content by OpenProtect(http://www.openprotect.com), and is believed to be clean.
You mean the whole Linux system is frozen and does not react to any user input
, not even on the console ?
Or just IDS does not work anymore and does not allow user interaction or
connections or anything?
If the first is true, I'd say it is not IDS but Linux and or the HBA. You
might want to try to provoke the problem when IDS is offline .
If the latter is true, you may try to kill -11 or (if that does not work kill
-9) the primary oninit process (ParentPID 1) instead shutting doen Linux box
the hard way.
Rgds
Tilman
Very funny, och aye le noo! Oi'll be sic'n old Nessie on ya! Tha' beastie in Loch Ness has been around agin! Nah, wotcha know aboot using Informix on a SAN?
CHARLES KIRBY wrote: > Very funny, och aye le noo! > > Oi'll be sic'n old Nessie on ya! > Tha' beastie in Loch Ness has been around agin! > > Nah, wotcha know aboot using Informix on a SAN? > > It's hardware. I don't do hardware. Hoots mon! Tattie scone? -- This message has been scanned for viruses and dangerous content by OpenProtect(http://www.openprotect.com), and is believed to be clean.
Right! No response on the console whatsoever, so it's impossible to do "kill -9" or "kill -kill" (what's kill -11?). I don't think it was possible to log in remotely, either, the last time it happened. We did try replacing the dual-channel HBA card (with another brand new one), as well as replacing the entire fibre-optic cable, and these things didn't help either. I think it happened originally with a single-channel HBA card, also. This has never happened with IDS offline. However, only the one database was in use on the SAN, so there's essentially no interaction with the SAN once IDS is shut down.
CHARLES KIRBY wrote: > Right! No response on the console whatsoever, so it's impossible to do "kill > -9" or "kill -kill" (what's kill -11?). I don't think it was possible to log > in remotely, either, the last time it happened. > > We did try replacing the dual-channel HBA card (with another brand new one), > as well as replacing the entire fibre-optic cable, and these things didn't > help either. I think it happened originally with a single-channel HBA card, > also. > > This has never happened with IDS offline. However, only the one database was > in use on the SAN, so there's essentially no interaction with the SAN once IDS > is shut down. > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > Perhaps you can provide some more details .... online.log from just after the freeze.. onconfig how are the dbspaces/chunks created. ( underlying filesystem etc...) rgds, Peter
Here we are... more than enough information in the online.log;
Certainly, you can see it lost contact with the logical logs drive space:
I wonder if Informix is overly quick to declare a dbspace off limits? Perhaps
the time it takes for SecurePath to flip still isn't adequate for Informix!
BTW: dbspaces are created in a tyical fashion, and linked with soft links.
02:40:54 Maximum server connections 42
08:44:10 Assert Warning: I/O error, Primary Chunk
'/opt/informix/links/nysd/llog_db1' -- Offline
08:44:10 IBM Informix Dynamic Server Version 9.40.UC8
08:44:10 Who: Session(19694, root@nysdlei2.nysd.circ2.dcn, 13784, 0x5aaf6df8)
Thread(19759, sqlexec, 5aace930, 1)
File: rsbuff.c Line: 4742
08:44:10 Results: Chunk is now unusable
08:44:10 Action: Repair and restore from mirror or archive
08:44:10 stack trace for pid 3206 written to /var/log/informix/af.51172a2a
08:44:10 See Also: /var/log/informix/af.51172a2a
08:44:13 I/O error, Primary Chunk '/opt/informix/links/nysd/llog_db1' --
Offline
08:44:14 Assert Failed: INFORMIX-OnLine Must ABORT
Critical media failure.
08:44:14 IBM Informix Dynamic Server Version 9.40.UC8
08:44:14 Who: Session(19694, root@nysdlei2.nysd.circ2.dcn, 13784, 0x5aaf6df8)
Thread(19759, sqlexec, 5aace930, 3)
File: rsmirror.c Line: 1812
08:44:14 stack trace for pid 3208 written to /var/log/informix/af.51172a2a
08:44:14 See Also: /var/log/informix/af.51172a2a
08:44:23 rsmirror.c, line 1812, thread 19759, proc id 3208, INFORMIX-OnLine
Must ABORT
Critical media failure..
08:44:23 Fatal error in ADM VP at mt.c:12419
08:44:23 Unexpected virtual processor termination, pid = 3208, exit = 0x100
08:44:24 PANIC: Attempting to bring system down
08:44:24 semctl: errno = 22
08:44:24 semctl: errno = 22
ROOTNAME rootdbs # Root dbspace name
ROOTPATH /opt/informix/links/nysd/root_db1
ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
ROOTSIZE 600000 # Size of root dbspace (Kbytes)
PHYSDBS physlogdb # Location (dbspace) of physical log
PHYSFILE 100000 # Physical log file size (Kbytes)
LOGFILES 16 # Number of logical log files
LOGSIZE 10000 # Logical log size (Kbytes)
SERVERNUM 10 # Unique id corresponding to a OnLine instance
DBSERVERNAME nysd_ecf # List of alternate dbservernames
DBSERVERALIASES nysd_shm # Name of default database server
NETTYPE soctcp,2,128,CPU # Configure poll thread(s) for nettype
NETTYPE ipcshm,1,128,CPU # Configure poll thread(s) for nettype
DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
NUMCPUVPS 3 # Number of user (cpu) vps
SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
NOAGE 0 # Process aging
AFF_SPROC 0 # Affinity start processor
AFF_NPROCS 0 # Affinity number of processors
LOCKS 550000 # Maximum number of locks
BUFFERS 150000 # Maximum number of shared buffers
NUMAIOVPS 32 # Number of IO vps
PHYSBUFF 90 # Physical log buffer size (Kbytes)
LOGBUFF 64 # Logical log buffer size (Kbytes)
CLEANERS 16 # Number of buffer cleaner processes
SHMBASE 0x44000000 # Shared memory base address
SHMVIRTSIZE 128000 # initial virtual shared memory segment size
SHMADD 64000 # Size of new shared memory segments (Kbytes)
SHMTOTAL 1200000 # Total shared memory (Kbytes). 0=>unlimited
CKPTINTVL 1800 # Check point interval (in sec)
LRUS 128 # Number of LRU queues
LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning limit
LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 32 # Stack size (Kbytes)
LTXHWM 50 # Long transaction high water mark percentage
LTXEHWM 60 # Long transaction high water mark (exclusive)
CHARLES KIRBY wrote:
> Here we are... more than enough information in the online.log;
> Certainly, you can see it lost contact with the logical logs drive space:
>
> I wonder if Informix is overly quick to declare a dbspace off limits? Perhaps
> the time it takes for SecurePath to flip still isn't adequate for Informix!
>
> BTW: dbspaces are created in a tyical fashion, and linked with soft links.
>
> 02:40:54 Maximum server connections 42
> 08:44:10 Assert Warning: I/O error, Primary Chunk
> '/opt/informix/links/nysd/llog_db1' -- Offline
> 08:44:10 IBM Informix Dynamic Server Version 9.40.UC8
> 08:44:10 Who: Session(19694, root@nysdlei2.nysd.circ2.dcn, 13784, 0x5aaf6df8)
>
> Thread(19759, sqlexec, 5aace930, 1)
>
> File: rsbuff.c Line: 4742
> 08:44:10 Results: Chunk is now unusable
> 08:44:10 Action: Repair and restore from mirror or archive
> 08:44:10 stack trace for pid 3206 written to /var/log/informix/af.51172a2a
> 08:44:10 See Also: /var/log/informix/af.51172a2a
> 08:44:13 I/O error, Primary Chunk '/opt/informix/links/nysd/llog_db1' --
> Offline
> 08:44:14 Assert Failed: INFORMIX-OnLine Must ABORT
>
> Critical media failure.
> 08:44:14 IBM Informix Dynamic Server Version 9.40.UC8
> 08:44:14 Who: Session(19694, root@nysdlei2.nysd.circ2.dcn, 13784, 0x5aaf6df8)
>
> Thread(19759, sqlexec, 5aace930, 3)
>
> File: rsmirror.c Line: 1812
> 08:44:14 stack trace for pid 3208 written to /var/log/informix/af.51172a2a
> 08:44:14 See Also: /var/log/informix/af.51172a2a
> 08:44:23 rsmirror.c, line 1812, thread 19759, proc id 3208, INFORMIX-OnLine
> Must ABORT
>
> Critical media failure..
> 08:44:23 Fatal error in ADM VP at mt.c:12419
> 08:44:23 Unexpected virtual processor termination, pid = 3208, exit = 0x100
>
> 08:44:24 PANIC: Attempting to bring system down
> 08:44:24 semctl: errno = 22
>
> 08:44:24 semctl: errno = 22
>
> ROOTNAME rootdbs # Root dbspace name
> ROOTPATH /opt/informix/links/nysd/root_db1
> ROOTOFFSET 0 # Offset of root dbspace into device (Kbytes)
> ROOTSIZE 600000 # Size of root dbspace (Kbytes)>
> PHYSDBS physlogdb # Location (dbspace) of physical log
> PHYSFILE 100000 # Physical log file size (Kbytes)
>
> LOGFILES 16 # Number of logical log files
> LOGSIZE 10000 # Logical log size (Kbytes)
>
> SERVERNUM 10 # Unique id corresponding to a OnLine instance
> DBSERVERNAME nysd_ecf # List of alternate dbservernames
> DBSERVERALIASES nysd_shm # Name of default database server
> NETTYPE soctcp,2,128,CPU # Configure poll thread(s) for nettype
> NETTYPE ipcshm,1,128,CPU # Configure poll thread(s) for nettype
> DEADLOCK_TIMEOUT 60 # Max time to wait of lock in distributed env.
> RESIDENT 0 # Forced residency flag (Yes = 1, No = 0)
>
> MULTIPROCESSOR 1 # 0 for single-processor, 1 for multi-processor
> NUMCPUVPS 3 # Number of user (cpu) vps
> SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps to one
>
> NOAGE 0 # Process aging
> AFF_SPROC 0 # Affinity start processor
> AFF_NPROCS 0 # Affinity number of processors
>
> LOCKS 550000 # Maximum number of locks
> BUFFERS 150000 # Maximum number of shared buffers
> NUMAIOVPS 32 # Number of IO vps
> PHYSBUFF 90 # Physical log buffer size (Kbytes)
> LOGBUFF 64 # Logical log buffer size (Kbytes)
> CLEANERS 16 # Number of buffer cleaner processes
> SHMBASE 0x44000000 # Shared memory base address
> SHMVIRTSIZE 128000 # initial virtual shared memory segment size
> SHMADD 64000 # Size of new shared memory segments (Kbytes)
> SHMTOTAL 1200000 # Total shared memory (Kbytes). 0=>unlimited
> CKPTINTVL 1800 # Check point interval (in sec)
> LRUS 128 # Number of LRU queues
> LRU_MAX_DIRTY 2.000000 # LRU percent dirty begin cleaning limit
> LRU_MIN_DIRTY 1.000000 # LRU percent dirty end cleaning limit
> TXTIMEOUT 0x12c # Transaction timeout (in sec)
> STACKSIZE 32 # Stack size (Kbytes)
> LTXHWM 50 # Long transaction high water mark percentage
> LTXEHWM 60 # Long transaction high water mark (exclusive)>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
Are all your filesystems on the SAN ? I mean your standard linux
filesytems as well not just your dbspaces/chunks?
Do you have any applications next to IDS doing any I/O on the SAN disks ?
To do a simple test ... just do a dd or cp to a location on one of
your SAN filesystem and change the channels on the HBA card...If that
fails then you know you have an I/O issue during channel switch of the
fibre controller and which is mostlikely an linux kernel /driver issue..
If not then from the top of my head....there is a setting the $ONCONFIG
where you can instruct IDS what to do when a chunk gets offline
...(could be wrong for version 9.xx though ))
HTH,
Peter
I've seen this happen on systems where the SAN was set up with RAID5 but never when RAID10 was the SAN configuration. Could that be the problem?? Art S. Kagel ----- Original Message ----- From: Charles Kirby <ids@iiug.org> At: 5/31 10:31:45 We want to run our Informix 9.40 database on a SAN, and we've tried this using an HBA fiber-optic (dual channel) card, with HP Secure-Path failover software. This has worked, to a point, except that each time a SAN controller fails for any reason, the Linux-based Informix server has simply locked up! The only way to get Informix going again is to suddenly shut down the power to the Linux server... even Ctrl-Alt-Delete won't work, as will no other Ctrl-X command work either (Ctrl-C, Ctrl-D, Ctrl-Z... no keyboard strokes will return control to the user. Attaching a Windows system to the SAN is never a problem. If the controller fails, and then comes back, the Windows system reattaches as if nothing had happened, because it continues to poll the controller while it's down. Netware fails, perhaps because it ceases to poll the controller. However, according to HP, our Linux (Enterprise 3) should be able to failover and reattach... but it doesn't! Of course, we can't tell whether 1) the HBA card is failing to flip channels, to pick up the other LUN, or 2) the Linux O.S. has a problem making this failover happen, or 3) perhaps it's the Informix engine that's incapable of a small burp of this nature. Has anyone else attempted a SAN database connection like this? O.S.: Linux Enterprise 3, 2.4.21-40.ELsmp #1 SMP Hardware: HP 380 G4, 4GB RAM Informix: IDS Version 9.40.UC8 HBA card: Qlogic QLA2300 Failover: SecurePath 3.0CFullUpdate-4.0.SP2 ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.