database meltdown on RAID
Posted in 1999
Topics: High Availability & Replication, Installation, Setup & Upgrades, Storage & Space Management, Platform-Specific Issues
We have a serious problem setting up two database servers with Informix Dynamic Server 7.31.uc2xf using raw (non-unix) disk space on RAID 5 devices. Root dbspace pages get corrupted intermittantly sometimes crashing the DB immediately, and sometimes just damaging it so that it can't be restarted if brought down. The corruption is sometimes extensive, like all (2k) pages from 4-100 set to zeroes, and sometimes more subtle. I am trying to set up Enterprise replication between the two servers, but don't see a clear pattern of ER causing the problem. The RAID is an Adjile Systems SCSI device. The disk format info shows: <CMDTECH-CRD-5440-1-C1-8 cyl 26047 alt 2 hd 16 sec 64> /sbus@3,0/SUNW,fas@3,8800000/sd@0,0 The corruption only occurrs in dbspace defined in partition 1 of the disk. A second DB instance defined in partition 2 on the same disk is rock steady. Our system administrator has applied all the Solaris 2.6 patches he could find, and upgraded the firmware in the RAID. Informix tech support says it's hardware - not their problem. The hardware vendors say it's Informix. If anyone has any ideas, or knows anybody who might have an idea please send me a note or post a reply. Thanks, Jay Haley Sent via Deja.com http://www.deja.com/ Before you buy.
In article <7ulji9$ur4$1@nnrp1.deja.com>, jshaley@my-deja.com writes >We have a serious problem setting up two database servers >with Informix Dynamic Server 7.31.uc2xf using raw (non-unix) disk space >on RAID 5 devices. Root dbspace pages get corrupted intermittantly >sometimes crashing the DB immediately, and sometimes just damaging it >so that it can't be restarted if brought down. > >The corruption is sometimes extensive, like all (2k) pages from 4-100 >set to zeroes, and sometimes more subtle. > >I am trying to set up Enterprise replication between the two servers, >but don't see a clear pattern of ER causing the problem. > >The RAID is an Adjile Systems SCSI device. The disk format info shows: > <CMDTECH-CRD-5440-1-C1-8 cyl 26047 alt 2 hd 16 sec 64> > /sbus@3,0/SUNW,fas@3,8800000/sd@0,0 > >The corruption only occurrs in dbspace defined in partition 1 >of the disk. A second DB instance defined in partition 2 on the >same disk is rock steady. > >Our system administrator has applied all the Solaris 2.6 patches he >could >find, and upgraded the firmware in the RAID. Informix tech support says >it's hardware - not their problem. The hardware vendors say it's >Informix. > >If anyone has any ideas, or knows anybody who might have >an idea please send me a note or post a reply. > >Thanks, Jay Haley > > >Sent via Deja.com http://www.deja.com/ >Before you buy. He have seen Online corruption on ICL machines..try setting KAIO_OFF=1 before starting online. We found KAIO+Logical Volume Manager does not work... PS We had a Solaris 2.6 box which got a corrupted root filesysytem after upgrading Informix recently...Restoring from backup+reupgrading fixed it..(??!!!). Try turning off KAIO as above.. -- David Williams
I believe this is a common issue with Sun. If you *know* where the
first 10K of disk is in the 'pseudo-disks' created by your volume/RAID
manager, skip it using the 'offset' part of onspaces. I've been told in
every Informix admin class to skip at least the first 4K of disk on
Sun systems (vtoc information?).
And just to beat a dead horse, Informix *always* recommends against
RAID5. Use hardware mirroring if possible. Good luck.
/bobr/
David Williams wrote:
>
> In article <7ulji9$ur4$1@nnrp1.deja.com>, jshaley@my-deja.com writes
> >We have a serious problem setting up two database servers
> >with Informix Dynamic Server 7.31.uc2xf using raw (non-unix) disk space
> >on RAID 5 devices. Root dbspace pages get corrupted intermittantly
> >sometimes crashing the DB immediately, and sometimes just damaging it
> >so that it can't be restarted if brought down.
> >
> >The corruption is sometimes extensive, like all (2k) pages from 4-100
> >set to zeroes, and sometimes more subtle.
> >
> >I am trying to set up Enterprise replication between the two servers,
> >but don't see a clear pattern of ER causing the problem.
> >
> >The RAID is an Adjile Systems SCSI device. The disk format info shows:
> > <CMDTECH-CRD-5440-1-C1-8 cyl 26047 alt 2 hd 16 sec 64>
> > /sbus@3,0/SUNW,fas@3,8800000/sd@0,0
> >
> >The corruption only occurrs in dbspace defined in partition 1
> >of the disk. A second DB instance defined in partition 2 on the
> >same disk is rock steady.
> >
> >Our system administrator has applied all the Solaris 2.6 patches he
> >could
> >find, and upgraded the firmware in the RAID. Informix tech support says
> >it's hardware - not their problem. The hardware vendors say it's
> >Informix.
> >
> >If anyone has any ideas, or knows anybody who might have
> >an idea please send me a note or post a reply.
> >
> >Thanks, Jay Haley
> >
> >
> >Sent via Deja.com http://www.deja.com/
> >Before you buy.
>
> He have seen Online corruption on ICL machines..try setting
>
> KAIO_OFF=1
>
> before starting online.
>
> We found KAIO+Logical Volume Manager does not work...
>
> PS We had a Solaris 2.6 box which got a corrupted root filesysytem
> after upgrading Informix recently...Restoring from backup+reupgrading
> fixed it..(??!!!).
>
> Try turning off KAIO as above..
>
> --
> David Williams