Problems on SUN with chunks being marked down
Posted in 2008
A DBA running six IDS 10 instances on Solaris with raw devices saw chunks repeatedly marked down with system error 2, including restore failures, while the Unix admin reported no disk errors; moving some LVs from local to SAN disk seemed to help, which puzzled everyone. Suggestions were: check raw-device offsets (offset 0), check that Solaris wasn't resetting raw device permissions at boot (errno 2 is usually permissions — the poster said it wasn't), and trace every dbspace/chunk through links, raw disks and LUN mappings looking for overlaps or size mismatches, plus a theory about mixed disk speeds. No resolution is recorded; the thread drifts into jokes.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Storage & Space Management
HELP, has anyone seen this before? After many years of only using HP, we put some Informix on Sun last year. We have experienced odd issues with chunks being marked off-line. The Unix admin tells us there are no disk errors. Informix reports system error 2. On one occasion (4 weeks ago) the Unix admin says all he did was add space to an unrelated file system (we use only raw spaces). I was then able to bring all the instances back on-line except the HDR. We had to mark the spaces on-line manually and then like went on for 3 weeks with no issues. Then it happened again. This time Informix dialed-in and marked the spaces on-line themselves. This was done over and over until all but one stayed up. But as soon as he logged out they all went down again except 1 instance. In the meantime we started a restore on the instance that would not come up at all. The restore failed about 6 times with error 2 also. This time we asked the Unix SA to moved some of our LV's from local disk to SAN disk. All the instances have mixed local and SAN disk. Informix thought it could be the difference between the disk speed. That worked but it makes no sense. He moved one space from each instance to SAN disk but not all. The restore then worked. This is only happening on one system. The HDR secondary has no issues. The Unix admin will not open a case with SUN because he says he has nothing to tell them. There are 6 instances on this system. v10.FC6 and v10.FC8X7 shm0 is HDR and came up after one lv was moved to SAN disk. shm1 is the one that came up and stayed up after Informix marked on on-line and no spaces were moved to SAN disk shm2 - shm5 all had one LV moved from local disk to SAN disk. There are more odd occurrences on other SUN servers but they have not reoccurred. On one the Unix admin says, "it looks like Informix is trying to take more raw space that it is allocated and it filled up the LUN to 100%. I moved things around to free up space on the LUN". Our eye brows are still raised on this one but it has not happened again! Thanks for any help you can provide. Alyx
Out of curiority, are you using an offset of 0 on your raw spaces? j. On Thu, Aug 14, 2008 at 3:58 PM, ALYX CHAVIS wrote: > HELP, has anyone seen this before? After many years of only using HP, we put some Informix on Sun last year. We have experienced odd issues with chunks being marked off-line. The Unix admin tells us there are no disk errors. Informix reports system error 2. On one occasion (4 weeks ago) the Unix admin says all he did was add space to an unrelated file system (we use only raw spaces). I was then able to bring all the instances back on-line except the HDR. We had to mark the spaces on-line manually and then like went on for 3 weeks with no issues. Then it happened again. This time Informix dialed-in and marked the spaces on-line themselves. This was done over and over until all but one stayed up. But as soon as he logged out they all went down again except 1 instance. In the meantime we started a restore on the instance that would not come up at all. The restore failed about 6 times with error 2 also. This time we asked the Unix SA to moved some of our LV's from local disk to SAN disk. All the instances have mixed local and SAN disk. Informix thought it could be the difference between the disk speed. That worked but it makes no sense. He moved one space from each instance to SAN disk but not all. The restore then worked. This is only happening on one system. The HDR secondary has no issues. The Unix admin will not open a case with SUN because he says he has nothing to tell them. There are 6 instances on this system. v10.FC6 and v10.FC8X7 shm0 is HDR and came up after one lv was moved to SAN disk. shm1 is the one that came up and stayed up after Informix marked on on-line and no spaces were moved to SAN disk shm2 - shm5 all had one LV moved from local disk to SAN disk. There are more odd occurrences on other SUN servers but they have not reoccurred. On one the Unix admin says, "it looks like Informix is trying to take more raw space that it is allocated and it filled up the LUN to 100%. I moved things around to free up space on the LUN". Our eye brows are still raised on this one but it has not happened again! Thanks for any help you can provide. Alyx ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Errno 2 is usually a permissions problem. Note that Solaris may rebuild certain devices at boot up with default permissions (usually owner root, group sys, privs 600) so your rc scripts may have to repermission the raw devices during the restart before starting IDS. Art On Thu, Aug 14, 2008 at 3:58 PM, ALYX CHAVIS <alyx_chavis@circuitcity.com>wrote: > HELP, has anyone seen this before? > > After many years of only using HP, we put some Informix on Sun last year. > We > have experienced odd issues with chunks being marked off-line. The Unix > admin > tells us there are no disk errors. Informix reports system error 2. > > On one occasion (4 weeks ago) the Unix admin says all he did was add space > to > an unrelated file system (we use only raw spaces). I was then able to bring > all the instances back on-line except the HDR. We had to mark the spaces > on-line manually and then like went on for 3 weeks with no issues. > > Then it happened again. This time Informix dialed-in and marked the spaces > on-line themselves. This was done over and over until all but one stayed > up. > But as soon as he logged out they all went down again except 1 instance. In > the meantime we started a restore on the instance that would not come up at > all. The restore failed about 6 times with error 2 also. > This time we asked the Unix SA to moved some of our LV's from local disk to > SAN disk. All the instances have mixed local and SAN disk. Informix thought > it > could be the difference between the disk speed. That worked but it makes no > sense. He moved one space from each instance to SAN disk but not all. The > restore then worked. > > This is only happening on one system. The HDR secondary has no issues. The > Unix admin will not open a case with SUN because he says he has nothing to > tell them. > > There are 6 instances on this system. v10.FC6 and v10.FC8X7 > shm0 is HDR and came up after one lv was moved to SAN disk. > shm1 is the one that came up and stayed up after Informix marked on on-line > and no spaces were moved to SAN disk > shm2 - shm5 all had one LV moved from local disk to SAN disk. > > There are more odd occurrences on other SUN servers but they have not > reoccurred. On one the Unix admin says, "it looks like Informix is trying > to > take more raw space that it is allocated and it filled up the LUN to 100%. > I > moved things around to free up space on the LUN". Our eye brows are still > raised on this one but it has not happened again! > > Thanks for any help you can provide. Alyx > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > -- Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves.
Yes, all of our servers use 0. I know Sun use to have an issue with this as did DEC. But that was a long time ago. Do you know something I don't? This system has been operational on SUN for over 1 yr and we just started having these issues 1.5 months ago. Thanks, Alyx
It's not a permission's issue and a reboot was not involved either time. That was the very first thing I checked anyway just in case. :-( The disk showed no errors, no waits, no queue lengths, nothing! Thanks,Alyx
Perhaps I'm just dating myself. j. Sane ego te vocavi. Forsitan capedictum tuum desit. -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org]On Behalf Of ALYX CHAVIS Sent: Thursday, August 14, 2008 5:06 PM To: ids@iiug.org Subject: Re: RE: Problems on SUN with chunks being mark.... [13112] Yes, all of our servers use 0. I know Sun use to have an issue with this as did DEC. But that was a long time ago. Do you know something I don't? This system has been operational on SUN for over 1 yr and we just started having these issues 1.5 months ago. Thanks, Alyx **************************************************************************** *** Forum Note: Use "Reply" to post a response in the discussion forum.
It is much more enjoyable to date another person -- you learn new things :) Take care. Clifton M. Bean Informix DBA / AIX System Admin Currency Technics & Metrics Main (972) 812-1411 x244 -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of Jack Parker Sent: Thursday, August 14, 2008 7:29 PM To: ids@iiug.org Subject: RE: RE: Problems on SUN with chunks being mark.... [13114] Perhaps I'm just dating myself. j. Sane ego te vocavi. Forsitan capedictum tuum desit. -----Original Message----- From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org]On Behalf Of ALYX CHAVIS Sent: Thursday, August 14, 2008 5:06 PM To: ids@iiug.org Subject: Re: RE: Problems on SUN with chunks being mark.... [13112] Yes, all of our servers use 0. I know Sun use to have an issue with this as did DEC. But that was a long time ago. Do you know something I don't? This system has been operational on SUN for over 1 yr and we just started having these issues 1.5 months ago. Thanks, Alyx ************************************************************************ **** *** Forum Note: Use "Reply" to post a response in the discussion forum. ************************************************************************ ******* Forum Note: Use "Reply" to post a response in the discussion forum.
2008/8/14 ALYX CHAVIS <alyx_chavis@circuitcity.com>: > It's not a permission's issue and a reboot was not involved either time. That > was the very first thing I checked anyway just in case. :-( > > The disk showed no errors, no waits, no queue lengths, nothing! > > Thanks,Alyx > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > Alyx Just some lateral random thoughts that might help. As you have several instances on the same box it might be worth tracking each (and every) dbspace through its chunks to the file system links (if you use them) and down to the raw disk and the LUN mappings just to ensure you havn't got any crossed links or are trying to use the same disk space in more than one instance and that all disk space sizes are consistant 'up the tree'. I have done this myself on AIX but with different results !! :-(( I don't see how informix can write beyond the end of a LUN. When a chunk is created infomrix checks the space is available and writable, if the LUN has been created at a certain size it will present that size to the O/S and should not allow writing outside that area (unless there is something very wrong with the SAN software !!). The disk speed question makes possible sense. I don't know the check informix undertakes as it starts, but f it expects all chunks to respond in a similar time-frame, and one or two are a bit slow (due to slow disk) I could see how they could be marked as down. As rootdbs would be the first accessed (and possibly set a standard for all other responses) you could try putting that on the slowest disk to see if that helps. Keith
Thanks, random thoughts can be very good. Alyx
ALYX CHAVIS wrote: > Thanks, random thoughts can be very good. Alyx > > Have you tried UPDATE STATISTICS? -- Cheers, Obnoxio the Clown http://obotheclown.blogspot.com
On Mon, Aug 18, 2008 at 3:24 PM, Obnoxio The Clown <obnoxio@serendipita.com> wrote: > ALYX CHAVIS wrote: >> Thanks, random thoughts can be very good. Alyx >> > Have you tried UPDATE STATISTICS? > You need a new random thought generator, OTC :D -- Jonathan Leffler #include <disclaimer.h> Email: jleffler@earthlink.net, jleffler@us.ibm.com Guardian of DBD::Informix v2008.0513 -- http://dbi.perl.org/ "Blessed are we who can laugh at ourselves, for we shall never cease to be amused." NB: Please do not use this email for correspondence. I don't necessarily read it every week, even.