Re: Raw Device IO Error
Posted in 2008
Topics: Backup & Restore, Storage & Space Management, Error Codes & Troubleshooting, Logging & Checkpoints, Platform-Specific Issues
Christian Schmidt wrote:
> Hi,
>
> I have a problem with my Storage Manager and Informix. Sometimes the
> Infomix DB crashes when I try to sync a mirrored rawdevice (not an
> informix mirrored chunck).
>
> Here is a passage from the online.log:
>
> 14:40:51 Logical Log 1320 - Backup Started
> 14:40:55 Logical Log 1320 - Backup Completed
> 14:41:07 Assert Warning: I/O error, Primary Chunk '/informix/baan/
> llogdbs' -- Offline
> 14:41:07 IBM Informix Dynamic Server Version 10.00.FC5
> 14:41:07 Who: Session(144, root@big-4, 9579, 10b0378e8)
> Thread(236, sqlexec, 10affeae8, 1)
> File: rsbuff.c Line: 5079
> 14:41:07 Results: Chunk is now unusable
> 14:41:07 Action: Repair and restore from mirror or archive
> 14:41:07 stack trace for pid 7429 written to /informix/baan/tmp/af.> 4d4d562
> 14:41:07 See Also: /informix/baan/tmp/af.4d4d562
> 14:41:16 I/O error, Primary Chunk '/informix/baan/llogdbs' -- Offline
> 14:41:17 Assert Failed: INFORMIX-OnLine Must ABORT
> Critical media failure.
> 14:41:17 IBM Informix Dynamic Server Version 10.00.FC5
> 14:41:17 Who: Session(144, root@big-4, 9579, 10b0378e8)
> Thread(236, sqlexec, 10affeae8, 1)
> File: rsmirror.c Line: 1847
> 14:41:17 stack trace for pid 7429 written to /informix/baan/tmp/af.> 4d4d562
> 14:41:17 See Also: /informix/baan/tmp/af.4d4d562
> 14:41:29 rsmirror.c, line 1847, thread 236, proc id 7429, INFORMIX-> OnLine Must ABORT
> Critical media failure..
> 14:41:30 The Master Daemon Died
> 14:41:30 PANIC: Attempting to bring system down>
> The rawdevice is accessible the whole time. But it can be, that the
> time taken for an IO operation is slightly longer. It seems to me,
> that Informix sends a signal (SIGTERM or SIGKILL) to the IO driver of
> the storage manager. There is a part in my messages file:
>
> Jul 22 14:41:05 big-4 vv: [ID 567783 kern.notice] av ixllogdbs_baan@0
> - vv_mirror: wait for clean track on ixllogdbs_baan@3 (66 / 66 / 66 -
> 1 / 0)
> Jul 22 14:41:07 big-4 vv: [ID 567783 kern.notice] NOTICE: av
> ixllogdbs_baan@0 - vv_devop_enter: waiting for operation 0x40
> interrupted
>
> It seems that there is an interrupt while the storage manager waits
> for a complete IO operation on the chunk with the logical logs.
>
> The system is Solaris 10 with Informix Dynamic Server Version
> 10.00.FC5.
>
> Is it possible to increase the IO timeout value for rawdevices?
>
>
> Thanks for any help!
> C. Schmidt
Post the top of the af file ...
http://www.research.ibm.com/journal/rd/524/hafner.html
When is an I/O not an I/O
> Post the top of the af file ... Have paste it here http://pastebin.com/m51f3b2a6
Hi, I think you should report this problem to IBM Informix Tech Support by opening a PMR. What looks somewhat strange to me is that it reports an errno of 4, which means interrupted system call. My impression was that this should not cause such a problem. You said during the mirror syncing the device is available all the time. But is this true for writing? Or is it only for reading? If the mirror sync procedure somehow sets some kind of "shared lock" on the raw device, then this could cause a problem when IDS tries to write to the raw device and gets an error due to the shared lock. (Though I'm not sure why/how that would map with the errno 4.) Regards, Martin -- Martin Fuerderer IBM Informix Development Munich, Germany Information Management IBM Deutschland Research & Development GmbH Chairman of the Supervisory Board: Martin Jetter Board of Management: Herbert Kircher Corporate Seat: Boeblingen, Germany Reg.-Gericht: Amtsgericht Stuttgart, HRB 243294 informix-list-bounces@iiug.org wrote on 23.07.2008 12:10:21: > > Post the top of the af file ... > > Have paste it here http://pastebin.com/m51f3b2a6 > > > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list
On 23 Jul., 13:24, Martin Fuerderer <MARTI...@de.ibm.com> wrote: > You said during the mirror syncing the device is > available all the time. But is this true for writing? > Or is it only for reading? It is true for both read and write. During the sync you can write on the volume but it could be that the response time for the volume is longer. The sync is done by the storage manager driver in the kernel mode. So Informix shouldn't see it. > If the mirror sync procedure somehow sets some > kind of "shared lock" on the raw device, then this > could cause a problem when IDS tries to write to > the raw device and gets an error due to the shared > lock. (Though I'm not sure why/how that would map > with the errno 4.) There are locks on the device but no shared locks or locks form the OS. The locks are made by the driver on the whole volume (even on the raw device). When Informix tries to write on a locked volume it would do it through the IO driver of the storage manager and the write operation will be hold back until the volume is unlocked. So I thought the problem could be an IO timeout. I don't know who sends the interrupt. The storage manager gets an interrupt too. (ixllogdbs_baan@0 - vv_devop_enter: waiting for operation 0x40 interrupted ) Regards, C. Schmidt > Regards, > Martin > -- > Martin Fuerderer > IBM Informix Development Munich, Germany > Information Management > > IBM Deutschland Research & Development GmbH > Chairman of the Supervisory Board: Martin Jetter > Board of Management: Herbert Kircher > Corporate Seat: Boeblingen, Germany > Reg.-Gericht: Amtsgericht Stuttgart, HRB 243294 >
Christian Schmidt wrote: > On 23 Jul., 13:24, Martin Fuerderer <MARTI...@de.ibm.com> wrote: >> You said during the mirror syncing the device is >> available all the time. But is this true for writing? >> Or is it only for reading? > > It is true for both read and write. During the sync you can write on > the volume but it could be that the response time for the volume is > longer. The sync is done by the storage manager driver in the kernel > mode. So Informix shouldn't see it. > >> If the mirror sync procedure somehow sets some >> kind of "shared lock" on the raw device, then this >> could cause a problem when IDS tries to write to >> the raw device and gets an error due to the shared >> lock. (Though I'm not sure why/how that would map >> with the errno 4.) > > There are locks on the device but no shared locks or locks form the > OS. The locks are made by the driver on the whole volume (even on the > raw device). When Informix tries to write on a locked volume it would > do it through the IO driver of the storage manager and the write > operation will be hold back until the volume is unlocked. So I thought > the problem could be an IO timeout. > > I don't know who sends the interrupt. The storage manager gets an > interrupt too. (ixllogdbs_baan@0 - vv_devop_enter: waiting for > operation 0x40 interrupted ) > > Regards, > C. Schmidt > > >> Regards, >> Martin >> -- >> Martin Fuerderer >> IBM Informix Development Munich, Germany >> Information Management >> >> IBM Deutschland Research & Development GmbH >> Chairman of the Supervisory Board: Martin Jetter >> Board of Management: Herbert Kircher >> Corporate Seat: Boeblingen, Germany >> Reg.-Gericht: Amtsgericht Stuttgart, HRB 243294 >> > > So, Informix get's an O/S error ... what should it do? Wait? Try again? Or error on a suspicious write return code for a critical dbspace :-/ But, wait .. this is meant to be "transparent"! i.e. the Storage Manager is doing some maintenance, and tells the O/S "Here is an error, so ... it is NOT transparent :-/