11.5 crashes after upgrade, data lost
Posted in 2009
After upgrading from IDS 9.4 (32-bit) to 11.50.FC3 (64-bit) on a reinstalled Red Hat box, the server crashed with an "Assert Failed: No Exception Handler ... MT_EX_OS, mem" message, and all BLOB/CLOB values written since the last checkpoint came back zero-length rather than null. The poster's outsourced DBAs blamed stack size or the newly enabled RTO_SERVER_RESTART. Respondents asked about storage layout; the smart blob spaces sat on cooked ext3 files with DIRECT_IO on, and Art Kagel suggested ext3 (especially writeback mode) is unsafe for chunks and could explain the corruption. No confirmed fix or root cause is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Installation, Setup & Upgrades, Error Codes & Troubleshooting, Server Administration, Logging & Checkpoints, Versions, Editions & End-of-Life
Hi, We've just upgraded our production db from 9.4 32-bit to 11.5 64-bit. (Same machine, but Red-Hat upgraded as well.) After a morning of stability, ifx crashed at about 3:15 in the afternoon. The dbas brought the db back up, we restarted our app, and immediately noticed strange messages in the app's logs. After some digging, it turns out that we had lost blob and clob data for rows written between 12:14 (the time of the last checkpoint) and the time of crash. To be more specific, the lob columns in the affected rows were not null, rather, the lobs were 0 length. Has anyone had a similar experience? After years of stability under 9.4, I'm now apprehensive. Particularly worrying is the loss of data integrity. I was of the understanding that inserts and updates of a row were atomic, regardless of whether lob columns are involved or not. Our db administration is outsourced, and I haven't heard an adequate explanation for either the crash or the data loss. One theory put forward is that stack size was too low -- whatever that means. An IBM representative recommended increasing it to 128, even though the online doco strongly recommends 64. Another theory involves the enabling of the RTO_SERVER_RESTART parameter. I'm told this is a new parameters for 11.5, and it was enabled after our testing of the 11.5 db had been completed. RTO_SERVER_RESTART has now been disabled and checkpoints are occurring at a more typical interval of about 5mins. Here's the log entry for the crash: 15:38:10 Assert Failed: No Exception Handler 15:38:10 IBM Informix Dynamic Server Version 11.50.FC3 15:38:10 Who: Session(17333, informix@aadcxs0001, 2329, 0x953a2ed8) Thread(17489, sqlexec, 8aabe170, 1) File: mtex.c Line: 481 15:38:10 Results: Exception Caught. Type: MT_EX_OS, Context: mem Any shedding of light would be appreciated. Thanks
Some questions: Are your chunks in filesystem space or device space? If a filesystem what type of filesystem? Is DIRECT_IO enabled? If DIRECT_IO is disabled how many AIO VPs versus how many chunks? If device space RAW (character) or COOKED (block) devices? Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, Apr 1, 2009 at 1:38 AM, ANTHONY PERSIC <tony.persic@amcor.com.au>wrote: > Hi, > > We've just upgraded our production db from 9.4 32-bit to 11.5 64-bit. (Same > machine, but Red-Hat upgraded as well.) After a morning of stability, ifx > crashed at about 3:15 in the afternoon. The dbas brought the db back up, we > restarted our app, and immediately noticed strange messages in the app's > logs. > After some digging, it turns out that we had lost blob and clob data for > rows > written between 12:14 (the time of the last checkpoint) and the time of > crash. > To be more specific, the lob columns in the affected rows were not null, > rather, the lobs were 0 length. > > Has anyone had a similar experience? After years of stability under 9.4, > I'm > now apprehensive. Particularly worrying is the loss of data integrity. I > was > of the understanding that inserts and updates of a row were atomic, > regardless > of whether lob columns are involved or not. > > Our db administration is outsourced, and I haven't heard an adequate > explanation for either the crash or the data loss. One theory put forward > is > that stack size was too low -- whatever that means. An IBM representative > recommended increasing it to 128, even though the online doco strongly > recommends 64. > > Another theory involves the enabling of the RTO_SERVER_RESTART parameter. > I'm > told this is a new parameters for 11.5, and it was enabled after our > testing > of the 11.5 db had been completed. RTO_SERVER_RESTART has now been disabled > and checkpoints are occurring at a more typical interval of about 5mins. > > Here's the log entry for the crash: > 15:38:10 Assert Failed: No Exception Handler > 15:38:10 IBM Informix Dynamic Server Version 11.50.FC3 > 15:38:10 Who: Session(17333, informix@aadcxs0001, 2329, 0x953a2ed8) > > Thread(17489, sqlexec, 8aabe170, 1) > > File: mtex.c Line: 481 > 15:38:10 Results: Exception Caught. Type: MT_EX_OS, Context: mem > > Any shedding of light would be appreciated. > > Thanks > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --0016361640ad2656a604667bf01b
Some things you didn't mention that might be important. Where are you storing the blobs/clobs (in tblspace, in blobspace, in smartblob space)? Does the database with this table that contains the blob have logging enabled? If you are using a smartblob space was it created with logging? Also, it's not clear if you mean every row inserted after the time of your last checkpoint has this blob problem, or just some of the rows inserted? Jacques Renaut IBM Informix support APD
Hi Jaques, I'll try to answer your questions as best I can: 1. Where are you storing the blobs/clobs (in tblspace, in blobspace, in smartblob space)? Smart blob space 2. Does the database with this table that contains the blob have logging enabled? Yes, but loss was over several tables, not isolated to one table. 3. If you are using a smartblob space was it created with logging. Don't know for sure, but I expect so. 4. Also, it's not clear if you mean every row inserted after the time of your last checkpoint has this blob problem, or just some of the rows inserted? All rows inserted or updated as far as we can tell. Thanks.
Hi Art, Here are the answers you requested: Are your chunks in filesystem space or device space? FS If a filesystem what type of filesystem? ext3 Is DIRECT_IO enabled? Yes If DIRECT_IO is disabled how many AIO VPs versus how many chunks? If device space RAW (character) or COOKED (block) devices? COOKED Thanks.
DId you see my post a few days ago about how EXT3 is NOT SAFE for database chunks? That may be the problem. Is the filesystem mounted with writeback mode enabled? That would put it at risk for the kind of corruption that I was reporting. Art Art S. Kagel Oninit (www.oninit.com) IIUG Board of Directors (art@iiug.org) Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Oninit, the IIUG, nor any other organization with which I am associated either explicitly or implicitly. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Thu, Apr 2, 2009 at 11:58 PM, ANTHONY PERSIC <tony.persic@amcor.com.au>wrote: > Hi Art, > > Here are the answers you requested: > > Are your chunks in filesystem space or device space? > FS > > If a filesystem what type of filesystem? > ext3 > > Is DIRECT_IO enabled? > Yes > > If DIRECT_IO is disabled how many AIO VPs versus how many chunks? > > If device space RAW (character) or COOKED (block) devices? > COOKED > > Thanks. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --00163691fe62a4a66c04669f292c