HDR - Alarmprogram when down
Posted in 2012
HDR replication failed with network errors on the primary, then the secondary crashed during logical recovery with error 126 (bad rowid). Responder suggested running oncheck -cDI on the primary to detect corruption before rebuilding HDR. User reported oncheck results showing index fragment warnings and leftover HPL conversion tables (_vio, _dia) with locking issues, proposing to drop these tables as a test.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: High Availability & Replication, Storage & Space Management, Error Codes & Troubleshooting, Triggers, Constraints & Referential Integrity, Logging & Checkpoints, Networking & sqlhosts Configuration, Versions, Editions & End-of-Life
Our HDR is up and running onstat -g dry all looks good. Then it drops and on
the primary server log we get this.
12:35:00 Maximum server connections 62
12:35:00 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
Plog used 200, Llog used 236
12:38:35 DR: Send error
12:38:35 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
Network connection is broken.
12:38:35 DR_ERR set to -2
12:38:37 DR: Turned off on primary server
12:38:37 DR: Cannot connect to secondary server
12:43:49 DR: Cannot connect to secondary server
Now these 2 servers are in a data centre in the same rack with a direct cable
connecting them so it is unlikely to be a physical break.
The secondary server is now down and when trying to start oninit -v log entry
is:
10:57:50 Started 2 B-tree scanners.
10:57:50 B-tree scanner threshold set at 500000.
10:57:50 B-tree scanner range scan size set to -1.
10:57:50 B-tree scanner ALICE mode set to 6.
10:57:50 B-tree scanner index compression level set to med.
10:57:50 Physical Recovery Started at Page (3:390209).
10:57:50 Physical Recovery Complete: 137 Pages Examined, 137 Pages Restored.
10:57:51 DR: Trying to connect to primary server = ol_arb_runix
10:57:51 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
secondary.
10:57:51 Dataskip is now OFF for all dbspaces
10:57:51 Restartable Restore has been ENABLED
10:57:51 Recovery Mode
10:58:01 DR: Secondary server connected
10:58:03 DR: Secondary server needs failure recovery
10:58:03 DR: Failure recovery from disk in progress ...
10:58:04 Logical Recovery Started.
10:58:04 10 recovery worker threads will be started.
10:58:04 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
secondary.
10:58:04 Start Logical Recovery - Start Log 262, End Log ?
10:58:04 Starting Log Position - 262 0x27e018
10:58:04 Started processing open transactions on secondary during startup
10:58:04 Finished processing open transactions on secondary during startup.
10:58:04 Rollforward of log record failed. iserrno = 126
10:58:04 Log Record: log = 262, pos = 0x3160a8, type = OLDRSAM:HUPAFT(43),
trans = 3547280096119226461
10:58:04 Assert Failed: Logical log replay error.
10:58:04 IBM Informix Dynamic Server Version 11.70.FC3GE
10:58:04 Who: Session(19, informix@arb-vdc-backix, 0, 0x9a831c20)
Thread(43, xchg_1.3, 9a7f7f30, 1)
File: rsprecvr.c Line: 7269
10:58:04 Results: The secondary server cannot continue.
10:58:04 Action: Reestablish the secondary server.
10:58:04 stack trace for pid 18108 written to /opt/IBM/informix/tmp/af.413bb9c
10:58:04 See Also: /opt/IBM/informix/tmp/af.413bb9c, shmem.413bb9c.0
10:58:22 Starting crash time check of:
10:58:22 1. memory block headers
10:58:22 2. stacks
10:58:22 Crash time checking found no problems
10:58:22 rsprecvr.c, line 7269, thread 43, proc id 18108, Logical log replay
error..
10:58:24 The Master Daemon Died
10:58:24 PANIC: Attempting to bring system down
It now looks like we will have to do a full backup on primary and restore on
secondary to get HDR up again. Does this scenario seem right?
Also is there a way to setup alarmprogram to email if HDR goes down for any
reason? I can not seem to find a trigger for this?
Many Thanks
Derrick Muller
I see this in the secondary's log:
10:58:04 Rollforward of log record failed. iserrno = 126
ISAM error 126 is bad rowid. I think that there is some corruption on theprimary and it would be good to run oncheck -cDI there before you do
anything else to check it out.
What was in the secondary's log at the time replication went down initially?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Tue, Apr 3, 2012 at 5:07 AM, DERRICK MULLER <derrick@xact.co.za> wrote:
> Our HDR is up and running onstat -g dry all looks good. Then it drops and
> on
> the primary server log we get this.
>
> 12:35:00 Maximum server connections 62
> 12:35:00 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked
> 0,
> Plog used 200, Llog used 236
>
> 12:38:35 DR: Send error
> 12:38:35 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
> Network connection is broken.
>
> 12:38:35 DR_ERR set to -2
> 12:38:37 DR: Turned off on primary server
> 12:38:37 DR: Cannot connect to secondary server
> 12:43:49 DR: Cannot connect to secondary server
>
> Now these 2 servers are in a data centre in the same rack with a direct
> cable
> connecting them so it is unlikely to be a physical break.
>
> The secondary server is now down and when trying to start oninit -v log
> entry
> is:
>
> 10:57:50 Started 2 B-tree scanners.
> 10:57:50 B-tree scanner threshold set at 500000.
> 10:57:50 B-tree scanner range scan size set to -1.
> 10:57:50 B-tree scanner ALICE mode set to 6.
> 10:57:50 B-tree scanner index compression level set to med.
> 10:57:50 Physical Recovery Started at Page (3:390209).
> 10:57:50 Physical Recovery Complete: 137 Pages Examined, 137 Pages
> Restored.
> 10:57:51 DR: Trying to connect to primary server = ol_arb_runix
> 10:57:51 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
> secondary.
> 10:57:51 Dataskip is now OFF for all dbspaces
> 10:57:51 Restartable Restore has been ENABLED
> 10:57:51 Recovery Mode
> 10:58:01 DR: Secondary server connected
> 10:58:03 DR: Secondary server needs failure recovery
>
> 10:58:03 DR: Failure recovery from disk in progress ...
> 10:58:04 Logical Recovery Started.
> 10:58:04 10 recovery worker threads will be started.
> 10:58:04 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
> secondary.
> 10:58:04 Start Logical Recovery - Start Log 262, End Log ?
> 10:58:04 Starting Log Position - 262 0x27e018
> 10:58:04 Started processing open transactions on secondary during startup
> 10:58:04 Finished processing open transactions on secondary during startup.
> 10:58:04 Rollforward of log record failed. iserrno = 126
> 10:58:04 Log Record: log = 262, pos = 0x3160a8, type = OLDRSAM:HUPAFT(43),
> trans = 3547280096119226461
> 10:58:04 Assert Failed: Logical log replay error.
> 10:58:04 IBM Informix Dynamic Server Version 11.70.FC3GE
> 10:58:04 Who: Session(19, informix@arb-vdc-backix, 0, 0x9a831c20)
>
> Thread(43, xchg_1.3, 9a7f7f30, 1)
>
> File: rsprecvr.c Line: 7269
> 10:58:04 Results: The secondary server cannot continue.
> 10:58:04 Action: Reestablish the secondary server.
> 10:58:04 stack trace for pid 18108 written to
> /opt/IBM/informix/tmp/af.413bb9c
> 10:58:04 See Also: /opt/IBM/informix/tmp/af.413bb9c, shmem.413bb9c.0
> 10:58:22 Starting crash time check of:
> 10:58:22 1. memory block headers
> 10:58:22 2. stacks
> 10:58:22 Crash time checking found no problems
> 10:58:22 rsprecvr.c, line 7269, thread 43, proc id 18108, Logical log
> replay
> error..
> 10:58:24 The Master Daemon Died
> 10:58:24 PANIC: Attempting to bring system down
>
> It now looks like we will have to do a full backup on primary and restore
> on
> secondary to get HDR up again. Does this scenario seem right?
>
> Also is there a way to setup alarmprogram to email if HDR goes down for any
> reason? I can not seem to find a trigger for this?
>
> Many Thanks
>
> Derrick Muller
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--14dae9340d5992d46404bcc4105b
Thanks Art
This informix stuff looks too smart. If I understand this all correctly the
seconadary will NOT accept corrupt or damaged looking data thereby protecting
the integrity of the last good data accepted.
Running oncheck give following:
ON ALL DB INDEXES WE GET:
Validating indexes for arb_db:wwwrun.sa22_qt_so_hd...
Index 224_345
Index fragment partition arb_dat_dbs in DBspace arb_dat_dbs
Index sa22_stk_enq
Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
Index sa22_dl_code_name
Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
Index sa22_loc_type_doc_no
Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
Is there any significance with "Index fragment............." message or is
this normal?
ON INDEXES LEFT OVER FROM HPL (_vio & _dia from unique constraints):
WARNING: index check requires a s-lock on tables whose lock level is page.
Validating indexes for arb_db:wwwrun.bo30_tran_dia...
WARNING: index check requires a s-lock on tables whose lock level is page.
Validating indexes for arb_db:wwwrun.bo30_tran_vio...
Should we be manually be dropping these _vio & _dia tables after HPL run?
HPL is a once of to convert from current DBISAM db to Informix DB.
ON TABLE SPACES (same _vio & _dia):
TBLspace data check for arb_db:wwwrun.bo30_tran_dia
WARNING: index check requires a s-lock on tables whose lock level is page.
TBLspace data check for arb_db:wwwrun.bo30_tran_vio
WARNING: index check requires a s-lock on tables whose lock level is page.
HERE IS THE SECONDARY LOG AT POINT OF FAILURE:
(same error as on primary at same time: 12:37:46 Rollforward of log record
failed. iserrno = 126)
12:34:27 Maximum server connections 0
12:34:27 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0,
Plog used 220, Llog used 0
12:37:46 Rollforward of log record failed. iserrno = 126
12:37:46 Log Record: log = 262, pos = 0x3160a8, type = OLDRSAM:HUPAFT(43),
trans = 3547280096119226461
12:37:46 Assert Failed: Logical log replay error.
12:37:46 IBM Informix Dynamic Server Version 11.70.FC3GE
12:37:46 Who: Session(24, informix@arb-vdc-backix, 0, 0x5d64f270)
Thread(51, xchg_1.3, 5d615fc0, 1)
File: rsprecvr.c Line: 7269
12:37:46 Results: The secondary server cannot continue.
If you agree I think we will drop the _vio & _dia to test.
Thanks again
OK, so the log message on the secondary is indicating that a bad rowid in a
log record could not be rolled forward. That's bad.
Yes, try dropping all of the violations tables, but first you have to STOP
VIOLATIONS FOR <tablename>; then you can DROP TABLE <tablename>_vio; DROP
TABLE <tablename>_dia;
However, the messages from oncheck are just warnings that because those
tables have page level locking enabled oncheck has to lock the pages of the
indexes as it is checking them which it tries to avoid doing unless you
tell it to.
So, the oncheck -cDI came up clean? I only see the index checks (-cI) did
you also do the data checks (-cD)?
If, after dropping the violations tables and rebuilding the secondary again
from an archive, you still can't keep replication up, I see two choices:
1 - open a support case with IBM
2 - unload all of your databases with dbexport or myexport, reinitialize
the instance, and reload the databases then reestablish replication again
with a clean engine.
I'm also going to suggest that you upgrade to 11.70.FC4GE. I don't know of
any specific problems in .FC3 that are related, but it certainly can't hurt
and .FC4 fixes some performance problems with the new readahead algorithm.
You say you used HPL to load the CISAM data into IDS. Just curious, did
you use DELUXE mode or EXPRESS mode?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Tue, Apr 3, 2012 at 8:23 AM, DERRICK MULLER <derrick@xact.co.za> wrote:
> Thanks Art
>
> This informix stuff looks too smart. If I understand this all correctly the
> seconadary will NOT accept corrupt or damaged looking data thereby
> protecting
> the integrity of the last good data accepted.
>
> Running oncheck give following:
>
> ON ALL DB INDEXES WE GET:
> Validating indexes for arb_db:wwwrun.sa22_qt_so_hd...
> Index 224_345
> Index fragment partition arb_dat_dbs in DBspace arb_dat_dbs
> Index sa22_stk_enq
> Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
> Index sa22_dl_code_name
> Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
> Index sa22_loc_type_doc_no
> Index fragment partition arb_idx_dbs in DBspace arb_idx_dbs
>
> Is there any significance with "Index fragment............." message or is
> this normal?
>
> ON INDEXES LEFT OVER FROM HPL (_vio & _dia from unique constraints):
>
> WARNING: index check requires a s-lock on tables whose lock level is page.
> Validating indexes for arb_db:wwwrun.bo30_tran_dia...
> WARNING: index check requires a s-lock on tables whose lock level is page.
> Validating indexes for arb_db:wwwrun.bo30_tran_vio...
>
> Should we be manually be dropping these _vio & _dia tables after HPL run?
> HPL is a once of to convert from current DBISAM db to Informix DB.
>
> ON TABLE SPACES (same _vio & _dia):
>
> TBLspace data check for arb_db:wwwrun.bo30_tran_dia
> WARNING: index check requires a s-lock on tables whose lock level is page.
> TBLspace data check for arb_db:wwwrun.bo30_tran_vio
> WARNING: index check requires a s-lock on tables whose lock level is page.
>
> HERE IS THE SECONDARY LOG AT POINT OF FAILURE:
> (same error as on primary at same time: 12:37:46 Rollforward of log record
> failed. iserrno = 126)
>
> 12:34:27 Maximum server connections 0
> 12:34:27 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked
> 0,
> Plog used 220, Llog used 0
>
> 12:37:46 Rollforward of log record failed. iserrno = 126
> 12:37:46 Log Record: log = 262, pos = 0x3160a8, type = OLDRSAM:HUPAFT(43),
> trans = 3547280096119226461
> 12:37:46 Assert Failed: Logical log replay error.
> 12:37:46 IBM Informix Dynamic Server Version 11.70.FC3GE
> 12:37:46 Who: Session(24, informix@arb-vdc-backix, 0, 0x5d64f270)
>
> Thread(51, xchg_1.3, 5d615fc0, 1)
>
> File: rsprecvr.c Line: 7269
> 12:37:46 Results: The secondary server cannot continue.
>
> If you agree I think we will drop the _vio & _dia to test.
>
> Thanks again
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--14dae93408159a870104bcc5c870
Thanks Art Dropped _vio & _dia tables. Restarted both primary and secondary db's Got a script which is flood testing primary by inserting 1,000,000 records to test. Logs auto backing up on primary Tables updating on secondary perfectly Logs updating on secondary perfectly HDR not failing any longer. Interestingly after 1 hour of inserts only 490,000 recs inserted (5 varchar fields) On primary %cached writes = 62.23 %cached reads = 99.36 On secondary %cached writes = 86.12 %cached reads = 99.86 I assume this requires some configuration tweaking to get insert speeds unto scratch. Of interest we used EXPRESS with HPL Question on backup of logs on secondary. Have set alarmprogram on secondary same as on primary but secondary does not auto backup logs. Is this correct? What would happen in this senario: 1.) Daily data backups are synced to a NAS server via alarmprogram. 2.) Primary logs are synced to a NAS server via alarmprogram. 3.) On Secondary logs from primary are received but roll over during day so secondary log table will lose some earlier logs files (overwritten). I assume if you had a large number of log files this would not happen (i.e. sufficient log files for 1 days usage). Is this practical? 4.) Primary crashes and we switch to secondary all good. Logs now synced from secondary to NAS server. Log no's are incremented and intact. 5.) Before Primary is repaired (dooms day senario) secondary also fails. 6.) When we get Primary back up we would restore prev nights db backup from NAS server. 7.) Now for the logs. We restore logs from NAS server (which were originally created on Primary) 8.) We restore logs from NAS server synced after We would have backups of original primary logs on another server after 4.) above. 9.) Primary back up with no data loss. Is this sort of correct. Thanks Derrick Muller
Answered privately. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Wed, Apr 4, 2012 at 4:16 AM, DERRICK MULLER <derrick@xact.co.za> wrote: > Thanks Art > > Dropped _vio & _dia tables. > Restarted both primary and secondary db's > Got a script which is flood testing primary by inserting 1,000,000 records > to > test. > Logs auto backing up on primary > Tables updating on secondary perfectly > Logs updating on secondary perfectly > HDR not failing any longer. > > Interestingly after 1 hour of inserts only 490,000 recs inserted (5 varchar > fields) > On primary %cached writes = 62.23 %cached reads = 99.36 > On secondary %cached writes = 86.12 %cached reads = 99.86 > I assume this requires some configuration tweaking to get insert speeds > unto > scratch. > > Of interest we used EXPRESS with HPL > > Question on backup of logs on secondary. > Have set alarmprogram on secondary same as on primary but secondary does > not > auto backup logs. Is this correct? > > What would happen in this senario: > 1.) Daily data backups are synced to a NAS server via alarmprogram. > 2.) Primary logs are synced to a NAS server via alarmprogram. > 3.) On Secondary logs from primary are received but roll over during day so > secondary log table will lose some earlier logs files (overwritten). I > assume > if you had a large number of log files this would not happen (i.e. > sufficient > log files for 1 days usage). Is this practical? > 4.) Primary crashes and we switch to secondary all good. Logs now synced > from > secondary to NAS server. Log no's are incremented and intact. > 5.) Before Primary is repaired (dooms day senario) secondary also fails. > 6.) When we get Primary back up we would restore prev nights db backup from > NAS server. > 7.) Now for the logs. We restore logs from NAS server (which were > originally > created on Primary) > 8.) We restore logs from NAS server synced after We would have backups of > original primary logs on another server after 4.) above. > 9.) Primary back up with no data loss. > > Is this sort of correct. > > Thanks > Derrick Muller > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --e89a8f3ba9593502f104bcd79d66
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g