Re: HDR - Alarmprogram when down [26640]
Posted in 2012
I see this in the secondary's log:
10:58:04 Rollforward of log record failed. iserrno = 126
ISAM error 126 is bad rowid. I think that there is some corruption on theprimary and it would be good to run oncheck -cDI there before you do
anything else to check it out.
What was in the secondary's log at the time replication went down initially?
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Tue, Apr 3, 2012 at 5:07 AM, DERRICK MULLER <derrick@xact.co.za> wrote:
> Our HDR is up and running onstat -g dry all looks good. Then it drops and
> on
> the primary server log we get this.
>
> 12:35:00 Maximum server connections 62
> 12:35:00 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked
> 0,
> Plog used 200, Llog used 236
>
> 12:38:35 DR: Send error
> 12:38:35 ASF Echo-Thread Server: asfcode = -25582: oserr = 0: errstr = :
> Network connection is broken.
>
> 12:38:35 DR_ERR set to -2
> 12:38:37 DR: Turned off on primary server
> 12:38:37 DR: Cannot connect to secondary server
> 12:43:49 DR: Cannot connect to secondary server
>
> Now these 2 servers are in a data centre in the same rack with a direct
> cable
> connecting them so it is unlikely to be a physical break.
>
> The secondary server is now down and when trying to start oninit -v log
> entry
> is:
>
> 10:57:50 Started 2 B-tree scanners.
> 10:57:50 B-tree scanner threshold set at 500000.
> 10:57:50 B-tree scanner range scan size set to -1.
> 10:57:50 B-tree scanner ALICE mode set to 6.
> 10:57:50 B-tree scanner index compression level set to med.
> 10:57:50 Physical Recovery Started at Page (3:390209).
> 10:57:50 Physical Recovery Complete: 137 Pages Examined, 137 Pages
> Restored.
> 10:57:51 DR: Trying to connect to primary server = ol_arb_runix
> 10:57:51 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
> secondary.
> 10:57:51 Dataskip is now OFF for all dbspaces
> 10:57:51 Restartable Restore has been ENABLED
> 10:57:51 Recovery Mode
> 10:58:01 DR: Secondary server connected
> 10:58:03 DR: Secondary server needs failure recovery
>
> 10:58:03 DR: Failure recovery from disk in progress ...
> 10:58:04 Logical Recovery Started.
> 10:58:04 10 recovery worker threads will be started.
> 10:58:04 Warning: Logging dbspace 'temp_dbs' ignored in DBSPACETEMP on DR
> secondary.
> 10:58:04 Start Logical Recovery - Start Log 262, End Log ?
> 10:58:04 Starting Log Position - 262 0x27e018
> 10:58:04 Started processing open transactions on secondary during startup
> 10:58:04 Finished processing open transactions on secondary during startup.
> 10:58:04 Rollforward of log record failed. iserrno = 126
> 10:58:04 Log Record: log = 262, pos = 0x3160a8, type = OLDRSAM:HUPAFT(43),
> trans = 3547280096119226461
> 10:58:04 Assert Failed: Logical log replay error.
> 10:58:04 IBM Informix Dynamic Server Version 11.70.FC3GE
> 10:58:04 Who: Session(19, informix@arb-vdc-backix, 0, 0x9a831c20)
>
> Thread(43, xchg_1.3, 9a7f7f30, 1)
>
> File: rsprecvr.c Line: 7269
> 10:58:04 Results: The secondary server cannot continue.
> 10:58:04 Action: Reestablish the secondary server.
> 10:58:04 stack trace for pid 18108 written to
> /opt/IBM/informix/tmp/af.413bb9c
> 10:58:04 See Also: /opt/IBM/informix/tmp/af.413bb9c, shmem.413bb9c.0
> 10:58:22 Starting crash time check of:
> 10:58:22 1. memory block headers
> 10:58:22 2. stacks
> 10:58:22 Crash time checking found no problems
> 10:58:22 rsprecvr.c, line 7269, thread 43, proc id 18108, Logical log
> replay
> error..
> 10:58:24 The Master Daemon Died
> 10:58:24 PANIC: Attempting to bring system down
>
> It now looks like we will have to do a full backup on primary and restore
> on
> secondary to get HDR up again. Does this scenario seem right?
>
> Also is there a way to setup alarmprogram to email if HDR goes down for any
> reason? I can not seem to find a trigger for this?
>
> Many Thanks
>
> Derrick Muller
>
>
>
> *******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>