Fast recovery slow
Posted in 2016
Topics: Server Administration, Platform-Specific Issues
Hi,
My environment is Informix 11.70 FC8W1 on HP-UX 11.31.
I shutdown a server with onmode -ky. There was a "heavy" transaction in
progress. After running oninit -v, and after an hour, I see that the instance
is still in "fast recovery".
With onstat -u, I see that there is a thread doing the recovery:
c000000050b62290 --RPRRD 34 informix - 0 0 0 376349 305152
My question is:
Are there any way to make "fast recovery" faster ?
Currently, these are the parameters I have set:
ON_RECVRY_THREADS 30
OFF_RECVRY_THREADS 80
Thanks in advance,
Roger
Roger - This behavior is typical if the engine dies with heavy tx's taking place. It can typically be two things: 1) a result of transactions being rolled forward that just takes a long time due to activity. But then it's possible that transactions are now being rolled back since they were not committed. 2) you're stuck in fast recovery since "physical log overflow" occurred. Less likely in later versions of the engines since the you have the oppty to set PLOG_OVERFLOW_PATH to a file on disk in case the plog wraps. This helps avoid getting stuck in fast recovery due to overflow but can be VERY slow since it's simply reading through pages on disk. (Unlike the logical logs, if the phys log wraps all the way around, we keep writing pages, and yes, have just overwritten pages needed for recovery. Thus the PLOG_OVERLOW_PATH config param.) The default for this I believe is $INFORMXDIR/tmp, which may or may not have enough room for a bad overflow. Either way, check your msg log for: 1) logical logs being walked through during fast recovery. You'll see them increment over time. 2) transactions being rolled back. 3) "physical log overflow has occurred" - if the engine died after this occured, you could be stuck in fast recovery and it'll never end. Since we don't see getting stuck in fast recovery as much as back in the day, I'd be curious to hear others weigh in on this too with 11.7+. If you'll send your msg log (or post part of it here) I'd interrogate it for the different behaviors I described above. Thanks - Mark Scranton The Mark Scranton Group mark@markscranton.com 765-729-2601
Thanks Mark, I send the relevant part of the engine log. You can see that the logical recovery is slow. 19:48:13 IBM Informix Dynamic Server Initialized -- Shared Memory Initialized. 19:48:13 Started 1 B-tree scanners. 19:48:13 B-tree scanner threshold set at 5000. 19:48:13 B-tree scanner range scan size set to -1. 19:48:13 B-tree scanner ALICE mode set to 6. 19:48:13 B-tree scanner index compression level set to med. 19:48:13 Physical Recovery Started at Page (221:448295). 19:48:13 Physical Recovery Complete: 1 Pages Examined, 1 Pages Restored. 19:48:14 Logical Recovery Started. 19:48:14 80 recovery worker threads will be started. 19:48:14 Logical Recovery has reached the transaction cleanup phase. 21:01:16 Logical Recovery Complete. 0 Committed, 1 Rolled Back, 0 Open, 0 Bad Locks 21:01:16 Dataskip is now OFF for all dbspaces 21:01:20 SCHAPI: Started dbScheduler thread. 21:01:21 Defragmenter cleaner thread now running 21:01:21 Defragmenter cleaner thread cleaned:0 partitions 21:01:49 Booting Language <spl> from module <> 21:01:49 Loading Module <SPLNULL> 21:01:49 Auto Registration is synced 21:01:51 SCHAPI: Started 2 dbWorker threads. 21:02:51 Loading Module <$INFORMIXDIR/extend/ifxmngr/ifxmngr.bld> 21:02:51 The C Language Module </opt/informix/extend/ifxmngr/ifxmngr.bld> loaded 21:04:03 Checkpoint Completed: duration was 163 seconds. 21:04:03 Tue Feb 9 - loguniq 331134, logpos 0x11bd015c, timestamp: 0x4eae9d79 Interval: 460956 21:04:03 Maximum server connections 13 21:04:03 Checkpoint Statistics - Avg. Txn Block Time 0.000, # Txns blocked 0, Plog used 383086, Llog used 57757 21:04:03 On-Line Mode Regards, Roger
Roger - My apologies...I thought you were saying that recovery was slow but it didn't complete. So my dissertation on getting stuck fast recovery is just good education now. ;) I see that it did complete (due to the "online mode" message. I would think that the section where I talked about transactions being applied taking a long time is normal and expected. I don't see any that got rolled back so that didn't add to the time (unless the msg log was cut out for brevity). Many factors come into play here - disk I/O, horsepower of engine, load on the system, etc. Heavy transactions get killed due to the engine stopping will always take "some time" depending on items above. Let me know if you have anymore questions. Fun topic. Thanks - Mark Scranton The Mark Scranton Group mark@markscranton.com
A very simple way to re-create this is to 1. create a table 2. do a checkpoint 3. load X million rows of data and commit 4. kill the engine (before the next checkpoint) 5. start the engine. Fast recovery has to replay the x million r= ow load. For this exact reason we created RTO=5FSERVER=5FRESTART # RTO=5FSERVER=5FRESTART - Specifies, in seconds, the Recovery Time # &= nbsp; Objective for Informi= x restart after a server # &nbs= p; &= nbsp; failure. Acceptable values are 0 (off) and # &nbs= p; &= nbsp; any number from 60-1800, inclusive. John F. Miller III STSM, Lead Architect miller3@us.ibm.com503-747-1366 IBM Informix Dynamic Server (IDS) ----- Original message ----- From: "MARK SCRANTON" <mark@marks= cranton.com> Sent by: ids-bounces@iiug.org To: ids@iiug.org Cc:= Subject: Re: Fast recovery slow [36600] Date: Mon, Feb 22, 2016 11:3= 4 AM Roger - My apologies...I thought you were saying that recovery = was slow but it didn't complete. So my dissertation on getting stuck fas= t recovery is just good education now. ;) I see that it did compl= ete (due to the "online mode" message. I would think that the section wh= ere I talked about transactions being applied taking a long time is norm= al and expected. I don't see any that got rolled back so that didn't add= to the time (unless the msg log was cut out for brevity). Many factors = come into play here - disk I/O, horsepower of engine, load on the system= , etc. Heavy transactions get killed due to the engine stopping will alw= ays take "some time" depending on items above. Let me know if you ha= ve anymore questions. Fun topic. Thanks - Mark Scranton The Ma= rk Scranton Group mark@markscranton.com *********************= ********************************************************** F= orum Note: Use "Reply" to post a response in the discussion forum. <= br>