onstat -u shows "blocked: LONGTXN"
Posted in 2000
Topics: General Discussion
Yesterday some program seemed to go a bit silly and caused the engine to
hang with a long transaction. onstat -u showed in the top line a message
similar
to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called support
believing that the roll back was not working. Informix suggested killing
engine and
bringing online with LTAPEDEV=/dev/null and then set back to the orinal tape
device.
So, I cannot find much doc on what blocked means. The docs I had seemed to
indicate that
the engine was blocked because it was rolling back a long transaction rather
than something
had happened to the engine meaning it had to be killed of.
Is this correct or what ?
Tam McLaughlin wrote:
> Yesterday some program seemed to go a bit silly and caused the engine to
> hang with a long transaction. onstat -u showed in the top line a message
> similar
> to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called support
> believing that the roll back was not working. Informix suggested killing
> engine and
> bringing online with LTAPEDEV=/dev/null and then set back to the orinal tape
> device.
> So, I cannot find much doc on what blocked means. The docs I had seemed to
> indicate that
> the engine was blocked because it was rolling back a long transaction rather
> than something
> had happened to the engine meaning it had to be killed of.
> Is this correct or what ?
The message indicates that, yes, but if you do not have enough free logical log
space
to allow the rollback to be documented before they all fill then the engine will
become
permanently hung. This can happen if you do not archive the logs so they can be
freed and reused or if there is just not enough logical log space to begin with
and all
of the logs are locked from reuse because that long transaction has spanned then
all.
The former case can be fixed by the method that tech support recommended. The
latter is usually unrecoverable and you may have had to restore from archive
after
truncating the last logfile or two (tech support has to do that for you) to
allow room for
the fast recovery rollback.
Bottom line?
1) Keep runaway transactions from happening and
2) Increase the logical log space, set LBU_PRESERVE to save at least one whole
logical logfile for recovery after a hang up, and increase LTXHWM and LTXEHWM
to provide more headroom for long transaction rollbacks.
Art S. Kagel
PC User wrote in message <388DCC93.8FBA28BD@bloomberg.net>...
>Tam McLaughlin wrote:
>
>> Yesterday some program seemed to go a bit silly and caused the engine to
>> hang with a long transaction. onstat -u showed in the top line a message
>> similar
>> to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called
support
>> believing that the roll back was not working. Informix suggested killing
>> engine and
>> bringing online with LTAPEDEV=/dev/null and then set back to the orinal
tape
>> device.
>> So, I cannot find much doc on what blocked means. The docs I had seemed
to
>> indicate that
>> the engine was blocked because it was rolling back a long transaction
rather
>> than something
>> had happened to the engine meaning it had to be killed of.
>> Is this correct or what ?
>
>The message indicates that, yes, but if you do not have enough free logical
log
>space
>to allow the rollback to be documented before they all fill then the engine
will
>become
>permanently hung. This can happen if you do not archive the logs so they
can be
>
>freed and reused or if there is just not enough logical log space to begin
with
>and all
>of the logs are locked from reuse because that long transaction has spanned
then
>all.
>
>The former case can be fixed by the method that tech support recommended.
The
>latter is usually unrecoverable and you may have had to restore from
archive
>after
>truncating the last logfile or two (tech support has to do that for you) to
>allow room for
>the fast recovery rollback.
>
>Bottom line?
>1) Keep runaway transactions from happening and
>2) Increase the logical log space, set LBU_PRESERVE to save at least one
whole
>logical logfile for recovery after a hang up, and increase LTXHWM and
LTXEHWM>to provide more headroom for long transaction rollbacks.
>
>Art S. Kagel
>
>
Our logs get archived continuously but I will increase the number just in
case. I do tell
the developers to watch out for long transaction. I also thought that if i
had
LTXHWM and LTXEHWM set to 45 and 50 then it will mean that a long
transaction will
have plenty of log space to roll back i.e. what if it uses 45 % of log space
and needs another 45% to roll back?
>
Tam McLaughlin wrote:
> PC User wrote in message <388DCC93.8FBA28BD@bloomberg.net>...
> >Tam McLaughlin wrote:
> >
> >> Yesterday some program seemed to go a bit silly and caused the engine to
> >> hang with a long transaction. onstat -u showed in the top line a message
> >> similar
> >> to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called
> support
> >> believing that the roll back was not working. Informix suggested killing
> >> engine and
> >> bringing online with LTAPEDEV=/dev/null and then set back to the orinal
> tape
> >> device.
> >> So, I cannot find much doc on what blocked means. The docs I had seemed
> to
> >> indicate that
> >> the engine was blocked because it was rolling back a long transaction
> rather
> >> than something
> >> had happened to the engine meaning it had to be killed of.
> >> Is this correct or what ?
> >
> >The message indicates that, yes, but if you do not have enough free logical
> log
> >space
> >to allow the rollback to be documented before they all fill then the engine
> will
> >become
> >permanently hung. This can happen if you do not archive the logs so they
> can be
> >
> >freed and reused or if there is just not enough logical log space to begin
> with
> >and all
> >of the logs are locked from reuse because that long transaction has spanned
> then
> >all.
> >
> >The former case can be fixed by the method that tech support recommended.
> The
> >latter is usually unrecoverable and you may have had to restore from
> archive
> >after
> >truncating the last logfile or two (tech support has to do that for you) to
> >allow room for
> >the fast recovery rollback.
> >
> >Bottom line?
> >1) Keep runaway transactions from happening and
> >2) Increase the logical log space, set LBU_PRESERVE to save at least one
> whole
> >logical logfile for recovery after a hang up, and increase LTXHWM and
> LTXEHWM> >to provide more headroom for long transaction rollbacks.
> >
> >Art S. Kagel
> >
> >
>
> Our logs get archived continuously but I will increase the number just in
> case. I do tell
> the developers to watch out for long transaction. I also thought that if i
> had
> LTXHWM and LTXEHWM set to 45 and 50 then it will mean that a long
> transaction will
> have plenty of log space to roll back i.e. what if it uses 45 % of log space
> and needs another 45% to roll back?
> >
That would work out OK since the LTXEHWM will stop all other activity until the
transaction has completed its rollback leaving it almost 50% of the logs for
rollback
records.
Art S. Kagel
PC User wrote:
> Tam McLaughlin wrote:
> ...
> > to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called support
> > believing that the roll back was not working. Informix suggested killing
> > engine and
> > bringing online with LTAPEDEV=/dev/null and then set back to the orinal tape
> > device.
> > So, I cannot find much doc on what blocked means. The docs I had seemed to
> > indicate that
"Blocked : LONGTXN" just means that the logical logs are now reserved for the "long
transaction" that is being rolled back. No other transactions have access to the
logical logs and, therefore, remain blocked. One way to confirm that Informix is
still rolling back the "long transaction" and is not stuck after running out of
logical log space is to examine the usage of logs (onstat -l). If you find that a
log is still being written to (% used gradually going up), the rollback is
progressing; if not, start calling for help. Rollbacks usually complete faster that
the time taken for the trx to reach the "long" stage - but not always; in fact,
there are times when a rollback, "long" or otherwise, takes longer to accomplish.
>
> > the engine was blocked because it was rolling back a long transaction rather
> > than something
> > had happened to the engine meaning it had to be killed of.
> > Is this correct or what ?
> ...
>
> The former case can be fixed by the method that tech support recommended. The
> latter is usually unrecoverable and you may have had to restore from archive
> after
> truncating the last logfile or two (tech support has to do that for you) to
> allow room for
> the fast recovery rollback.
>
> Bottom line?
> 1) Keep runaway transactions from happening and
> 2) Increase the logical log space, set LBU_PRESERVE to save at least one whole
> logical logfile for recovery after a hang up, and increase LTXHWM and LTXEHWM
> to provide more headroom for long transaction rollbacks.
You mean DECREASE the parameters LTXHWM and LTXEHWM, don't you? Increasing them
means that Informix has less headroom to rollback because it marks a transaction as
"long" after it uses up a higher % of the logical logs - it increases the danger of
a "lock-up". (Increasing LTXHWM and LTXEHWM does reduce the possibility of a
transaction being marked as "long", but isn't recommended, I'm sure you will
agree).
Finally, as far as I can see, keeping LTXEHWM at 50 should almost guarantee that a
Lock-up (due to Long Trx Rollback running out of log space) will never occur.
Rudy
Rudy Fernandes wrote:
> PC User wrote:
>
> > Tam McLaughlin wrote:
> > ...
> > > to "Blocked: LONGTXN". The engine was stuck for 1/2 hr so I called support
> > > believing that the roll back was not working. Informix suggested killing
> > > engine and
> > > bringing online with LTAPEDEV=/dev/null and then set back to the orinal tape
> > > device.
> > > So, I cannot find much doc on what blocked means. The docs I had seemed to
> > > indicate that
>
> "Blocked : LONGTXN" just means that the logical logs are now reserved for the "long
> transaction" that is being rolled back. No other transactions have access to the
> logical logs and, therefore, remain blocked. One way to confirm that Informix is
> still rolling back the "long transaction" and is not stuck after running out of
> logical log space is to examine the usage of logs (onstat -l). If you find that a
> log is still being written to (% used gradually going up), the rollback is
> progressing; if not, start calling for help. Rollbacks usually complete faster that
> the time taken for the trx to reach the "long" stage - but not always; in fact,
> there are times when a rollback, "long" or otherwise, takes longer to accomplish.
>
> >
> > > the engine was blocked because it was rolling back a long transaction rather
> > > than something
> > > had happened to the engine meaning it had to be killed of.
> > > Is this correct or what ?
> > ...
> >
> > The former case can be fixed by the method that tech support recommended. The
> > latter is usually unrecoverable and you may have had to restore from archive
> > after
> > truncating the last logfile or two (tech support has to do that for you) to
> > allow room for
> > the fast recovery rollback.
> >
> > Bottom line?
> > 1) Keep runaway transactions from happening and
> > 2) Increase the logical log space, set LBU_PRESERVE to save at least one whole
> > logical logfile for recovery after a hang up, and increase LTXHWM and LTXEHWM
> > to provide more headroom for long transaction rollbacks.
>
> You mean DECREASE the parameters LTXHWM and LTXEHWM, don't you?
I meant increase the headroom. Does one scroll up or down? Should the down
arrow move the window down over the page or move the page down beneath the
window? This is just semantics and unfortunately not clear. My fault obviously for
not being clear.
Art S. Kagel
> Increasing them
> means that Informix has less headroom to rollback because it marks a transaction as
> "long" after it uses up a higher % of the logical logs - it increases the danger of
> a "lock-up". (Increasing LTXHWM and LTXEHWM does reduce the possibility of a
> transaction being marked as "long", but isn't recommended, I'm sure you will
> agree).
>
> Finally, as far as I can see, keeping LTXEHWM at 50 should almost guarantee that a
> Lock-up (due to Long Trx Rollback running out of log space) will never occur.
>
> Rudy