RE: long checkpoints
Posted in 2003
Topics: High Availability & Replication, Performance & Tuning, Storage & Space Management, Logging & Checkpoints, Platform-Specific Issues
Just a stab in the dark, but are you running HDR in synchronous mode
(DRINTERVAL=-1)? If so, that *could* slow things down. If so, I suggest
asynch mode.
Another thing -- since you're running an HCx version, I assume that this is
on HP-UX, but the HC version runs in 32-bit mode. So you're still subject
to the wonderful limitations of HP-UX memory management. Do you have
multiple virtual memory segments (onstat -g seg)? That definitely hurts
performance on HP-UX.
HTH,
Paul Mosser
-----Original Message-----
From: schmotom@yahoo.com [mailto:schmotom@yahoo.com]
Sent: Thursday, June 12, 2003 1:42 PM
To: informix-list@iiug.org
Subject: Ree: long checkpoints
muchos gracias!
we have 250,000 buffers - starting to think that maybe that is too
many....whatever is exactly going on - we are getting this blocking
information on (it seems) every checkpoint (running onstat -u and
onstat -R while checkpoints run): seems i will need to somehow drilldown into what exactly is causing this BLOCKING...?
: Informix Dynamic Server 2000 Version 9.21.HC5 -- On-Line (Prim)
(CKPT REQ) -- Up 1 days 18:08:56 -- 1004948 Kbytes
Blocked:CKPT
# Shared Memory Parameters
LOCKS 500000 # Maximum number of locks (8,000,000informix max; will allocate max of 15 100,000 additional blocks)
# LOCKS 300000 # Maximum number of locks mb 12/4/01
BUFFERS 250000 # Maximum number of shared buffers
NUMAIOVPS 2 # Number of IO vps
PHYSBUFF 32 # Physical log buffer size (Kbytes)
LOGBUFF 32 # Logical log buffer size (Kbytes)LOGSMAX 40 # Maximum number of logical log files
CLEANERS 50 # Number of buffer cleaner processes
SHMBASE 0x0 # Shared memory base address
SHMVIRTSIZE 196608 # initial virtual shared memorysegment size
SHMADD 32768 # Size of new shared memory segments
(Kbytes)
SHMTOTAL 0 # Total shared memory (Kbytes).
0=>unlimited
CKPTINTVL 300 # Check point interval (in sec)
LRUS 50 # Number of LRU queues
LRU_MAX_DIRTY 4 # LRU percent dirty begin cleaning
limit -- 30 changed 6/6/03
LRU_MIN_DIRTY 2 # LRU percent dirty end cleaning limit
-- 25 changed 6/6/03
LTXHWM 50 # Long transaction high water markpercentage
LTXEHWM 60 # Long transaction high water mark
(exclusive)
TXTIMEOUT 0x12c # Transaction timeout (in sec)
STACKSIZE 128 # Stack size (Kbytes) ---- was
32 changed 5/10/03
"David Williams" <djw@smooth1.fsnet.co.uk> wrote in message
news:<bc8dph$dt3$1@news6.svr.pol.co.uk>...
> "tomL" <schmotom@yahoo.com> wrote in message
> news:e70357eb.0306111425.75612c7e@posting.google.com...
> > thanks, Rob
> > here is our onstat -F:
> >
> > Fg Writes LRU Writes Chunk Writes
> > 0 58117 306213
> >
> > can anyone tell me more about the relevance of our
> > Fg Writes and LRU Writes and Chunk Writes?
> >
>
> Fg Write= No free buffers found so a 'foreground' as oppose to
> 'background'
> write of a dirty buffer was done. i.e. the sessions
> blocked whilst the
> dirty buffer was written to disk. You have 0 which is
> perfect!
>
> LRU Write = write done in background as far as sessions are concerned.
> Sessions
> do not block whilst the write occurs. This is what
> the cleaner threads
> do, they start queuing I/O when LRU_MAX_DIRTY
> percentage of
> buffers are dirty and stop queuing I/O when
> LRU_MIN_DIRTY> percentage of buffers are dirty. This is best in
> terms of response
> times for users since they do not block at
> checkpoint time. It is
> less efficient from the machine perspective since
> writes are not
> ordered (i.e. random writes rather then writes
> ordered by position
> in the chunk).
>
> Checkpoint writes = writes done at checkpoint time. At checkpoint time
> each cleaner
> a) is assigned to the next chunk in
chunk
> id order
> b)sorts all dirty buffers by offset
within
> the chunk (to reduce
> disk seeksm hence more efficient).
> c) queues all writes for the dirty
buffers
> d) waits for KAIO or AIO threads to
> complete the writes
> e) is assigned the next chunk to work
on.
>
> However whilst a checkpoint is occuring all sessions which require
> writes are frozen.
> You want almost all LRU writes to reduce checkpoint time.
>
> Get the last time from onstat -b to your buffer size in bytes
> (2048/4096).
>
> LRU_MIN_DIRTY * BUFFERS * <buffer size> = how many meg ?
> LRU_MAX_DIRTY * BUFFERS * <buffer size> = how many meg ?
>
>
>
>
>
> How many buffers do you have?
>
> > thanks again.....
> >
> >
> > Rob Wilson <rob_wilson@ameritech.net> wrote in message
> news:<DuGFa.2662$87.1640483@newssrv26.news.prodigy.com>...
> > > tomL wrote:
> > > > Thanks for all the help. But we have lowered our
> > > > LRU_MAX_DIRTY/MIN_DIRTY from 30/25 to 4/2 and still receive
> > > > checkpoints in the 25-50 second range - way too long. We notice that
> > > > checkpoints still seem to occur at the 5 minute CKPTINTVL interval
> > > > (rather than the % MIN/MAX), which is surprising to me. We are
> > > > thinking now about concentrating on whether we have too many buffers
> > > > or if it would help to lower our CKPTINTVL setting to something like
2
> > > > minutes. Any ideas?
> > > >
> > > >
> > >
> > > It is better to tune the time between checkpoints by tinkering with
the
> > > physical log size.
> > >
> > > The LRU_MIN_DIRTY/LRU_MAX_DIRTY parameters cause LRU writes which
occur
> > > in the background. If you look at onstat -F then you wlil see FG
writes
> > > (there were no clean buffers when a clean buffer was needed), LRU
writes
> > > (caused my min/max dirty), and chunk writes (occurs at checkpoint
time).
> > > LRU writes should be higher with the lower min/max - unless
something
> > > else is going on.
> > >
> > > A very good discussion about checkpoint tuning can be found at:
> > > http://www.smooth1.demon.co.uk/ifaq06.htm#6.23
Paul,
this is the output from onstat -g seg:
Segment Summary:
id key addr size ovhd class blkused
blkfree
7680 1381386241 80000000 592850944 245712 R* 144714 25
(shared) 1381386241 a3563000 201334784 6760 V 45433 3721
46597 1381386242 c2101000 192512 624 M 47 0
15366 1381386243 c2c0a000 33554432 1640 V 4479 3713
519 1381386244 c4d30000 33554432 1640 V 3297 4895
2056 1381386245 c6d30000 33554432 1640 V 982 7210
9 1381386246 c8d30000 33554432 1640 V 19 8173
10 1381386247 cad30000 33554432 1640 V 32 8160
11 1381386248 ccd30000 33554432 1640 V 69 8123
12 1381386249 ced30000 33554432 1640 V 2332 5860
79885 1381386250 d0dde000 33554432 1640 V 21 8171
15 1381386251 d2dde000 33554432 1640 V 352 7840
16 1381386252 d4dde000 33554432 1640 V 139 8053
526 1381386253 d6dde000 33554432 1640 V 34 8158
17 1381386254 d8dde000 33554432 1640 V 243 7949
Total: - - 1197031424 - - 202193 90051
(* segment locked in memory)
but i am unsure what i am looking for to determine if we have multiple
virtual memory segments ?
Thanks!
Tom
mosserp@wellsfargo.com wrote in message news:<bcbq0i$esc$1@terabinaries.xmission.com>...
> Just a stab in the dark, but are you running HDR in synchronous mode
> (DRINTERVAL=-1)? If so, that *could* slow things down. If so, I suggest
> asynch mode.
>
> Another thing -- since you're running an HCx version, I assume that this is
> on HP-UX, but the HC version runs in 32-bit mode. So you're still subject
> to the wonderful limitations of HP-UX memory management. Do you have
> multiple virtual memory segments (onstat -g seg)? That definitely hurts
> performance on HP-UX.
>
> HTH,
> Paul Mosser
On 13 Jun 2003 14:18:57 -0700, schmotom@yahoo.com (tomL) wrote:
Aaaugh, that might explain a few things . . . .
More than 4 SHM segments on HPUX is DEATH. I think that HPUX 11i
might give a few more, but the first thing I'd do is reset SHMVIRTSIZE
to SHMVIRTSIZE = current value in ONCONFIG + ( ( segments added + 2)
* SHMADD
The + 2 is just a fudge factor, YMMV.
>Paul,
>this is the output from onstat -g seg:
>
>Segment Summary:
>id key addr size ovhd class blkused
>blkfree
>7680 1381386241 80000000 592850944 245712 R* 144714 25
>(shared) 1381386241 a3563000 201334784 6760 V 45433 3721
>46597 1381386242 c2101000 192512 624 M 47 0
>15366 1381386243 c2c0a000 33554432 1640 V 4479 3713
>519 1381386244 c4d30000 33554432 1640 V 3297 4895
>2056 1381386245 c6d30000 33554432 1640 V 982 7210
>9 1381386246 c8d30000 33554432 1640 V 19 8173
>10 1381386247 cad30000 33554432 1640 V 32 8160
>11 1381386248 ccd30000 33554432 1640 V 69 8123
>12 1381386249 ced30000 33554432 1640 V 2332 5860
>79885 1381386250 d0dde000 33554432 1640 V 21 8171
>15 1381386251 d2dde000 33554432 1640 V 352 7840
>16 1381386252 d4dde000 33554432 1640 V 139 8053
>526 1381386253 d6dde000 33554432 1640 V 34 8158
>17 1381386254 d8dde000 33554432 1640 V 243 7949
>Total: - - 1197031424 - - 202193 90051
>
> (* segment locked in memory)
>
>but i am unsure what i am looking for to determine if we have multiple
>virtual memory segments ?
>
>Thanks!
>Tom
>
>
>mosserp@wellsfargo.com wrote in message news:<bcbq0i$esc$1@terabinaries.xmission.com>...
>> Just a stab in the dark, but are you running HDR in synchronous mode
>> (DRINTERVAL=-1)? If so, that *could* slow things down. If so, I suggest
>> asynch mode.
>>
>> Another thing -- since you're running an HCx version, I assume that this is
>> on HP-UX, but the HC version runs in 32-bit mode. So you're still subject
>> to the wonderful limitations of HP-UX memory management. Do you have
>> multiple virtual memory segments (onstat -g seg)? That definitely hurts
>> performance on HP-UX.
>>
>> HTH,
>> Paul Mosser
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g