Re: long checkpoint problem
Posted in 2004
Looking at your stats:
Alex V. Spirin wrote:
> andykent.bristol@virgin.net (Andy Kent) wrote in message news:<a3525a11.0401290307.42f6e997@posting.google.com>...
>
>>It's quite plausible that page cleaning would lag behind buffer writes
>>during a bulk load. Disk i/o is obviously inherently much, much slower
>>than buffer i/o. It's likely to be slower if:
>>
>>- Everything's being written out to one disk or volume (e.g. the table
>>doesn't use fragmentation)
>
> Yes! It was disk with temporary dbspace.
> device %busy avque r+w/s blks/s avwait avserv (-d)
> 16:20:00 Sdsk-4 99.21 1.00 98.09 722.96 0.00 10.11
> 16:40:00 Sdsk-4 100.00 1.00 90.63 742.17 0.00 15.09
> 17:20:00 Sdsk-4 100.00 1.00 185.21 1486.84 0.00 9.00
> 17:40:00 Sdsk-4 100.00 1.00 225.13 1217.06 0.00 9.65
These drives are being overwhelmed! You need to add more spindles to
the system and turn your descrete mirrors into a RAID10 array or if
mirrors is as far as you can go then fragment the tables (including
temp tables) across these and the new mirrors. I'd say you need at
least two more pair but perhaps three.
> and long checkpoint:
> 16:23:38 Checkpoint Completed: duration was 17 seconds.
> 17:10:34 Checkpoint Completed: duration was 20 seconds.
> 17:23:41 Checkpoint Completed: duration was 55 seconds.
> 17:31:17 Checkpoint Completed: duration was 28 seconds.
Increase CLEANERS to at least 4 (the manuals recommend, as Malcolm
suggested, that CLEANERS == spindles) and preferably the same as LRUS
(my own preference). So I'd set CLEANERS 128.
>>- CLEANERS != no. of disks it's writing to
>
> Yes. There are 3 disks, each is RAID0 (mirror), for data and 1 - for
> tempdbs. But only 2 physical CPU...
>
<SNIP>
Notice below that the io/wup for EVERY aio VP is > 1.0. This
indicates that 10-20% of the time an aio VP is awakened to work it
finds another page to work on once it finishes its initial assignment.
This means that that second page had to wait for an available aio VP
slowing IO performance. Increase NUMAIOVPS until at least one always
shows an io/wup value in onstat -g iov that is less than 1.0. When
more than a few show <1.0 you have more than you need.
> AIO I/O vps:
> class/vp s io/s totalops dskread dskwrite dskcopy wakeups io/wup
> errors
> msc 0 i 0.0 439 0 0 0 438 1.0
> 0
> aio 0 i 13.2 1000139 922709 77286 0 875052 1.1
> 0
> aio 1 i 12.1 913225 838889 74326 0 827833 1.1
> 0
> aio 2 i 11.3 857368 788895 68469 0 780730 1.1
> 0
> aio 3 i 10.7 808501 742739 65760 0 738572 1.1
> 0
> aio 4 i 10.1 765309 702871 62437 0 699782 1.1
> 0
> aio 5 i 9.7 730822 671013 59809 0 668452 1.1
> 0
> aio 6 i 9.3 705079 647127 57951 0 644157 1.1
> 0
> aio 7 i 9.0 680232 623834 56398 0 624003 1.1
> 0
> aio 8 i 8.7 661005 605972 55033 0 605157 1.1
> 0
> aio 9 i 8.5 640230 587107 53122 0 583304 1.1
> 0
> aio 10 i 8.1 614319 562517 51802 0 532505 1.2
> 0
> aio 11 i 7.7 579855 529708 50147 0 455235 1.3
> 0
> aio 12 i 7.2 547176 498112 49063 0 411269 1.3
> 0
> aio 13 i 6.7 509861 461989 47872 0 416278 1.2
> 0
> aio 14 i 6.5 494899 447238 47661 0 419991 1.2
> 0
> aio 15 i 6.5 494212 444279 49933 0 407697 1.2
> 0
> pio 0 i 0.0 852 0 852 0 852 1.0
> 0
> lio 0 i 0.0 1985 0 1985 0 1985 1.0
> 0
<SNIP>
Art S. Kagel