Long checkpoint woe
Posted in 2006
Topics: Storage & Space Management, Stored Procedures & SPL, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
IDS 10.0FC3 on Solaris 9. OLTP system. Our system shows very long checkpoints (typically up to 10s, all day. Can be up to a minute) for quite small amounts of dirty pages. We have lru_min/max set so that the number of dirty ones never exceeds about 10,000. As a checkpoint occurs we see through our monitor script the dirty page count descend very slowly to 0 at a rate of about 1,000/s. (When I issued a checkpoint during an ALTER FRAGMENT of a 4g table after I'd set lru_min/max to 60 and 70, it took over half an hour!!). I've had this open with IBM tech support for months. They've had me check hardware, OS configuration etc etc, all of which have checked out fine, but can't see a problem with Informix. Until the weekend we had one large 80g dbspaces, but we took an outage at the weekend to split the busiest tables in terms of writes out to separate dbspaces and now have a fairly even balance over 5 dbspaces. It's made no difference at all The IBM tech guy got me to create 2 20m slices then copy between them with a 2k page size. It does take a surprising amount of time with such a small blocksize: # time dd if=/dev/md/rdsk/d121 of=/dev/md/rdsk/d122 bs=2k 10240+0 records in 10240+0 records out real 1m4.04s user 0m0.03s sys 0m0.64s But id this really indicative of a problem? Any other observations about the bigger Informix issue? thanks -- Neil Truby t:01932 724027 Director m:07798 811708 Ardenta Limited e:neil.truby@ardenta.com
Neil Truby said: > real 1m4.04s Bloody Norah!!! A minute to copy 20M from one disk to another? Good grief! I'd guess that your disks are the problem from that... -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
Obnoxio The Clown wrote: > Neil Truby said: >> real 1m4.04s > > Bloody Norah!!! A minute to copy 20M from one disk to another? Good grief! > > I'd guess that your disks are the problem from that... > I concur with The Clown - you have one major hardware/performance problem - I had an 8086 that could do better than that :-) Care to share Storage technology Storage topology What else is running apart from dd when you did the test -- Clive
Use /dev/null for if and then for of. This will tell you if it is a read issue or write issue. If they are the same then it is a real problem with the disks. Some of these storage arrays have the cache turned on for reads but off or tuned low for writes. I have also noticed that if the backup battery for the cache is low or being charged then the disk write cache is disabled. Post more information about the storage array. Thanks
"Roy Mercer" <roy.mercer@gmail.com> wrote in message news:1143575685.269708.17040@v46g2000cwv.googlegroups.com... > Use /dev/null for if and then for of. > This will tell you if it is a read issue or write issue. > If they are the same then it is a real problem with the disks. # time dd if=/dev/md/rdsk/d121 of=/dev/null bs=2k 10240+0 records in 10240+0 records out real 0m14.24s user 0m0.01s sys 0m0.36s # time dd of=/dev/md/rdsk/d121 if=/dev/null bs=2k 0+0 records in 0+0 records out real 0m0.02s user 0m0.00s sys 0m0.01s > Post more information about the storage array. It's a Sun StorEdge 3310 in JBOD format. I don;t think it is a hardware problem (as I keep telling IBM too) because it's the same result on two servers/arrays 300 miles apart.