Re: Long checkpoint woe
Posted in 2006
Neil Truby reported very long Informix checkpoints on a Solaris box with a Sun StorEdge 3310 JBOD: dirty pages flushed at only ~1,000/s, so ~20MB took ~10s. Others pushed him to test the disks directly with dd; his raw-device 2K-block reads/writes took around a minute for 20MB, while Clive Eisen's laptop (SATA, raw partition and filesystem) did the same in well under a second, suggesting a storage/firmware issue rather than an Informix one. Neil noted 256K blocks were fast and still suspected Informix. No resolution is recorded in the thread.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Logging & Checkpoints
Neil Truby wrote: > "Roy Mercer" <roy.mercer@gmail.com> wrote in message > news:1143575685.269708.17040@v46g2000cwv.googlegroups.com... >> Use /dev/null for if and then for of. >> This will tell you if it is a read issue or write issue. >> If they are the same then it is a real problem with the disks. > It's a Sun StorEdge 3310 in JBOD format. I don;t think it is a hardware > problem (as I keep telling IBM too) because it's the same result on two > servers/arrays 300 miles apart. > > I'm afraid I don't agree ( assuming that nothing else is reading/writing to these spindles ) I just ran a test on my shiny new laptop ( sata ) dd if=/dev/zero of=test bs=2k count=10240 10240+0 records in 10240+0 records out 20971520 bytes (21 MB) copied, 0.199701 seconds, 105 MB/s clive@linux:/tmp> dd if=test of=/dev/null bs=2k 10240+0 records in 10240+0 records out 20971520 bytes (21 MB) copied, 0.032379 seconds, 648 MB/s You are at least a factor of 100 out in performance -and this is a laptop - call Sun? I do know that there are disk firmware problems with the 3510 series - yes the disks, not the controllers. They aren't by any chance around 2 years old are they? nNil - If you contact me off-usenet I can give chapter and verse :-) -- Clive
"Clive Eisen" <clive@serendipita.com> wrote in message news:4429aedf$0$5003$db0fefd9@news.zen.co.uk... > Neil Truby wrote: >> "Roy Mercer" <roy.mercer@gmail.com> wrote in message >> news:1143575685.269708.17040@v46g2000cwv.googlegroups.com... >>> Use /dev/null for if and then for of. >>> This will tell you if it is a read issue or write issue. >>> If they are the same then it is a real problem with the disks. >> It's a Sun StorEdge 3310 in JBOD format. I don;t think it is a hardware >> problem (as I keep telling IBM too) because it's the same result on two >> servers/arrays 300 miles apart. > I'm afraid I don't agree ( assuming that nothing else is reading/writing > to these spindles ) > > I just ran a test on my shiny new laptop ( sata ) > > dd if=/dev/zero of=test bs=2k count=10240 > 10240+0 records in > 10240+0 records out > 20971520 bytes (21 MB) copied, 0.199701 seconds, 105 MB/s > clive@linux:/tmp> dd if=test of=/dev/null bs=2k > 10240+0 records in > 10240+0 records out > 20971520 bytes (21 MB) copied, 0.032379 seconds, 648 MB/s > > You are at least a factor of 100 out in performance -and this is a > laptop - call Sun? That's different though - these are file syste files. On the Sun: # time dd if=/dev/zero of=/scratch/test bs=2k count=10240 10240+0 records in 10240+0 records out real 0m0.31s user 0m0.01s sys 0m0.24s # dd if=/scratch/test of=/dev/null bs=2k 10240+0 records in 10240+0 records out real 0m0.51s user 0m0.00s sys 0m0.19s
Neil Truby said: >> You are at least a factor of 100 out in performance -and this is a >> laptop - call Sun? > > That's different though - these are file syste files. On the Sun: > > # time dd if=/dev/zero of=/scratch/test bs=2k count=10240 > 10240+0 records in > 10240+0 records out > > real 0m0.31s > user 0m0.01s > sys 0m0.24s > # dd if=/scratch/test of=/dev/null bs=2k > 10240+0 records in > 10240+0 records out > > real 0m0.51s > user 0m0.00s > sys 0m0.19s So you're saying that cooked file access is faster than raw disk access? And you still think the problem is in Informix? -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
Neil Truby wrote: > "Clive Eisen" <clive@serendipita.com> wrote in message > news:4429aedf$0$5003$db0fefd9@news.zen.co.uk... >> Neil Truby wrote: > > That's different though - these are file syste files. On the Sun: > > # time dd if=/dev/zero of=/scratch/test bs=2k count=10240 > 10240+0 records in > 10240+0 records out > > real 0m0.31s > user 0m0.01s > sys 0m0.24s > # dd if=/scratch/test of=/dev/null bs=2k > 10240+0 records in > 10240+0 records out > > real 0m0.51s > user 0m0.00s > sys 0m0.19s > > > Fair point - however dd if=/dev/zero of=/dev/sda2 bs=2k count=10240 10240+0 records in 10240+0 records out 20971520 bytes (21 MB) copied, 0.706315 seconds, 29.7 MB/s linux:~ # dd if=/dev/sda2 of=/dev/null bs=2k count=10240 10240+0 records in 10240+0 records out 20971520 bytes (21 MB) copied, 0.627349 seconds, 33.4 MB/s Stil my little sata based laptop knocking the spots off your Sun - and the same spindle is running the operating system -- Clive
Obnoxio The Clown wrote: > Neil Truby said: > So you're saying that cooked file access is faster than raw disk access? > And you still think the problem is in Informix? > yes - it is faster for 20M - you are just writing to/reading from the cache(s) if you did the same test with 2G to 20G I suspect you would see a different result depending how much ram is in the box - however see my next post.
> Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message > news:mailman.201.1143584365.18205.informix-list@iiug.org... > > Neil Truby said: >>> You are at least a factor of 100 out in performance -and this is a >>> laptop - call Sun? >> >> That's different though - these are file syste files. On the Sun: >> >> # time dd if=/dev/zero of=/scratch/test bs=2k count=10240 >> 10240+0 records in >> 10240+0 records out >> >> real 0m0.31s >> user 0m0.01s >> sys 0m0.24s >> # dd if=/scratch/test of=/dev/null bs=2k >> 10240+0 records in >> 10240+0 records out >> >> real 0m0.51s >> user 0m0.00s >> sys 0m0.19s > > So you're saying that cooked file access is faster than raw disk access? > And you still think the problem is in Informix? Actually, deep down, I do. This config restoes an 80g database in about 20-25m - pretty good I think. It's the long checkpoints that are the worry.
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message news:mailman.201.1143584365.18205.informix-list@iiug.org... > And you still think the problem is in Informix? To recap: because of LRU min and max settings the number of dirty pages never gets much beyond 10,000. A checkpoint at this point will take ~10s - we see the dirty page copunt descend sloooooowwwwllllly from 10,000 to 0 at about 1,000/s. 10s to write 20m - that's awful.
Neil Truby said: >> Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message >> news:mailman.201.1143584365.18205.informix-list@iiug.org... >> >> Neil Truby said: >>>> You are at least a factor of 100 out in performance -and this is a >>>> laptop - call Sun? >>> >>> That's different though - these are file syste files. On the Sun: >>> >>> # time dd if=/dev/zero of=/scratch/test bs=2k count=10240 >>> 10240+0 records in >>> 10240+0 records out >>> >>> real 0m0.31s >>> user 0m0.01s >>> sys 0m0.24s >>> # dd if=/scratch/test of=/dev/null bs=2k >>> 10240+0 records in >>> 10240+0 records out >>> >>> real 0m0.51s >>> user 0m0.00s >>> sys 0m0.19s >> >> So you're saying that cooked file access is faster than raw disk access? >> And you still think the problem is in Informix? > > Actually, deep down, I do. This config restoes an 80g database in about > 20-25m - pretty good I think. It's the long checkpoints that are the > worry. But a 20MB disk to disk copy without touching Informix takes a minute. That's a crazy long time. -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
Neil Truby said: > "Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message > news:mailman.201.1143584365.18205.informix-list@iiug.org... > >> And you still think the problem is in Informix? > > To recap: because of LRU min and max settings the number of dirty pages > never gets much beyond 10,000. A checkpoint at this point will take ~10s > - > we see the dirty page copunt descend sloooooowwwwllllly from 10,000 to 0 > at > about 1,000/s. 10s to write 20m - that's awful. But the OS is taking 61 s to read then write 20M? -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message news:mailman.203.1143610892.18205.informix-list@iiug.org... > > Neil Truby said: >> "Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message >> news:mailman.201.1143584365.18205.informix-list@iiug.org... >> >>> And you still think the problem is in Informix? >> >> To recap: because of LRU min and max settings the number of dirty pages >> never gets much beyond 10,000. A checkpoint at this point will take >> ~10s >> - >> we see the dirty page copunt descend sloooooowwwwllllly from 10,000 to 0 >> at >> about 1,000/s. 10s to write 20m - that's awful. > > But the OS is taking 61 s to read then write 20M? With a 2k block size, yes. With a 256k block size it takes less than a second. Which sort of goes back to my question: is the unusual/relevant?
Neil Truby said: > With a 2k block size, yes. With a 256k block size it takes less than a > second. Which sort of goes back to my question: is the unusual/relevant? Well, your mighty database server's disk IO rates have been trashed by a couple of laptops using the same block size of 2K. I can't really see why we're still having this discussion, to be honest. -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.