Re: Long checkpoint woe
Posted in 2006
Thread continues an investigation into very long IDS checkpoints (up to ~258s) on Solaris/Sun hardware. Testing with dd showed raw-device writes at 2K block size are extremely slow (~60s for 20MB), but Sun reproduced identical results on several servers (E250 to E4900) and on other sites' EMC-attached kit, so it looks like default Solaris raw-write behaviour rather than a fault. Writes via character devices, or with 16K pages (checkpoints dropped to ~37s), were much faster. Participants disagreed over whether this explains the checkpoint problem; no definitive fix or root cause is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Logging & Checkpoints
"TBP" <TheBigPotato@NotHere.Co.Uk> wrote in message news:AL5Xf.30370$5B4.982@newsfe6-gui.ntli.net... > Neil Truby wrote: >> "Nog" <noel@gfm.co.uk> wrote in message >> news:1143751527.646087.192390@e56g2000cwe.googlegroups.com... >>> Are both primary and secondary raided the same way? are these the >>> machines that are miles apart? The 3310 are directly attached (scsi?) >>> aren't they? >> >> The live one is exclusively RAID 1+0. The DR/DEV one has a mixture of >> RAID 1+0 and RAID 0. But the dd output is the same whether I use a >> mirrored logical volume, unmirrored LV, write directly to a device on the >> external array or write directly to one of the internal disks. >> >> I've raised a case with Sun. My hunch is they will come back and say >> it's expected behaviour, > > What, crap I/O rates at 2k block size is "expected behaviour"? Well, they did come back and say exactly what I'd expected: that they have tested it on 3 servers there (as I have tested it on 2 here), from a E250 up to an E4900, and the results were the same as mine, ie over a minute in every case: # time dd if=/dev/zero of=/dev/md/rdsk/d122 bs=2k count=10240 10240+0 records in 10240+0 records out real 1m0.07s user 0m0.01s sys 0m0.32s They did raise a couple of interesting points: 1. The dd is much faster (by a factor of 8) if the character, rather than raw device is used. Obviously we've been through that one at length and are sure we should we be using the raw devices. 2, The dd is also much, much faster (by a factor of 20). I did try creating a 16k page dbspace, and doing an ALTER FRAGMENT of a 3g table was much faster (8m instead of 16) and checkpoint maximums of 37s instead of 258, than the 2k dbspace. But you might expect a sequential operation to be faster with a larger blocksize, and it possibly (probably?) would not be applicable to an OLTP system. And anyway, you'd expect hundreds of other users to be similarly affected, which they aren't. They also said that you don;t see this on IDE. So I conclude that it's a red herring.
Neil Truby said: > Well, they did come back and say exactly what I'd expected: that they have > tested it on 3 servers there (as I have tested it on 2 here), from a E250 > up > to an E4900, and the results were the same as mine, ie over a minute in > every case: > > # time dd if=/dev/zero of=/dev/md/rdsk/d122 bs=2k count=10240 > 10240+0 records in > 10240+0 records out > > real 1m0.07s > user 0m0.01s > sys 0m0.32s > > They did raise a couple of interesting points: > 1. The dd is much faster (by a factor of 8) if the character, rather than > raw device is used. Obviously we've been through that one at length and > are > sure we should we be using the raw devices. > 2, The dd is also much, much faster (by a factor of 20). I did try > creating > a 16k page dbspace, and doing an ALTER FRAGMENT of a 3g table was much > faster (8m instead of 16) and checkpoint maximums of 37s instead of 258, > than the 2k dbspace. But you might expect a sequential operation to be > faster with a larger blocksize, and it possibly (probably?) would not be > applicable to an OLTP system. And anyway, you'd expect hundreds of other > users to be similarly affected, which they aren't. > > They also said that you don;t see this on IDE. > > So I conclude that it's a red herring. It's a red herring that the disk performs badly with a block size of 2K, even if Informix isn't any part of the equation? What kind of herring would that be, salted or unsalted? -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message news:mailman.223.1143820711.18205.informix-list@iiug.org... >> They did raise a couple of interesting points: >> 1. The dd is much faster (by a factor of 8) if the character, rather >> than >> raw device is used. Obviously we've been through that one at length and >> are >> sure we should we be using the raw devices. >> 2, The dd is also much, much faster (by a factor of 20). I did try >> creating >> a 16k page dbspace, and doing an ALTER FRAGMENT of a 3g table was much >> faster (8m instead of 16) and checkpoint maximums of 37s instead of 258, >> than the 2k dbspace. But you might expect a sequential operation to be >> faster with a larger blocksize, and it possibly (probably?) would not be >> applicable to an OLTP system. And anyway, you'd expect hundreds of other >> users to be similarly affected, which they aren't. >> >> They also said that you don;t see this on IDE. >> >> So I conclude that it's a red herring. > > It's a red herring that the disk performs badly with a block size of 2K, > even if Informix isn't any part of the equation? What kind of herring > would that be, salted or unsalted? It's a red herring in the sense that it's behaved the same on both boxes I've tried, and all 3 that the Sun engineer tried, so it's a reasonable chance that it's default behaviour, yet there are not armies of other IDS users jumping up and down complaining about performance on Sun. So, therefore, whilst interesting in itself it is not an indicator of poor Informix performance. I'd be interested to see the dd results people with a caxhed array get from running the test on internal disk (presumably the same as me) against a raw device on their SAN. IDE drives do not exhibit this behaviour btw.
"Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message news:mailman.223.1143820711.18205.informix-list@iiug.org... > >> So I conclude that it's a red herring. > > It's a red herring that the disk performs badly with a block size of 2K, > even if Informix isn't any part of the equation? What kind of herring > would that be, salted or unsalted? Here's the result from one of the UK's largest Informix users, with huge Sun boxes attached to an EMC Clariion: time dd if=/dev/zero of=/dev/informix_links/appdbs4_6 bs=2k count=10240 10240+0 records in 10240+0 records out real 0m21.213s user 0m0.010s sys 0m0.700s 3 times the speed of mine (due to RAID cacheing I'd guess) but still damned slow. Still think it's just my kit ...?
Neil Truby said: > IDE drives do not exhibit this behaviour btw. Well, then... -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
Neil Truby wrote: > "Obnoxio The Clown" <obnoxio@serendipita.com> wrote in message > news:mailman.223.1143820711.18205.informix-list@iiug.org... > >>> So I conclude that it's a red herring. >> It's a red herring that the disk performs badly with a block size of 2K, >> even if Informix isn't any part of the equation? What kind of herring >> would that be, salted or unsalted? > > Here's the result from one of the UK's largest Informix users, with huge Sun > boxes attached to an EMC Clariion: > > time dd if=/dev/zero of=/dev/informix_links/appdbs4_6 bs=2k count=10240 > 10240+0 records in > 10240+0 records out > > real 0m21.213s > user 0m0.010s > sys 0m0.700s > > 3 times the speed of mine (due to RAID cacheing I'd guess) but still damned > slow. > > Still think it's just my kit ...? Try threatening them with a migration to Oracle. That ought to get you the support you need. Those checkpoint times are outrageous. -- Daniel A. Morgan http://www.psoug.org damorgan@x.washington.edu (replace x with u to respond)
Neil Truby said: > Still think it's just my kit ...? I don't think it's just YOUR kit, but if the basic write performance with a 2K page size is slow to start with, how do you think that writing to the disk using a database will be any faster? -- Bye now, Obnoxio Information within this post contains forward looking statements within the meaning of Section 27A of the Securities Act of 1933 and Section 21B of the S E C Act of 1934. Statements that involve discussions with respect to projections of future events are not statements of historical fact and may be forward looking statements. Don't rely on them to make a decision. The poster is not a reporting company registered under the Exchange Act of 1934. I have received a life peerage from Her Majesty, who is not an officer, minister or affiliate Labour party member. I intend to recover my loan now, which could cause the parliamentary majority to go down, resulting in losses for you. Today's Labour party has: an accumulated deficit and a reliance on loans from officers and affiliates to pay expenses. It is not an operating political party. The party is going to need financing to continue as a going concern. A failure to finance could cause the party to go out of business. This report shall not be construed as any kind of investment advice or solicitation. You can lose all your money by investing in this party.
With Solaris 5.8, Sun Enterprise 250. Same results than you, Neil. # time dd if=/dev/rdsk/c0t8d0s6 of=/dev/null bs=2k count=10240 10240+0 registros dentro 10240+0 registros fuera real 1.7 user 0.0 sys 0.3 # time dd of=/dev/rdsk/c0t8d0s6 if=/dev/zero bs=2k count=10240 10240+0 registros dentro 10240+0 registros fuera real 1:02.0 user 0.0 sys 1.0 I can probe with Solstice DiskSuite metadevices if you want. Guillermo Gómez Valcárcel GRAFIA S.A. -----Mensaje original----- De: informix-list-bounces@iiug.org [mailto:informix-list-bounces@iiug.org] En nombre de Neil Truby Enviado el: viernes, 31 de marzo de 2006 17:47 Para: informix-list@iiug.org Asunto: Re: Long checkpoint woe "TBP" <TheBigPotato@NotHere.Co.Uk> wrote in message news:AL5Xf.30370$5B4.982@newsfe6-gui.ntli.net... > Neil Truby wrote: >> "Nog" <noel@gfm.co.uk> wrote in message >> news:1143751527.646087.192390@e56g2000cwe.googlegroups.com... >>> Are both primary and secondary raided the same way? are these the >>> machines that are miles apart? The 3310 are directly attached (scsi?) >>> aren't they? >> >> The live one is exclusively RAID 1+0. The DR/DEV one has a mixture of >> RAID 1+0 and RAID 0. But the dd output is the same whether I use a >> mirrored logical volume, unmirrored LV, write directly to a device on the >> external array or write directly to one of the internal disks. >> >> I've raised a case with Sun. My hunch is they will come back and say >> it's expected behaviour, > > What, crap I/O rates at 2k block size is "expected behaviour"? Well, they did come back and say exactly what I'd expected: that they have tested it on 3 servers there (as I have tested it on 2 here), from a E250 up to an E4900, and the results were the same as mine, ie over a minute in every case: # time dd if=/dev/zero of=/dev/md/rdsk/d122 bs=2k count=10240 10240+0 records in 10240+0 records out real 1m0.07s user 0m0.01s sys 0m0.32s They did raise a couple of interesting points: 1. The dd is much faster (by a factor of 8) if the character, rather than raw device is used. Obviously we've been through that one at length and are sure we should we be using the raw devices. 2, The dd is also much, much faster (by a factor of 20). I did try creating a 16k page dbspace, and doing an ALTER FRAGMENT of a 3g table was much faster (8m instead of 16) and checkpoint maximums of 37s instead of 258, than the 2k dbspace. But you might expect a sequential operation to be faster with a larger blocksize, and it possibly (probably?) would not be applicable to an OLTP system. And anyway, you'd expect hundreds of other users to be similarly affected, which they aren't. They also said that you don;t see this on IDE. So I conclude that it's a red herring. _______________________________________________ Informix-list mailing list Informix-list@iiug.org http://www.iiug.org/mailman/listinfo/informix-list __________ Informacisn de NOD32, revisisn 1.1465 (20060331) __________ Este mensaje ha sido analizado con NOD32 antivirus system http://www.nod32.com
Thanks. I'm quite torn, because the IBM UK team (and Obnoxio!) have highlighted this very slow write rate as an issue we must get resolved with Sun, and I fully understand why they are saying that. Yet it's clearly default behaviour so why aren't Sun / Informix users around the world squealling ....? regards -- Neil Truby t:01932 724027 Director m:07798 811708 Ardenta Limited e:neil.truby@ardenta.com "Guillermo G'mez Valc'rcel" <guillermo.gomez@grafia.es> wrote in message news:mailman.239.1144084760.18205.informix-list@iiug.org... With Solaris 5.8, Sun Enterprise 250. Same results than you, Neil. # time dd if=/dev/rdsk/c0t8d0s6 of=/dev/null bs=2k count=10240 10240+0 registros dentro 10240+0 registros fuera real 1.7 user 0.0 sys 0.3 # time dd of=/dev/rdsk/c0t8d0s6 if=/dev/zero bs=2k count=10240 10240+0 registros dentro 10240+0 registros fuera real 1:02.0 user 0.0 sys 1.0 I can probe with Solstice DiskSuite metadevices if you want. Guillermo G'mez Valc'rcel GRAFIA S.A. -----Mensaje original----- De: informix-list-bounces@iiug.org [mailto:informix-list-bounces@iiug.org] En nombre de Neil Truby Enviado el: viernes, 31 de marzo de 2006 17:47 Para: informix-list@iiug.org Asunto: Re: Long checkpoint woe "TBP" <TheBigPotato@NotHere.Co.Uk> wrote in message news:AL5Xf.30370$5B4.982@newsfe6-gui.ntli.net... > Neil Truby wrote: >> "Nog" <noel@gfm.co.uk> wrote in message >> news:1143751527.646087.192390@e56g2000cwe.googlegroups.com... >>> Are both primary and secondary raided the same way? are these the >>> machines that are miles apart? The 3310 are directly attached (scsi?) >>> aren't they? >> >> The live one is exclusively RAID 1+0. The DR/DEV one has a mixture of >> RAID 1+0 and RAID 0. But the dd output is the same whether I use a >> mirrored logical volume, unmirrored LV, write directly to a device on the >> external array or write directly to one of the internal disks. >> >> I've raised a case with Sun. My hunch is they will come back and say >> it's expected behaviour, > > What, crap I/O rates at 2k block size is "expected behaviour"? Well, they did come back and say exactly what I'd expected: that they have tested it on 3 servers there (as I have tested it on 2 here), from a E250 up to an E4900, and the results were the same as mine, ie over a minute in every case: # time dd if=/dev/zero of=/dev/md/rdsk/d122 bs=2k count=10240 10240+0 records in 10240+0 records out real 1m0.07s user 0m0.01s sys 0m0.32s They did raise a couple of interesting points: 1. The dd is much faster (by a factor of 8) if the character, rather than raw device is used. Obviously we've been through that one at length and are sure we should we be using the raw devices. 2, The dd is also much, much faster (by a factor of 20). I did try creating a 16k page dbspace, and doing an ALTER FRAGMENT of a 3g table was much faster (8m instead of 16) and checkpoint maximums of 37s instead of 258, than the 2k dbspace. But you might expect a sequential operation to be faster with a larger blocksize, and it possibly (probably?) would not be applicable to an OLTP system. And anyway, you'd expect hundreds of other users to be similarly affected, which they aren't. They also said that you don;t see this on IDE. So I conclude that it's a red herring. _______________________________________________ Informix-list mailing list Informix-list@iiug.org http://www.iiug.org/mailman/listinfo/informix-list __________ Informacisn de NOD32, revisisn 1.1465 (20060331) __________ Este mensaje ha sido analizado con NOD32 antivirus system http://www.nod32.com