Performance Impact after storage migration
Posted in 2013
After moving IDS 11.50 on HP-UX to new servers and a new SAN, the poster saw intermittent slowdowns, long blocking checkpoints and few scheduled checkpoints after setting RTO_SERVER_RESTART=180. Art Kagel suggested checking onstat -F (chunk vs LRU writes), lowering lru_min/max_dirty to increase LRU writing, and questioned the SAN layout: RAID 5 for database VGs, over-large 128K/256K stripe sizes, 10K drives, and everything wide-striped across 168 spindles; he also explained reading service times from onstat -g iof (should be under ~20ms). The fix came from HP: hardcoding the port speed to 4Gb and enabling Adaptive Optimization restored fast performance, though which change helped was not determined.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Performance & Tuning, Installation, Setup & Upgrades, Triggers, Constraints & Referential Integrity, Logging & Checkpoints, Migration, Import/Export & Data Conversion, Platform-Specific Issues
Hello All,
IDS 11.50.FC8W2 on HP-UX B.11.31 U ia64
Recently we have upgraded our storage and the hp boxes and since then we see
intermittent performance issues.
onstat -g ckpAUTO_CKPTS=On RTO_SERVER_RESTART=180 seconds Estimated recovery time 1304
seconds
We observed that the checkpoint was taking long time and blocking the
transactions hence we set the RTO_SERVER_RESTART to 180
post this the AUTO_CKPTS was ON and now I see there is very very less number
of checkpoint except for the checkpoints triggered
by backup so I believe most of the writing is done by LRU flushing
Is it normal not to have regular interval checkpoints ? could the excessing
LRU flusing is causing the intermittent performance problems?
onstat -c |grep ^BUFFERBUFFERPOOL
size=2K,buffers=9000000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
BUFFERPOOL
size=4K,buffers=3000000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
BUFFERPOOL
size=8K,buffers=1200000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
BUFFERPOOL
size=16K,buffers=3500000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
onstat -g segSegment Summary:
id key addr size ovhd class blkused blkfree
32789 52564803 c000000015380000 348160 5296 M 84 1
65551 52564801 c000000100000000 77832777728 912540720 R* 19002141 2
393236 52564802 c000001400000000 7372800000 86401888 V* 1127853 672147
Total: - - 85205925888 - - 20130078 672150
onstat -pProfile
dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
3533376905 40019324337 129263240373 97.27 119250498 619708349 2931928150 95.94
isamtot open start read write rewrite delete commit rollbk
168657611197 15194194925 5119943273 78536749289 1844472464 29194093 56456814
32472838 8777
gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
0 0 0 0 0 0 0
ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
0 0 0 987815.69 148729.15 1685 3891180
bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
420005993 316802 28144601035 0 0 61682 104707403 2348872108
ixda-RA idx-RA da-RA RA-pgsused lchwaits
1053637354 85987054 623849831 1727787827 84014387
Vikas:
Onstat -F will tell you whether LRU writes or Chunk writes are dominating.
With your lru_min/max_dirty settings at 1&3 I'd also guess that LRU writes
are dominating. Normally on v11.50+ with non-blocking checkpoints, most
sessions are not affected by long checkpoints, but it does happen. Before
you turned on the auto checkpointing, did the onstat -g ckp show long block
times?
You say that you "upgraded" your storage. That usually means a bigger
badder SAN. If you now have all of your chunks, including logical and
physical logs, rootdb, temp spaces, and your data, all on a single
monstrous structure on the SAN you may be experiencing IO bottlenecks
during the checkpoints and, also during periods of more intense LRU
flushing, that is affecting the engine's ability to move data into the
cache when it is needed. Similarly, how was(were) the SAN structure(s) for
the instance configured? What RAID level? How big is the stripe blocking?
How many spindles are in each structure? Is it possible that the same
sets of spindles are being shared with other applications that have
different access patterns than the database (Windows filesystems and
especially mail servers are especially bad to share with for example). How
are your read and write IO service times as the engine sees them (onstat -g
iof)?
Art
Art S. Kagel, Principal Consultant
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Wed, Oct 9, 2013 at 3:47 AM, VIKAS HIVARKAR <vikas.hivarkar@gmail.com>wrote:
> Hello All,
>
> IDS 11.50.FC8W2 on HP-UX B.11.31 U ia64
>
> Recently we have upgraded our storage and the hp boxes and since then we
> see
> intermittent performance issues.
>
> onstat -g ckp> AUTO_CKPTS=On RTO_SERVER_RESTART=180 seconds Estimated recovery time 1304
> seconds
>
> We observed that the checkpoint was taking long time and blocking the
> transactions hence we set the RTO_SERVER_RESTART to 180
> post this the AUTO_CKPTS was ON and now I see there is very very less
> number
> of checkpoint except for the checkpoints triggered
> by backup so I believe most of the writing is done by LRU flushing
>
> Is it normal not to have regular interval checkpoints ? could the excessing
> LRU flusing is causing the intermittent performance problems?
>
> onstat -c |grep ^BUFFER> BUFFERPOOL>
>
size=2K,buffers=9000000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
> BUFFERPOOL>
>
size=4K,buffers=3000000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
> BUFFERPOOL>
>
size=8K,buffers=1200000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
> BUFFERPOOL>
>
size=16K,buffers=3500000,lrus=512,lru_min_dirty=1.000000,lru_max_dirty=3.000000
>
> onstat -g seg> Segment Summary:
> id key addr size ovhd class blkused blkfree
> 32789 52564803 c000000015380000 348160 5296 M 84 1
> 65551 52564801 c000000100000000 77832777728 912540720 R* 19002141 2
> 393236 52564802 c000001400000000 7372800000 86401888 V* 1127853 672147
> Total: - - 85205925888 - - 20130078 672150
>
> onstat -p> Profile
> dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
> 3533376905 40019324337 129263240373 97.27 119250498 619708349 2931928150
> 95.94
>
> isamtot open start read write rewrite delete commit rollbk
> 168657611197 15194194925 5119943273 78536749289 1844472464 29194093
> 56456814
> 32472838 8777
>
> gp_read gp_write gp_rewrt gp_del gp_alloc gp_free gp_curs
> 0 0 0 0 0 0 0
>
> ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
> 0 0 0 987815.69 148729.15 1685 3891180
>
> bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
> 420005993 316802 28144601035 0 0 61682 104707403 2348872108
>
> ixda-RA idx-RA da-RA RA-pgsused lchwaits
> 1053637354 85987054 623849831 1727787827 84014387
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--089e013d14cadc215f04e84c96de
Hello Art,
Sorry for the delay in my reponse, I was waiting for the information from HP
Art - did the onstat -g ckp show long block times?
Yes, infact lately we have increased the BUFFER size for 16K buffer as we saw
high BTR and I believe that since that change I am seeing checkpoints taking
long time with fairly longer blocking time.
Configured with group of 168 FC 10k RPM drives and dedicated four front end
ports for Informix cluster.
What RAID level?
79 VGs in total out of which 61 are on RAID 1 and 18 are RAID 5
How big is the stripe blocking?
For Raid 5 stripe size is 128Kb
For Raid 1 stripe size is 256Kb
How many spindles are in each structure?
All luns are wide striped across 168 spindles.
onstat -F
IBM Informix Dynamic Server Version 11.50.FC8W2 -- On-Line (Prim) -- Up 9 days
10:14:13 -- 83208912 Kbytes
Fg Writes LRU Writes Chunk Writes
0 51712213 95335894
address flusher state data # LRU Chunk Wakeups Idle Tim
c000001408415888 0 I 0 0 3819542 4621106 813303.587
c0000014084160e0 1 I 0 0 3557889 4359644 813343.843
c000001408416938 2 I 0 0 3274470 4075463 812531.368
c000001408417190 3 I 0 0 3166100 3966471 811934.702
c0000014084179e8 4 I 0 0 3002538 3804280 813324.752
c000001408418240 5 I 0 0 2834845 3636310 813031.860
c000001408418a98 6 I 0 0 2686597 3486253 811202.465
c0000014084192f0 7 I 0 0 2557634 3358741 812642.450
c000001408419b48 8 I 0 0 2467357 3268927 813121.692
c00000140841a3a0 9 I 0 0 2423648 3225405 813342.546
c00000140841abf8 10 I 0 0 2375710 3175721 811535.485
c00000140841b450 11 I 0 0 2301315 3101566 811795.799
c00000140841bca8 12 I 0 0 2267774 3068964 812759.581
c00000140841c500 13 I 0 0 2195100 2996698 813156.343
c00000140841cd58 14 I 0 0 2075059 2876297 812764.397
c00000140841d5b0 15 I 0 0 1774294 2576683 813943.028
c00000140841de08 16 I 0 0 1757574 2560138 814113.722
c00000140841e660 17 I 0 0 1740367 2542914 814100.952
c00000140841eeb8 18 I 0 0 1715677 2518217 814078.609
c00000140841f710 19 I 0 0 1678186 2480560 813929.036
c00000140841ff68 20 I 0 0 1628619 2430915 813850.870
c0000014084207c0 21 I 0 0 1578856 2380834 813509.546
c000001408421018 22 I 0 0 1557423 2359718 813847.838
c000001408421870 23 I 0 0 1524951 2326763 813363.393
c0000014084220c8 24 I 0 0 1489123 2291543 813976.098
c000001408422920 25 I 0 0 1439096 2241394 813848.700
c000001408423178 26 I 0 0 1408349 2210684 813881.037
c0000014084239d0 27 I 0 0 1346739 2149156 813965.500
c000001408424228 28 I 0 0 1284856 2087248 813931.687
c000001408424a80 29 I 0 0 1212814 2015075 813809.127
c0000014084252d8 30 I 0 0 1148024 1950013 813522.639
c000001408425b30 31 I 0 0 1116655 1918692 813586.337
c000001408426388 32 I 0 0 1074229 1876200 813506.475
c000001408426be0 33 I 0 0 1026402 1828330 813454.487
c000001408427438 34 I 0 0 975538 1777522 813516.306
c000001408427c90 35 I 0 0 935815 1737636 813347.500
c0000014084284e8 36 I 0 0 867591 1669514 813447.230
c000001408428d40 37 I 0 0 818066 1619650 813138.467
c000001408429598 38 I 0 0 789886 1591734 813375.733
c000001408429df0 39 I 0 0 757005 1558901 813434.546
c00000140842a648 40 I 0 0 712937 1514772 813384.725
c00000140842aea0 41 I 0 0 654107 1455916 813360.628
c00000140842b6f8 42 I 0 0 639123 1440821 813233.113
c00000140842bf50 43 I 0 0 585283 1386602 812831.072
c00000140842c7a8 44 I 0 0 538219 1339512 812784.338
c00000140842d000 45 I 0 0 513150 1314294 812633.153
c00000140842d858 46 I 0 0 449294 1250415 812619.052
c00000140842e0b0 47 I 0 0 413833 1215221 812911.429
c00000140842e908 48 I 0 0 396667 1198171 813026.016
c00000140842f160 49 I 0 0 366924 1168836 813432.231
c00000140842f9b8 50 I 0 0 352521 1154531 813523.150
c000001408430210 51 I 0 0 334947 1136473 813038.820
c000001408430a68 52 I 0 0 317619 1119994 813893.778
c0000014084312c0 53 I 0 0 309327 1111570 813759.278
c000001408431b18 54 I 0 0 299992 1101864 813378.746
c000001408432370 55 I 0 0 281463 1083137 813189.738
c000001408432bc8 56 I 0 0 267427 1069605 813678.658
c000001408433420 57 I 0 0 257462 1059482 813543.164
c000001408433c78 58 I 0 0 233999 1036539 814069.121
c0000014084344d0 59 I 0 0 216667 1019007 813869.390
c000001408434d28 60 I 0 0 208636 1011150 814033.976
c000001408435580 61 I 0 0 199558 1002101 814067.447
c000001408435dd8 62 I 0 0 191899 994587 814209.011
c000001408436630 63 I 0 0 187226 989824 814125.894
c000001408436e88 64 I 0 0 181928 984465 814061.368
c0000014084376e0 65 I 0 0 177257 979813 814086.271
c000001408437f38 66 I 0 0 171837 974313 814006.752
c000001408438790 67 I 0 0 148002 950578 814103.220
c000001408438fe8 68 I 0 0 142534 945130 814124.732
c000001408439840 69 I 0 0 137263 939803 814061.076
c00000140843a098 70 I 0 0 126512 928923 813930.265
c00000140843a8f0 71 I 0 0 124673 926876 813726.060
c00000140843b148 72 I 0 0 104677 906280 813120.425
c00000140843b9a0 73 I 0 0 100333 902258 813452.502
c00000140843c1f8 74 I 0 0 98501 901189 814216.083
c00000140843ca50 75 I 0 0 86576 889207 814164.492
c00000140843d2a8 76 I 0 0 85357 887859 814038.557
c00000140843db00 77 I 0 0 83300 885932 814165.726
c00000140843e358 78 I 0 0 75558 877991 813950.859
c00000140843ebb0 79 I 0 0 63186 865910 814255.623
c00000140843f408 80 I 0 0 63039 865806 814295.974
c00000140843fc60 81 I 0 0 62133 864630 814020.365
c0000014084404b8 82 I 0 0 61008 863136 813651.043
c000001408440d10 83 I 0 0 52552 855215 814184.712
c000001408441568 84 I 0 0 52336 854973 814159.424
c000001408441dc0 85 I 0 0 40774 843556 814316.028
c000001408442618 86 I 0 0 34525 837235 814229.431
c000001408442e70 87 I 0 0 34327 837111 814311.499
c0000014084436c8 88 I 0 0 28096 830800 814226.639
c000001408443f20 89 I 0 0 27830 830520 814221.625
c000001408444778 90 I 0 0 27724 830457 814258.723
c000001408444fd0 91 I 0 0 27645 830370 814254.519
c000001408445828 92 I 0 0 27415 830089 814208.406
c000001408446080 93 I 0 0 27360 829859 814018.141
c0000014084468d8 94 I 0 0 13939 816265 813852.939
c000001408447130 95 I 0 0 13897 816490 814115.386
c000001408447988 96 I 0 0 3793 806212 813938.983
c0000014084481e0 97 I 0 0 3745 806441 814216.083
c000001408448a38 98 I 0 0 3700 806557 814382.270
c000001408449290 99 I 0 0 3640 806389 814276.829
c000001408449ae8 100 I 0 0 3411 806188 814305.173
c00000140844a340 101 I 0 0 1109 803833 814243.488
c00000140844ab98 102 I 0 0 1082 803407 813848.433
c00000140844b3f0 103 I 0 0 1058 803911 814376.240
c00000140844bc48 104 I 0 0 1038 803782 814269.765
c00000140844c4a0 105 I 0 0 1020 803781 814283.682
c00000140844ccf8 106 I 0 0 992 803779 814317.611
c00000140844d550 107 I 0 0 967 803688 814243.383
c00000140844dda8 108 I 0 0 951 803747 814323.007
c00000140844e600 109 I 0 0 924 803729 814327.078
c00000140844ee58 110 I 0 0 883 803648 814287.418
c00000140844f6b0 111 I 0 0 855 803512 814181.788
c00000140844ff08 112 I 0 7641 827 809705 812751.497
c000001408450760 113 I 0 4125 805 806920 813510.738
c000001408450fb8 114 I 0 10575 773 811401 811519.681
c000001408451810 115 I 0 7 748 803544 814311.271
c000001408452068 116 I 0 10319 718 811707 812178.285
c0000014084528c0 117 I 0 9742 696 811593 812652.787
c000001408453118 118 I 0 3249 661 806361 813981.246
c000001408453970 119 I 0 9357 637 811009 812524.289
c0000014084541c8 120 I 0 10162
See my comments below:
Art
Art S. Kagel, Principal Consultant
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Fri, Oct 11, 2013 at 7:38 AM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello Art,
>
> Sorry for the delay in my reponse, I was waiting for the information from
> HP
>
> Art - did the onstat -g ckp show long block times?
>
> Yes, infact lately we have increased the BUFFER size for 16K buffer as we
> saw
> high BTR and I believe that since that change I am seeing checkpoints
> taking
> long time with fairly longer blocking time.
>
OK, so although the lru_min/max_dirty levels or 1,3 looks OK on the
surface, because of the increased buffer pool there isn't enough LRU
writing going on. The -F output shows about 2/3 of writes are chunk writes
which is good for efficiency but not for checkpoint performance. I would
drop it again to 0.5 and 2.
>
> Configured with group of 168 FC 10k RPM drives and dedicated four front end
> ports for Informix cluster.
>
Single monstrous structure indeed! 15K drives would have been MUCH better.
Que sera sera!
>
> What RAID level?
> 79 VGs in total out of which 61 are on RAID 1 and 18 are RAID 5
>
You knew I was going to say this: NO RAID5!!!!!! Hopefully those VGs are
ONLY being used for low activity no-write dbspaces. If they are housing
your logs or high activity tables that's one source of the checkpoint
performance problem. You are aware that RAID5 is NOT SAFE for your data?
Check out my presentation from the 2012 IIUG Conference "Doing Storage
Right!" for details on the latest research on the subject which bears out
and support this assertion I've been screaming about for over 20 years. So
you are using single RAID1 pairs? That's good for isolation, but you could
probably improve performance by gathering those 61 RAID1 pairs into 12 x 5
pair RAID10 arrays and get improved performance with little degradation of
data safety if any.
>
> How big is the stripe blocking?
> For Raid 5 stripe size is 128Kb
> For Raid 1 stripe size is 256Kb
>
Whoa! Way to big. Informix normally writes either one page or 8 pages so
even your 16K page chunks are never writing more than 128K and the 2K
chunks never more than 16K at a time. That means that you could be
rewriting each logical block up to 8 times on the RAID5 and up to 16 times
on the RAID1. That 128K block for the RAID5 array is really writing 0.5GB
across the array (assuming a 5 drive array) for every 2, 4, 8, or 16K
logical write with the accompanying RAID5 write penalty no matter how much
cache you have.
>
> How many spindles are in each structure?
> All luns are wide striped across 168 spindles.
>
Wait, the storage techs have carved out the RAID1 and RAID5 VGs out of a
single 168 spindle stripe (RAID0)? Are they insane? So the RAID1 and
RAID5 redundancy is only VIRTUAL redundancy? That's like protecting
against unwanted pregnancy with withdrawal! Also, if this is so, you
really have no isolation at all since everything is on the same massive
stripe.
Am I missing something?
>
> onstat -F>
> IBM Informix Dynamic Server Version 11.50.FC8W2 -- On-Line (Prim) -- Up 9> days
> 10:14:13 -- 83208912 Kbytes
>
> Fg Writes LRU Writes Chunk Writes
> 0 51712213 95335894
>
> address flusher state data # LRU Chunk Wakeups Idle Tim
> c000001408415888 0 I 0 0 3819542 4621106 813303.587
> c0000014084160e0 1 I 0 0 3557889 4359644 813343.843
> c000001408416938 2 I 0 0 3274470 4075463 812531.368
> c000001408417190 3 I 0 0 3166100 3966471 811934.702
> c0000014084179e8 4 I 0 0 3002538 3804280 813324.752
> c000001408418240 5 I 0 0 2834845 3636310 813031.860
> c000001408418a98 6 I 0 0 2686597 3486253 811202.465
> c0000014084192f0 7 I 0 0 2557634 3358741 812642.450
> c000001408419b48 8 I 0 0 2467357 3268927 813121.692
> c00000140841a3a0 9 I 0 0 2423648 3225405 813342.546
> c00000140841abf8 10 I 0 0 2375710 3175721 811535.485
> c00000140841b450 11 I 0 0 2301315 3101566 811795.799
> c00000140841bca8 12 I 0 0 2267774 3068964 812759.581
> c00000140841c500 13 I 0 0 2195100 2996698 813156.343
> c00000140841cd58 14 I 0 0 2075059 2876297 812764.397
> c00000140841d5b0 15 I 0 0 1774294 2576683 813943.028
> c00000140841de08 16 I 0 0 1757574 2560138 814113.722
> c00000140841e660 17 I 0 0 1740367 2542914 814100.952
> c00000140841eeb8 18 I 0 0 1715677 2518217 814078.609
> c00000140841f710 19 I 0 0 1678186 2480560 813929.036
> c00000140841ff68 20 I 0 0 1628619 2430915 813850.870
> c0000014084207c0 21 I 0 0 1578856 2380834 813509.546
> c000001408421018 22 I 0 0 1557423 2359718 813847.838
> c000001408421870 23 I 0 0 1524951 2326763 813363.393
> c0000014084220c8 24 I 0 0 1489123 2291543 813976.098
> c000001408422920 25 I 0 0 1439096 2241394 813848.700
> c000001408423178 26 I 0 0 1408349 2210684 813881.037
> c0000014084239d0 27 I 0 0 1346739 2149156 813965.500
> c000001408424228 28 I 0 0 1284856 2087248 813931.687
> c000001408424a80 29 I 0 0 1212814 2015075 813809.127
> c0000014084252d8 30 I 0 0 1148024 1950013 813522.639
> c000001408425b30 31 I 0 0 1116655 1918692 813586.337
> c000001408426388 32 I 0 0 1074229 1876200 813506.475
> c000001408426be0 33 I 0 0 1026402 1828330 813454.487
> c000001408427438 34 I 0 0 975538 1777522 813516.306
> c000001408427c90 35 I 0 0 935815 1737636 813347.500
> c0000014084284e8 36 I 0 0 867591 1669514 813447.230
> c000001408428d40 37 I 0 0 818066 1619650 813138.467
> c000001408429598 38 I 0 0 789886 1591734 813375.733
> c000001408429df0 39 I 0 0 757005 1558901 813434.546
> c00000140842a648 40 I 0 0 712937 1514772 813384.725
> c00000140842aea0 41 I 0 0 654107 1455916 813360.628
> c00000140842b6f8 42 I 0 0 639123 1440821 813233.113
> c00000140842bf50 43 I 0 0 585283 1386602 812831.072
> c00000140842c7a8 44 I 0 0 538219 1339512 812784.338
> c00000140842d000 45 I 0 0 513150 1314294 812633.153
> c00000140842d858 46 I 0 0 449294 1250415 812619.052
> c00000140842e0b0 47 I 0 0 413833 1215221 812911.429
> c00000140842e908 48 I 0 0 396667 1198171 813026.016
> c00000140842f160 49 I 0 0 366924 1168836 813432.231
> c00000140842f9b8 50 I 0 0 352521 1154531 813523.150
> c000001408430210 51 I 0 0 334947 1136473 813038.820
> c000001408430a68 52 I 0 0 317619 1119994 813893.778
> c0000014084312c0 53 I 0 0 309327 1111570 813759.278
> c000001408431b18 54 I 0 0 299992 1101864 813378.746
> c000001408432370 55 I 0 0 281463 1083137 813189.738
> c000001408432bc8 56 I 0 0 267427 1069605 813678.658
> c000001408433420 57 I 0 0 257462 1059482 813543.164
> c000001408433c78 58 I 0 0 233999 1036539 814069.121
> c0000014084344d0 59 I 0 0 216667 1019007 813869.390
> c000001408434d28 60 I 0 0 208636 1011150 814033.976
> c000001408435580 61 I 0 0 19
Hello Art,
We are still facing the timely issue of high response and not sure
It is because of LRU flusing as comapred to Chunk writes OR
It is because of the new disk setup we have implemented.
HP team is working on the issue and as always the HP team says that you need to
check the database bottleneck and the IBM AVL support says you need to check
with HP :(
I took up some of the you queries with the HP team and the response is as
Why 10k rmp and not 15k disk was selected?
They say the new disk has some kind of 2 GB SSD which will enhance the
performance
of disk and it is called as Adaptive optimization configuration, BTW they have
not yet enabled this feature as we are having performance issues now.
Raid 1 Vs 5
Yes, 6 DATABASE vgs are on RAID5 but rest all are on RAID 1
I have also asked them to move these 6 to RAID 1 and not have anything
releated to database on RAID 5
They do not understand my question on one single spindle being virtual
redundancy
and if it fails then everything fails
Until they break their heads I want to quicky check few things on informix side
I have even rebooted the servers primary and sds just to check if it helps
For sure chunck writes are better than LRU rights but after reboot I see more
LRU Writes
prod:/home/vikash $ onstat -F
IBM Informix Dynamic Server Version 11.50.FC8W2 -- On-Line (Prim) -- Up
06:37:30 -- 83208912 Kbytes
Fg Writes LRU Writes Chunk Writes
0 1647079 59511
I am not sure if these intensive writing could be or problem, you recommended
that 0.5 and 2 setting against our current 1 and 5, would it reduce the LRU
rights
The onstat -g iof output is too big to post here can you help me understand
what portion
in that output will show me the delay in read/writes or a benchmark value
below which
we are in trouble
sar -a
HP-UX prodtu B.11.31 U ia64 10/16/13
00:00:00 iget/s namei/s dirbk/s
00:05:00 2 803 34
00:10:00 1 770 32
00:15:00 94 772 47
00:20:00 1 782 31
00:25:00 1 776 37
00:30:00 1 745 24
00:35:00 1 782 40
00:40:00 1 755 37
00:45:00 1 865 34
00:50:00 1 765 26
00:55:00 1 781 36
01:00:00 1 753 22
01:05:00 1 771 24
01:10:00 3 821 42
01:15:00 87 753 32
01:20:00 2 785 42
01:25:00 2 813 49
01:30:00 2 836 67
01:35:01 3 870 74
01:40:00 1 794 38
01:45:00 1 756 32
01:50:00 2 817 48
01:55:00 5 1238 50
02:00:00 1 750 20
02:05:00 1 775 24
02:10:00 1 741 22
02:15:00 99 743 21
02:20:00 2 740 26
02:25:00 1 762 24
02:30:00 1 731 20
02:35:00 6 941 51
02:40:00 9 1680 139
02:45:00 4 1006 222
02:50:00 8 1000 144
02:55:00 3 828 58
03:00:00 2 793 38
03:05:00 1 771 22
03:10:00 1 763 24
03:15:00 110 756 21
03:20:00 1 769 28
03:25:00 4 748 26
03:30:00 1 740 20
03:35:00 1 758 23
03:40:00 2 823 41
03:45:00 1 771 22
03:50:00 5 825 47
03:55:00 1 772 24
04:00:00 1 724 20
04:05:00 1 774 24
04:10:00 1 764 25
04:15:00 100 421 25
04:20:00 1 377 31
04:25:00 1 393 34
04:30:00 1 229 13
04:35:00 1 285 29
04:40:00 1 212 18
05:05:02 102 64 99
05:10:00 2 596 47
05:15:00 111 674 58
05:20:00 3 883 45
05:25:00 2 793 40
05:30:00 1 743 37
05:35:00 1 750 43
05:40:00 1 798 44
05:45:00 5 781 48
05:50:00 1 760 39
05:55:00 1 768 36
06:00:00 1 726 37
06:05:00 1 750 38
06:10:00 1 760 40
06:15:00 97 738 37
06:20:00 2 757 40
06:25:00 1 716 36
06:30:00 1 752 37
06:35:00 1 740 37
06:40:00 1 768 39
06:45:00 1 727 42
06:50:00 1 758 41
06:55:00 1 739 36
07:00:00 1 744 36
07:05:00 1 771 43
07:10:00 1 756 40
07:15:00 64 726 39
07:20:00 1 772 41
07:25:00 1 794 41
07:30:00 1 779 38
07:35:00 1 787 38
07:40:00 1 773 40
07:45:00 1 785 35
07:50:00 1 802 39
07:55:00 1 791 37
08:00:00 1 809 37
08:05:01 1 837 43
08:10:00 3 931 56
08:15:00 72 851 37
08:20:00 2 867 42
08:25:00 1 825 40
08:30:00 1 756 39
08:35:00 1 762 45
08:40:00 1 767 41
08:45:00 1 719 35
08:50:00 1 757 40
08:55:00 1 752 38
09:00:00 3 752 40
09:05:00 1 806 46
09:10:00 1 764 45
09:15:00 84 767 37
09:20:00 1 781 43
09:25:00 1 733 38
09:30:00 1 761 38
09:35:00 1 773 40
09:40:00 1 790 43
09:45:00 1 761 36
09:50:01 1 790 42
09:55:00 1 778 39
10:00:00 1 835 45
10:05:00 1 792 40
10:10:00 1 821 46
10:15:00 72 781 43
10:20:00 1 814 41
10:25:00 1 824 41
10:30:00 1 782 37
10:35:00 1 800 40
10:40:00 1 802 43
10:45:00 1 778 36
10:50:00 1 807 42
10:55:00 1 802 40
11:00:00 1 813 39
11:05:00 1 804 42
11:10:00 1 850 49
11:15:00 70 868 42
11:20:00 2 907 45
11:25:00 1 836 40
11:30:00 1 848 41
11:35:00 1 824 41
11:40:00 1 836 44
Average 102 65 99
VIkas:
First, here's some of my onstat -g iof output:
> onstat -g iof
IBM Informix Dynamic Server Version 11.70.FC7 -- On-Line (Prim) -- Up 4
days 07:52:09 -- 6781312 Kbytes
AIO global files:
gfd pathname bytes read page reads bytes write page writes
io/s
3 rootdbs.1 169037824 82538 74778624 36513 915.3
op type count avg. time
seeks 0 N/A
reads 0 N/A
writes 7 0.0011
kaio_reads 16308 0.0012
kaio_writes 21607 0.0010
4 physlog.1 34816 17 1404880896 685977 513.4
op type count avg. time
seeks 0 N/A
reads 0 N/A
writes 0 N/A
kaio_reads 15 0.0016
kaio_writes 11901 0.0019
5 logicallog.1 4621144064 2256418 3662395392 1788279
1129.7
op type count avg. time
seeks 0 N/A
reads 0 N/A
writes 0 N/A
kaio_reads 432663 0.0005
kaio_writes 220925 0.0016
9 datadbs1.1 24896223232 12156387 1337100288 652880
478.4
op type count avg. time
seeks 0 N/A
reads 0 N/A
writes 1893 0.0149
kaio_reads 1071662 0.0008
kaio_writes 331983 0.0061
10 datadbs2.1 18248091648 8910256 53866496 26302
1179.4
op type count avg. time
seeks 0 N/A
reads 0 N/A
writes 0 N/A
kaio_reads 380858 0.0008
kaio_writes 7589 0.0013
This server uses RAW devices for chunks so KAIO is enabled. If you are
using cooked devices or filesystem files for chunks and have not enabled
DIRECT_IO then the data will appear in the seeks, reads, & writes lines
instead, but the interpretation is the same. The "avg. time" column
represents the service time for the device/file or how long it took, on
average, for a read or write (or seek - KAIO eliminates the seek step) to
return. This should be less than 20ms or 0.0200. If it gets higher than
that, then this device is being a bottleneck - however, ignore chunks that
do not have a significant number of IOs since their service times will be
skewed.
Other responses below:
I can continue to answer specific questions for you here, but you have a
complex environment so I don't know how much help I can be without getting
hands on with your system and monitor it over time to locate the
bottlenecks. I can do that if you need me to, but that would be a paid
consulting gig. If your company wants that kind of help, contact me
privately.
Art
Art S. Kagel, Principal Consultant
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions
and do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other organization with which I am associated either explicitly,
implicitly, or by inference. Neither do those opinions reflect those of
other individuals affiliated with any entity with which I am affiliated nor
those of the entities themselves.
On Wed, Oct 16, 2013 at 3:26 AM, VIKAS HIVARKAR
<vikas.hivarkar@gmail.com>wrote:
> Hello Art,
>
> We are still facing the timely issue of high response and not sure
> It is because of LRU flusing as comapred to Chunk writes OR
> It is because of the new disk setup we have implemented.
>
OK, so chunks writes versus lru writes is a thorny issue. Chunks writes
are more efficient for the system as a whole and for the SAN in particular
because they allow the OS's IO drivers and the SAN to coalesce multiple
contiguously placed writes into fewer larger IO operations and to do what
are known as elevator sorts to reorder IO operations to avoid random seeks
and minimize disk head location seeking. However, large chunk writes
during synchronous checkpoints (yes, the engine still does blocking
checkpoints under certain conditions) will cause serious performance
problems in an OLTP environment. So, LRU writes are not necessarily a bad
thing and if you are experiencing checkpoint blocking issues then you want
more LRU writes not fewer.
>
> HP team is working on the issue and as always the HP team says that you
> need to
> check the database bottleneck and the IBM AVL support says you need to
> check with HP :(
>
> I took up some of the you queries with the HP team and the response is as
>
> Why 10k rmp and not 15k disk was selected?
> They say the new disk has some kind of 2 GB SSD which will enhance the
> performance
> of disk and it is called as Adaptive optimization configuration, BTW they
> have
> not yet enabled this feature as we are having performance issues now.
>
OK! I just LOVE SAN geeks, they think that their technology is the end-all
and be-all of performance solutions. So, you have a hybrid SSD/spindle
system that stages high throughput files in the SSD drives for speed and
persists the data on the spindles for safety. Note that these HP SANs will
also magically relocate your data from volume group to volume group and
from RAID10 to RAID5 when it decides it needs to. I have another client
who bought into this line of BS and their IO performance on the new
super-SAN is not meeting their expectations. Just a warning. I've said it
elsewhere, SANs and big disks were the absolute worst thing to happen to
database systems. I can build a faster database server using many
singleton disks and many controllers than any SAN.
>
> Raid 1 Vs 5
>
> Yes, 6 DATABASE vgs are on RAID5 but rest all are on RAID 1
> I have also asked them to move these 6 to RAID 1 and not have anything
> releated to database on RAID 5
>
OK, good. Note again, that if this is the HP SAN system I'm thinking of,
it can reconfigure your RAID1/RAID10 drives as RAID5 if free space gets too
low!
>
> They do not understand my question on one single spindle being virtual
> redundancy and if it fails then everything fails
>
Is it RAID1 or RAID10? RAID1 has redundancy as does RAID10, but RAID10 has
better performance as it incorporates the ability of RAID0 to spread the
IOs over multiple RAID1 pairs.
>
> Until they break their heads I want to quicky check few things on informix
> side
> I have even rebooted the servers primary and sds just to check if it helps
>
> For sure chunck writes are better than LRU rights but after reboot I see
> more
> LRU Writes
>
> prod:/home/vikash $ onstat -F
>
> IBM Informix Dynamic Server Version 11.50.FC8W2 -- On-Line (Prim) -- Up
> 06:37:30 -- 83208912 Kbytes>
> Fg Writes LRU Writes Chunk Writes
> 0 1647079 59511
>
> I am not sure if these intensive writing could be or problem, you
> recommended
> that 0.5 and 2 setting against our current 1 and 5, would it reduce the LRU
> rights
>
Reducing the lru_min/max_dirty settings will actually increase the LRU
writes and decrease the chunk writes. If you want to try moving more of
the IO to chunk writes during checkpoint time then you have to increase
these values instead.
>
> The onstat -g iof output is too big to post here can you help me understand
> what portion
> in that output will show me the delay in read/writes or a benchmark value
> below which
> we are in troubl
Hi Art, The issue is finally resolved by the HP team, they made two changes 1. Hardcode the portspeed to 4gig and 2. Enable Adaptive Optimization I am yet to know which change resolved the issue but since these two changes we are getting a blazing fast performance. and yes I will not leave their back until we move the db disk to RAID 1 from RAID5 Many thanks for your support and as always got to leart new thing from you! Regards Vikas
Great news! Art Art S. Kagel, Principal Consultant Advanced DataTools (www.advancedatatools.com) Blog: http://informix-myview.blogspot.com/ Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Fri, Oct 18, 2013 at 8:09 AM, VIKAS HIVARKAR <vikas.hivarkar@gmail.com>wrote: > Hi Art, > > The issue is finally resolved by the HP team, they made two changes > 1. Hardcode the portspeed to 4gig and > 2. Enable Adaptive Optimization > > I am yet to know which change resolved the issue but since these two > changes > we are getting a blazing fast performance. > > and yes I will not leave their back until we move the db disk to RAID 1 > from > RAID5 > > Many thanks for your support and as always got to leart new thing from you! > > Regards > Vikas > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --001a11c3c9d885f9a104e902d733
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g