LRU *less* aggressive??
Posted in 2013
Hi ,
ifx 11.50FC9X6 , AIX 6.1
*My questions* :
1) The engine auto tunning, is smarter to make LRU less aggressive?
(making dirty LRU higher).
2) Is possible make the auto tunning "more aggressive" ?? (please, read
bellow to understand the context of my question)
This last week we have problem with the battery of one controller of our
storage , naturally we lost the write cache.
All this was solved by IBM support where become to replace the battery this
weekend
The problem occurred for 6 days and during this days , was reflected
directly on the performance of our database flush.
When the write cache of the storage works fine , our checkpoints average is
15 seconds (non-blocking) , when we lost the write cache this flush
downgrade to 5-15 minutes.... blocking sometimes.
This make our RTO (120 seconds) not be reachable and for each checkpoint an
alert of performance advise + LRU auto tunning was showed , something like
:
243678 04/28/13 00:05:27 Performance Advisory: The time to flush the
bufferpool 205.1 is longer than RTO_SERVER_RESTART 120.
243679 04/28/13 00:05:27 Results: The server cannot meet the RTO policy.
243680 04/28/13 00:05:27 Action: Automatically adjusting LRU flushing to
be more aggressive.
243969 04/28/13 06:12:08 Performance Advisory: The time to flush the
bufferpool 221.3 is longer than RTO_SERVER_RESTART 120.
243970 04/28/13 06:12:08 Results: The server cannot meet the RTO policy.
243971 04/28/13 06:12:08 Action: Automatically adjusting LRU flushing to
be more aggressive.
Now the battery already OK and our checkpoints back to 15 seconds average.
The engine will be smart to notice that? and start come back the LRU to old
configuration? And if the case, trying avoid flush data by LRU ?
Although we suffer 6 days with this problematic checkpoints, I do not
change any configuration "on-fly" , keep the configurations to see if the
engine will try solve by self the problem. (RTO = 120, AUTO_LRU = on ,
AUTO_CKPT = on , LRU_min_dirty=40, max_dirty=60 ).
After 6 days with constant problem, for each checkpoint what trespass the
RTO mark the message was printed on online.log and the LRU is downgrade.
What I mean is, the engine auto tunning not able to solve the problem in 6
days running 24x7, check bellow.
Our configuration have 512 LRUs and buffers of 110 GB 4k + 40 GB of 8k.
During this 6 days , the engine downgrade the LRU dirty to 6-10% and
neither this way able to do a unique LRU write .
Now the battery already OK , this is the output of onstat -F :
| Fg Writes LRU Writes Chunk Writes
| 2 0 572093179
(the FG write occurred when the battery "died" )
So... *My questions* :
1) The engine auto tunning, is smart to make LRU less aggressive? (to
back the original configuration)
2) Is possible make the auto tunning "more aggressive" ?? (lowering the LRU
dirty faster ...)
The reason of my question 1 is because, IF the engine able to start do a
LRU flushing (which not occurred in this case) and avoid the performance
problem during the checkpoint, when we solve the battery , they will keep
the configuration? since I not able to back my LRU_dirty to 40-60%
manually....without bounce the engine.
The reason of my question 2 is because the auto tuning not reach they
objective in 6 days... this is too long... should be done in hours....
Regards
Cesar
--001a11c1b942252e0504db7e5bc3