CLEANERS
Posted in 2000
A DBA on IDS 7.30 (Solaris 2.6) raised CLEANERS from 64 to 128 (LRU also at 128) and then saw startup warnings that the physical log size was reported as 0 / too small and that the logical log layout could lock the server, plus a CDR_NIFRETRY message and checkpoints stretching to about a minute. Replies suggested backing CLEANERS out to 64 to confirm cause, noting onstat -u showed only 77 cleaners actually writing (so ~77 would suffice), and argued long checkpoints usually stem from disk contention or skewed write distribution across spindles, LRU_MIN/MAX tuning and poor application/schema design rather than cleaner count. No confirmed resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Server Administration, Logging & Checkpoints, Platform-Specific Issues, Versions, Editions & End-of-Life
Hi,
I only increased 'CLEANERS' from 64 to 128 and restarted the informix
server. I got the following warning message in log file:
00:42:53 WARNING! Physical Log size 0 is too small. Physical Log overflows may occur during peak activity.
Recommended minimum Physical Log size is 20 times maximum
concurrent user threads.
00:42:54 WARNING! Logical log layout may cause Dynamic Server to getinto
a locked state. Recommended smallest logical log size
is 2 times maximum concurrent user threads.
Is it serious message? What should I do? Change CLEANERS back to 64?
Before I changed CLEANERS, I never had such message.
Thanks in advance.
OS is Solaris 2.6. IDS 7.30.uc6. Other onconfig related parameters are:
LRU 128
PHYSFILE 100000
PHYSBUFF 64
LOGFILES 38
LOGSIZE 500
PHYSBUFF 64
LOGBUFF 64LOGSMAX 60
...
Also after I changed CLEANERS, there is such a line:
00:42:59 Onconfig parameter CDR_NIFRETRY modified from 0 to 301.
in the log file. What does it mean?
Sent via Deja.com http://www.deja.com/
Before you buy.
Thanks for help from everyone.
Thanks, Frank. I heard the bug by setting LRU to 128 in Art's article.
But I didn't know what the bug looked like. Now I know it. It really
scared me.
Why I set LRU to 128? Because I found that the BWR in the server was 11-
13%. After I increased LRU from 64 to 128(the only thing I changed at
that time), BWR was down to 2-4%. And I heard that it's better to set
CLEANERS greater than or equal to LRU. So this time I increased
CLEANERS to 128.
But I found that checkpoint duration is increased after CLEANERS
increased to 128. It looks like that checkpoint take 1 minute every 6
hours. I never had the problem before. It's really scary for
performance is the most important thing in the server.(The reponse time
should be 2-3 seconds.)
Some other related parameters:
LRU_MAX_DIRTY 2
LRU_MIN_DIRTY 1
CKPTINTVL 300
We have about 30 disks, 37 dbspaces, 88 chunks...
In article <90gti5$gso$1@nnrp1.deja.com>,
hanna_shaw@my-deja.com wrote:
> Hi,
>
> I only increased 'CLEANERS' from 64 to 128 and restarted the informix
> server. I got the following warning message in log file:
>
> 00:42:53 WARNING! Physical Log size 0 is too small.> Physical Log overflows may occur during peak activity.
> Recommended minimum Physical Log size is 20 times maximum
> concurrent user threads.
>
> 00:42:54 WARNING! Logical log layout may cause Dynamic Server to get> into
> a locked state. Recommended smallest logical log size
> is 2 times maximum concurrent user threads.
>
> Is it serious message? What should I do? Change CLEANERS back to 64?
> Before I changed CLEANERS, I never had such message.
>
> Thanks in advance.
>
> OS is Solaris 2.6. IDS 7.30.uc6. Other onconfig related parameters
are:
>
> LRU 128
> PHYSFILE 100000
> PHYSBUFF 64
> LOGFILES 38
> LOGSIZE 500
> PHYSBUFF 64
> LOGBUFF 64> LOGSMAX 60
> ...
>
> Also after I changed CLEANERS, there is such a line:
>
> 00:42:59 Onconfig parameter CDR_NIFRETRY modified from 0 to 301.>
> in the log file. What does it mean?
>
> Sent via Deja.com http://www.deja.com/
> Before you buy.
>
Sent via Deja.com http://www.deja.com/
Before you buy.
Hard to argue with a production system...
As much as I hate to say it, if you are sure something in your
processing has not changed, and this is not some transaction in a
critical section, I would be changing CLEANERS back to 64 and seeing if
that solves the problem.
I would be interested in hearing what happens,
Will
In article <90m20c$m3k$1@nnrp1.deja.com>,
hanna_shaw@my-deja.com wrote:
> Thanks for help from everyone.
>
> Thanks, Frank. I heard the bug by setting LRU to 128 in Art's article.
> But I didn't know what the bug looked like. Now I know it. It really
> scared me.
>
> Why I set LRU to 128? Because I found that the BWR in the server was
11-
> 13%. After I increased LRU from 64 to 128(the only thing I changed at
> that time), BWR was down to 2-4%. And I heard that it's better to set
> CLEANERS greater than or equal to LRU. So this time I increased
> CLEANERS to 128.>
> But I found that checkpoint duration is increased after CLEANERS
> increased to 128. It looks like that checkpoint take 1 minute every 6
> hours. I never had the problem before. It's really scary for
> performance is the most important thing in the server.(The reponse
time
> should be 2-3 seconds.)
>
> Some other related parameters:
> LRU_MAX_DIRTY 2
> LRU_MIN_DIRTY 1
> CKPTINTVL 300>
> We have about 30 disks, 37 dbspaces, 88 chunks...
>
> In article <90gti5$gso$1@nnrp1.deja.com>,
> hanna_shaw@my-deja.com wrote:
> > Hi,
> >
> > I only increased 'CLEANERS' from 64 to 128 and restarted the
informix
> > server. I got the following warning message in log file:
> >
> > 00:42:53 WARNING! Physical Log size 0 is too small.> > Physical Log overflows may occur during peak activity.
> > Recommended minimum Physical Log size is 20 times maximum
> > concurrent user threads.
> >
> > 00:42:54 WARNING! Logical log layout may cause Dynamic Server toget
> > into
> > a locked state. Recommended smallest logical log size
> > is 2 times maximum concurrent user threads.
> >
> > Is it serious message? What should I do? Change CLEANERS back to 64?
> > Before I changed CLEANERS, I never had such message.
> >
> > Thanks in advance.
> >
> > OS is Solaris 2.6. IDS 7.30.uc6. Other onconfig related parameters
> are:
> >
> > LRU 128
> > PHYSFILE 100000
> > PHYSBUFF 64
> > LOGFILES 38
> > LOGSIZE 500
> > PHYSBUFF 64
> > LOGBUFF 64> > LOGSMAX 60
> > ...
> >
> > Also after I changed CLEANERS, there is such a line:
> >
> > 00:42:59 Onconfig parameter CDR_NIFRETRY modified from 0 to 301.> >
> > in the log file. What does it mean?
> >
> > Sent via Deja.com http://www.deja.com/
> > Before you buy.
> >
>
> Sent via Deja.com http://www.deja.com/
> Before you buy.
>
Sent via Deja.com http://www.deja.com/
Before you buy.
I typed 'onstat -u' and found that 'nwrites' of 77 cleaner thread were
not zero, the other 51 of the 128 were 0. Does it mean I'd better to
set 'CLEANERS' to 77?
What's the cost for setting LRU and CLEANERS very high?
In article <90m79t$r3d$1@nnrp1.deja.com>,
William Rice <ricew@operamail.com> wrote:
> Hard to argue with a production system...
>
> As much as I hate to say it, if you are sure something in your
> processing has not changed, and this is not some transaction in a
> critical section, I would be changing CLEANERS back to 64 and seeing
if
> that solves the problem.
>
> I would be interested in hearing what happens,
> Will
>
> In article <90m20c$m3k$1@nnrp1.deja.com>,
> hanna_shaw@my-deja.com wrote:
> > Thanks for help from everyone.
> >
> > Thanks, Frank. I heard the bug by setting LRU to 128 in Art's
article.
> > But I didn't know what the bug looked like. Now I know it. It really
> > scared me.
> >
> > Why I set LRU to 128? Because I found that the BWR in the server was
> 11-
> > 13%. After I increased LRU from 64 to 128(the only thing I changed
at
> > that time), BWR was down to 2-4%. And I heard that it's better to
set
> > CLEANERS greater than or equal to LRU. So this time I increased
> > CLEANERS to 128.> >
> > But I found that checkpoint duration is increased after CLEANERS
> > increased to 128. It looks like that checkpoint take 1 minute every
6
> > hours. I never had the problem before. It's really scary for
> > performance is the most important thing in the server.(The reponse
> time
> > should be 2-3 seconds.)
> >
> > Some other related parameters:
> > LRU_MAX_DIRTY 2
> > LRU_MIN_DIRTY 1
> > CKPTINTVL 300> >
> > We have about 30 disks, 37 dbspaces, 88 chunks...
> >
> > In article <90gti5$gso$1@nnrp1.deja.com>,
> > hanna_shaw@my-deja.com wrote:
> > > Hi,
> > >
> > > I only increased 'CLEANERS' from 64 to 128 and restarted the
> informix
> > > server. I got the following warning message in log file:
> > >
> > > 00:42:53 WARNING! Physical Log size 0 is too small.> > > Physical Log overflows may occur during peak activity.
> > > Recommended minimum Physical Log size is 20 times
maximum
> > > concurrent user threads.
> > >
> > > 00:42:54 WARNING! Logical log layout may cause Dynamic Server to> get
> > > into
> > > a locked state. Recommended smallest logical log size
> > > is 2 times maximum concurrent user threads.
> > >
> > > Is it serious message? What should I do? Change CLEANERS back to
64?
> > > Before I changed CLEANERS, I never had such message.
> > >
> > > Thanks in advance.
> > >
> > > OS is Solaris 2.6. IDS 7.30.uc6. Other onconfig related parameters
> > are:
> > >
> > > LRU 128
> > > PHYSFILE 100000
> > > PHYSBUFF 64
> > > LOGFILES 38
> > > LOGSIZE 500
> > > PHYSBUFF 64
> > > LOGBUFF 64> > > LOGSMAX 60
> > > ...
> > >
> > > Also after I changed CLEANERS, there is such a line:
> > >
> > > 00:42:59 Onconfig parameter CDR_NIFRETRY modified from 0 to 301.> > >
> > > in the log file. What does it mean?
> > >
> > > Sent via Deja.com http://www.deja.com/
> > > Before you buy.
> > >
> >
> > Sent via Deja.com http://www.deja.com/
> > Before you buy.
> >
>
> Sent via Deja.com http://www.deja.com/
> Before you buy.
>
Sent via Deja.com http://www.deja.com/
Before you buy.
Two days ago I would have been sure that your problem did not have to do
with the cleaners.
I do not believe that there are any issues with setting LRU's to 128, I
have never seen problems with setting CLEANERS=LRU's.
That all said, I have never worked on your system :)
If I change cleaners, and then start having problems, then I suspect
cleaners.
If it is causing serious production issues, I would change it back and
see if the problem continued. If problems continue and you know that
cleaners are not your problem. If it stops, then you can wait a week,
then try to change CLEANERS again. If you encounter the same problem
again, then I would start researching why CLEANERS are causing the
issue.
If it takes you longer to flush the same amount of BUFFERS with more
CLEANERS, then the number of CLEANERS might be causing disk thrashing
during the checkpoint, or for some reason a transaction might be in a
critical section. Using onstat -R you can look at the number of dirty
buffers. Collect some data for lenth of checkpoints flushing certain
numbers of buffers, and work fro there.
Hope this help's some,
Will
In article <90m9aj$t1h$1@nnrp1.deja.com>,
hanna_shaw@my-deja.com wrote:
> I typed 'onstat -u' and found that 'nwrites' of 77 cleaner thread were
> not zero, the other 51 of the 128 were 0. Does it mean I'd better to
> set 'CLEANERS' to 77?
>
> What's the cost for setting LRU and CLEANERS very high?
>
> In article <90m79t$r3d$1@nnrp1.deja.com>,
> William Rice <ricew@operamail.com> wrote:
> > Hard to argue with a production system...
> >
> > As much as I hate to say it, if you are sure something in your
> > processing has not changed, and this is not some transaction in a
> > critical section, I would be changing CLEANERS back to 64 and seeing
> if
> > that solves the problem.
> >
> > I would be interested in hearing what happens,
> > Will
> >
> > In article <90m20c$m3k$1@nnrp1.deja.com>,
> > hanna_shaw@my-deja.com wrote:
> > > Thanks for help from everyone.
> > >
> > > Thanks, Frank. I heard the bug by setting LRU to 128 in Art's
> article.
> > > But I didn't know what the bug looked like. Now I know it. It
really
> > > scared me.
> > >
> > > Why I set LRU to 128? Because I found that the BWR in the server
was
> > 11-
> > > 13%. After I increased LRU from 64 to 128(the only thing I changed
> at
> > > that time), BWR was down to 2-4%. And I heard that it's better to
> set
> > > CLEANERS greater than or equal to LRU. So this time I increased
> > > CLEANERS to 128.> > >
> > > But I found that checkpoint duration is increased after CLEANERS
> > > increased to 128. It looks like that checkpoint take 1 minute
every
> 6
> > > hours. I never had the problem before. It's really scary for
> > > performance is the most important thing in the server.(The reponse
> > time
> > > should be 2-3 seconds.)
> > >
> > > Some other related parameters:
> > > LRU_MAX_DIRTY 2
> > > LRU_MIN_DIRTY 1
> > > CKPTINTVL 300> > >
> > > We have about 30 disks, 37 dbspaces, 88 chunks...
> > >
> > > In article <90gti5$gso$1@nnrp1.deja.com>,
> > > hanna_shaw@my-deja.com wrote:
> > > > Hi,
> > > >
> > > > I only increased 'CLEANERS' from 64 to 128 and restarted the
> > informix
> > > > server. I got the following warning message in log file:
> > > >
> > > > 00:42:53 WARNING! Physical Log size 0 is too small.> > > > Physical Log overflows may occur during peak activity.
> > > > Recommended minimum Physical Log size is 20 times
> maximum
> > > > concurrent user threads.
> > > >
> > > > 00:42:54 WARNING! Logical log layout may cause Dynamic Serverto
> > get
> > > > into
> > > > a locked state. Recommended smallest logical log size
> > > > is 2 times maximum concurrent user threads.
> > > >
> > > > Is it serious message? What should I do? Change CLEANERS back to
> 64?
> > > > Before I changed CLEANERS, I never had such message.
> > > >
> > > > Thanks in advance.
> > > >
> > > > OS is Solaris 2.6. IDS 7.30.uc6. Other onconfig related
parameters
> > > are:
> > > >
> > > > LRU 128
> > > > PHYSFILE 100000
> > > > PHYSBUFF 64
> > > > LOGFILES 38
> > > > LOGSIZE 500
> > > > PHYSBUFF 64
> > > > LOGBUFF 64> > > > LOGSMAX 60
> > > > ...
> > > >
> > > > Also after I changed CLEANERS, there is such a line:
> > > >
> > > > 00:42:59 Onconfig parameter CDR_NIFRETRY modified from 0 to
301.> > > >
> > > > in the log file. What does it mean?
> > > >
> > > > Sent via Deja.com http://www.deja.com/
> > > > Before you buy.
> > > >
> > >
> > > Sent via Deja.com http://www.deja.com/
> > > Before you buy.
> > >
> >
> > Sent via Deja.com http://www.deja.com/
> > Before you buy.
> >
>
> Sent via Deja.com http://www.deja.com/
> Before you buy.
>
Sent via Deja.com http://www.deja.com/
Before you buy.
hanna_shaw@my-deja.com wrote in message <90m20c$m3k$1@nnrp1.deja.com>... > >But I found that checkpoint duration is increased after CLEANERS >increased to 128. It looks like that checkpoint take 1 minute every 6 >hours. I never had the problem before. It's really scary for >performance is the most important thing in the server.(The reponse time >should be 2-3 seconds.) You'll probably find that the outrageous checkpoint time is due to massive contention for disk activity. Are you aware of the little Informix utility pwrite that is handed out in tuning courses? Pay attention to the consequences of the measures from that... It's immutable.
hanna_shaw@my-deja.com wrote in message <90m9aj$t1h$1@nnrp1.deja.com>...
>I typed 'onstat -u' and found that 'nwrites' of 77 cleaner thread were
>not zero, the other 51 of the 128 were 0. Does it mean I'd better to
>set 'CLEANERS' to 77?
That's basically the situation. As an aside, I'm stunned that you say you
get checkpoints upto a minute sometimes! There is definitely a problem
there, and maybe you need to track down the process causing that, seal it in
a steel box and throw it out to sea.
Y'know, checkpoints are not really the enemy. F'instance, we've found that
checkpoints of upto a few seconds are definitely not noticed by users in our
applications. When the checkpoints creep up to 8 seconds or more, they start
to get antsy, and if they reach 15 seconds or more, they are definitely
unhappy. Setting massive numbers of LRU's, CLEANERS and outrageously low
water marks to force FG writes has never been considered here as a method,
but it's an interesting approach and I might just try it for the hell of it
one day. Just have to find a friendly user site or the right opportunity to
test-n-snoop.
One of the worst reasons for long checkpoints that I've found in the past is
seriously skewed distribution of writes across the spindles. Once I've
analysed the sysmaster statistics and come up with a plan to scatter the
tables around by their read and write loads (and it's basically a 4
dimensional problem) I've gotten much better checkpoint times due to an
increased and appropriate number of active cleaner threads kicking in during
checkpoints.
When driven the right way, disks can get massive throughput, and it's the
fact that FG writes don't accommodate the strengths of disks is the reason
that I'd always be attempting to reduce FG writes and improve chunk and
write distributions for best checkpoint performance.
"Quality" of application code is also very important, because inappropriate
indexes, queries, etc can cause too much work during queries and that leads
to increased bemand for buffers due to excess sequential reads.
And if you have a DSS situation, look seriously at star schema's etc because
your traditional normalised database that we all know and love for OLTP is
almost but not quite the worst possible shape for DSS and warehousing style
of work. Mixing large DSS in with OLTP is fast becoming considered to be a
mortal sin.
Sometimes I think we need to look at the big picture and put some effort
into reorganising the applications and/or the way they work. Nobody can
claim perfect design throughout a large application. Pressures to deliver
ensure that there is some crap design in every system. We can spend 100% of
our tuning time messing with the engine, but then we'll miss the time and
effort that should be put into tuning the application.
In article <3a2edbf6@news.iprimus.com.au>,
"Andrew Hamm" <ahamm@sanderson.net.au> wrote:
> hanna_shaw@my-deja.com wrote in message
<90m9aj$t1h$1@nnrp1.deja.com>...
> >I typed 'onstat -u' and found that 'nwrites' of 77 cleaner thread
were
> >not zero, the other 51 of the 128 were 0. Does it mean I'd better to
> >set 'CLEANERS' to 77?
>
> That's basically the situation. As an aside, I'm stunned that you say
you
> get checkpoints upto a minute sometimes! There is definitely a problem
> there, and maybe you need to track down the process causing that, seal
it in
> a steel box and throw it out to sea.
>
> Y'know, checkpoints are not really the enemy. F'instance, we've found
that
> checkpoints of upto a few seconds are definitely not noticed by users
in our
Long checkpoints are most definitely an enemy to be avoided at
considerable cost on an OLTP system, for exactly the reason you stated,
the users get antsy if you have longer full checkpoints(over 4 seconds
has been my experience on when users start to notice).
> applications. When the checkpoints creep up to 8 seconds or more, they
start
> to get antsy, and if they reach 15 seconds or more, they are
definitely
> unhappy. Setting massive numbers of LRU's, CLEANERS and outrageously
low
> water marks to force FG writes has never been considered here as a
method,
To force LRU writes. I do not think anybody likes FG writes, but LRU
writes aren't really a bad thing. I mean you might not get optimal head
movement, but you do get to spread your I/O over a longer period of
time, and the users are not waiting for that I/O to finish.
> but it's an interesting approach and I might just try it for the hell
of it
> one day. Just have to find a friendly user site or the right
opportunity to
> test-n-snoop.
>
> One of the worst reasons for long checkpoints that I've found in the
past is
> seriously skewed distribution of writes across the spindles. Once I've
> analysed the sysmaster statistics and come up with a plan to scatter
the
> tables around by their read and write loads (and it's basically a 4
> dimensional problem) I've gotten much better checkpoint times due to
an
> increased and appropriate number of active cleaner threads kicking in
during
> checkpoints.
Maybe that has been your experience, but if the disks are not fast
enough to handle the checkpoint, the only way to get reasonable
checkpoints is to lower LRU_MIN/LRU_MAX, and hope your CLEANERS can keep
up.
>
> When driven the right way, disks can get massive throughput, and it's
the
> fact that FG writes don't accommodate the strengths of disks is the
reason
> that I'd always be attempting to reduce FG writes and improve chunk
and
> write distributions for best checkpoint performance.
>
Once again I imagine you mean LRU writes. I think spreading the writes
over a period of time, in which the user is not put on hold could be
considered accommodating the strengths of the disk. I mean if you have
a really high read cache the disks are not getting used much between
checkpoints. So you have the disk sit idle until the checkpoint then
you make it work hard. Why not spread that work over time, and not be
limited by the amount of throughput the disks can do have in 3 seconds?
> "Quality" of application code is also very important, because
inappropriate
> indexes, queries, etc can cause too much work during queries and that
leads
> to increased bemand for buffers due to excess sequential reads.
Quality of code is a major determinant.
>
> And if you have a DSS situation, look seriously at star schema's etc
because
> your traditional normalised database that we all know and love for
OLTP is
> almost but not quite the worst possible shape for DSS and warehousing
style
> of work. Mixing large DSS in with OLTP is fast becoming considered to
be a
> mortal sin.
Didn't see the relevance of this.
>
> Sometimes I think we need to look at the big picture and put some
effort
> into reorganising the applications and/or the way they work. Nobody
can
> claim perfect design throughout a large application. Pressures to
deliver
> ensure that there is some crap design in every system. We can spend
100% of
> our tuning time messing with the engine, but then we'll miss the time
and
> effort that should be put into tuning the application.
>
>
Sometimes we do not have the influence to force the applications to be
tuned, and somtimes is appropriate to dirty lots of buffers.
Will
Sent via Deja.com http://www.deja.com/
Before you buy.
William Rice wrote in message <90o4p8$bj3$1@nnrp1.deja.com>...
> Long checkpoints are most definitely an enemy to be avoided at
> considerable cost on an OLTP system, for exactly the reason you stated,
> the users get antsy if you have longer full checkpoints(over 4 seconds
> has been my experience on when users start to notice).
>
We quote similarish numbers - it's all a perceptual thing from users.
General agreement from all on bad checkpoint lengths it seems. Yup - Long
checkpoints are the enemy. Short checkpoints are cool.
> [you mean]To force LRU writes.
Oops yes - I noticed that slysdexic mention of FG instead of LRU meself
(twice in the same document dammit) after sending it. Total agreement that
FG writes are really hard to justify. Our experience with LRU writes is that
too much of that starts to impact select performance and gives a feeling of
molasses to users. It's a fine balancing act, and would be dependent on the
style of the application too.
> I do not think anybody likes FG writes, but LRU
> writes aren't really a bad thing. I mean you might not get optimal
> head movement, but you do get to spread your I/O over a longer
> period of time, and the users are not waiting for that I/O to finish.
> I think spreading the writes over a period of time,
> in which the user is not put on hold could be
>
Emm (yes). My only qualification for our systems is as above. I'm sure the
issue of spreading the load evenly across all spindles would also be of
great assistance to LRU writes as well. [More on this is in another reply to
people asking about the pwrite utility] That's a theoretical "sure" because
I've always chased minimal LRU and short CHKPT, but it seems bleedingly
obvious. One important factor may be to take care not to overload the
spindles with a few cleaners AND several query threads, since the sum of
them impacts the total throughput.
This may point to a benefit in keeping writes on one set of spindles, and
reads on another set, and makes me feel less uncomfortable about the
weaknesses in my load-sharing algorithm. I still am not satisified that I'm
simultaneously spreading read AND write equally, but a skew across that
dimension may well be useful to keep control on the actual busy threads on
each disk. Ahhh - another opportunity to put my feet up on the desk, stare
into space and claim to be working! Tables busy with reads AND writes would
interfere with that theory.
> Maybe that has been your experience, but if the disks are not fast
> enough to handle the checkpoint, the only way to get reasonable
> checkpoints is to lower LRU_MIN/LRU_MAX, and hope your
> CLEANERS can keep up.>
Fair call. Our technical management is now starting to push thru contract
that customers throw more transistors, cycles, busses and spindles at the
problem should they have a problem due to their workload. I hope the
customers sign on the dotted.
>Once again I imagine you mean LRU writes.
Dammit yes. Slysdexia strikes again ;-) No, it's plain old sleep deprivation
I think. Senility?
> Why not spread that work over time, and not be limited by
> the amount of throughput the disks can do have in 3 seconds?
>
Well, if the I/O system CAN get the checkpoint complete in 3 seconds, then
I'd be quite happy. That's a decent longest checkpoint time, and daily
average in that case would be 1-2 seconds. Nice. All without LRU writes. For
customers with grunty enough boxes. Can we blame price-shaving salesreps (at
hardware companies of course, not our own) for our performance problems on
some boxes? Yeah - I like the sound of that.
>> And if you have a DSS situation, look seriously at star schema's
>> etc because your traditional normalised database that we all
>> know and love for OLTP is almost but not quite the worst possible
>> shape for DSS and warehousing style of work. Mixing large
>> DSS in with OLTP is fast becoming considered to be a
>> mortal sin.
>
> Didn't see the relevance of this.
>
The relevance is, whether we recognise to label it DSS work or not, all OLTP
systems have big reports and now users wielding Crystal Reports, Business
Objects, Cognos etc which hammer the transaction processing. Whether or not
you switch on PDQ etc for their benefit, the rapidly improving science of
warehousing will be another approach to removing contentious activity by
offloading it to warehouse machinery. The trick is to start recognising DSS
style work and agitate for its removal. All dependent on whether you have
that much control of your application of course. I think one or two members
of this news thread have implied that their world is principally DSS? So my
OLTP-centric view is not going to be too relevant to them.
> Sometimes we do not have the influence to force the applications
> to be tuned, and somtimes is appropriate to dirty lots of buffers.
>
That is a sticking point, and it would be even worse if you can't influence
the supplier of the software. My sympathy to people trying to achieve
performance under such constraints. To tell the truth, it's even difficult
to tick all the boxes when you "own" the system, because there are always
more important things to do.