Fw: Informix thread management
Posted in 2004
Topics: Storage & Space Management, SQL Development & Query Writing, Versions, Editions & End-of-Life
I suppose I deserve this for switching machines and not correctly updating
my mail account.
cheers
j.
----- Original Message -----
From: "Jack Parker" <jack.parker4@verizon.net>
To: "Andy Kent" <andykent.bristol@virgin.net>; <informix-list@iiug.org>
Sent: Tuesday, January 06, 2004 8:52 PM
Subject: Re: Informix thread management
>
> ----- Original Message -----
> From: "Andy Kent" <andykent.bristol@virgin.net>
> To: <informix-list@iiug.org>
> Sent: Tuesday, January 06, 2004 7:57 AM
> Subject: Re: Informix thread management
>
>
> > Thanks Jack, I've read your article now and it's certainly quite
> > informative.
> >
> > I'm still not clear on exactly what DS memory stores and what it
>
> Think of DS memory as 'working' or 'scratch' memory. A large section of
> memory devoted to things that can take advantage of it - hash joins, sort
> areas, group area, etc.
>
> > allows you to do, and how come parallelisation is still able to happen
> > on the system I am currently working on where IDS 7.3 has defaulted it
> > to the ludicrously low value of 768k so it never gets used by any
> > queries?
>
> They're different - PDQPRIORITY =1 will turn on parallelism - i.e.
parallel
> scans where appropriate - whether they have the memory to fit into to do
> things like useful hash joins is another matter.
>
> Do I take it that DS and parallelism are in fact two separate
> > concepts? I'd be interested if you could clarify this.
>
> Parallelism should be an element to DSS processing.
>
> While you can do DSS without parallelism, it would not make sense. There
> are cases where parallelism would make sense without doing DSS queries or
> processing. Traditional OLTP will not benefit greatly from parallelism -
if
> you are accessing a row with a b-tree index, then there isn't much sense
in
> reading chunks in parallel - you are going to navigate down a tree - one
> thread is sufficient.
>
> cr*p - still didn't differentiate. Parallelism, multiple processes
> performing discrete elements of a larger task. DSS: Processing massive
> quantities of data to derive information.
>
> DSS is useful when you need to 'boil the ocean' - read ALL, or large
> portions of the data. Hence it is also useful for index builds. Here are
> some other cute things you can do with DSS -
>
http://www-106.ibm.com/developerworks/db2/zones/informix/library/techarticle/parker/0502parker.html
>
> >
> > I'm interested in your comment "Because insufficient memory was
> > allocated, the hash join is stalled while it swaps hash table pages
> > from the temp disk." How did you tell this was what had happened?
>
> Kind of an intuitive thing with X-tree. You see it running merrily along
> with the speedometer needles pegged at whatever their rate is - and then
> suddenly they drop off to zero for 10-20-30 (or whatnot) seconds - you
know
> something else is happening, but it's not happening with your query. ergo
> it must be the swap - especially since you can calculate how much memory
> your join will take, and about when (volume and time) it will overflow.
You
> can time these drops, they will take roughly the same amount of time each
> time, until they hit the 200% mark at which point they will double* - I
> wager you will also be able to watch your temp disk spindles light up
during
> this period (i.e. I've never actually bothered to check on the timing). I
> admit, I can get a bit anal with X-tree and will spend a great deal of
time
> entranced by the needles and the clock, feeling the engine underneath -
sort
> of the way you might drive a standard transmission on a mountain road. If
> you know what the engine has to do, you can 'feel' it do each piece. XPS
is
> much better in that it will show you realtime what is going on under the
> covers. Nowadays I don't bother, with a sizing spreadsheet I can predict
> how things will run before they happen and don't have to check unless I'm
> taken by surprise.
>
> * If you need 1GB of memory for a join and you give it 250MB, for the
first
> 250MB things will go along swimmingly. For the next 250MB you will see
> these needle drop offs when the engine swaps to temp - perhaps at 10sec a
> pop every 100MB. At 500MB those needle drops will suddenly start taking
> perhaps 30-60 seconds more frequently.
>
> I would highly recommend taking the internals class and the masters series
> classes if you need to get into this, it won't 'give' you all of the
> knowledge you need, but it will form a solid foundation from which you can
> build the rest.
>
> Mind you, I'm no expert - merely an observer.
>
> cheers
> j.
>
> >
> > Thanks,
> >
> > Andy
> >
> >
> >
> > "Jack Parker" <vze2qjg5@verizon.net> wrote in message
> news:<bt58cu$5s1$1@terabinaries.xmission.com>...
> > > Thanks you Norton for shutting that down.
> > >
> > > cheers
> > > j.
> > > -.-- --- ..- / -. . . -.. / - --- / --. . - / .- / .-.. .. ..-. .
.-.-.-
> /
> > > ... --- / -.. --- / .. .-.-.-
> > > ----- Original Message -----
> > > From: "Jack Parker" <vze2qjg5@verizon.net>
> > > To: "Andy Kent" <andykent.bristol@virgin.net>;
<informix-list@iiug.org>
> > > Sent: Thursday, January 01, 2004 8:44 AM
> > > Subject: Re: Informix thread management
> > >
> > >
> > > >
> > > > Decision Support has evolved from the data warehouse community. It
> > > > generally implies that you intend to read entire tables, join them
and
> > > > whatever.
> > > >
> > > > A query which meets the criteria to be processed in parallel means a
> query
> > > > where you intend to pull more than one row from whatever tables you
> happen
> > > > to be reading - and that the multiple rows reside on different
> dbspaces.
> > > I
> > > > suppose that's what we call horizontal parallelism, vertical
> parallelism
> > > > would be where elements of the query tree can be run in parallel -
> maybe
> > > > I've got that backwards (who cares).
> > > >
> > > > Imagine that you have something like:
> > > >
> > > > select count(*) (or whatnot)
> > > > from tab1, tab2
> > > > where tab1.key=tab2.key
> > > > and ......> > > >
> > > > where tab1 is fragmented across 4 dbspaces and tab2 across another
4.
> > > >
> > > > If you are running in parallel then the engine can read all 4 of the
> > > chunks
> > > > in parallel.
> > > >
> > > > -- h*ll - I've already written this article. See:
> > > >
> > > >
> > >
>
http://www7b.boulder.ibm.com/dmdd/zones/informix/library/techarticle/parker/
> > > > part-1.pdf
> > > >
> > > > cheers
> > > > j.
> > > >
> > > > ----- Original Message -----
> > > > From: "Andy Kent" <andykent.bristol@virgin.net>
> > > > To: <informix-list@iiug.org>
> > > > Sent: Thursday, December 18, 2003 5:01 AM
> > > > Subject: Re: Informix thread management
> > > >
> > > >
> > > > > Does "DSS query" mean the same thing as "Query which meets the
> > > > > criteria to be processed in parallel" - or not? Is this what you@
> > Kind of an intuitive thing with X-tree. You see it running merrily along > > with the speedometer needles pegged at whatever their rate is - and then > > suddenly they drop off to zero for 10-20-30 (or whatnot) seconds - you > know > > something else is happening, but it's not happening with your query. Is there any *proper* documentation for xtree anywhere? It's pretty hit-and-miss trying to work out which table it's scanning. The query I'm trying to analyse *looks* as though it's spending all its time repeatedly filtering on the results of a subquery. It only seems to have allocated one thread to this, so I am getting zero benefit from throwing masses of PDQ and DS resources at it. The weird thing is the subquery is uncorrelated, so I'd expect the engine just to run it the once, do an autoindex on the resultant TT and then do indexes lookups to the TT. Why would it be spending all its time on the filter to the subquery? But of course in the absence of any decent xtree documentation, there's an element of guesswork in the above. We're on 7.31.UC5. Thanks Andy
Related threads
- IDS not writing to online.log
- RE: How to find usage for sbspace
- Writing an iterator function in SPL
- log files location