Paralleling UDR which handles BLOB
Posted in 2000
Topics: Storage & Space Management
Hi; We are developing a "Face detection over video" blade. Our Compare() function should do a massive IO and CPU work. For this reason we wish to make it work parallel. (Obviously) The problem is that BLOB (smart blob which holds the UDT) can't be fragmented over few SBSpaces (as we understand from the Ans. OnLine (3878.pdf) "1-248 Informix Guide to SQL: Syntax") How can we reach our goal? Any idea? TIA, Noam & Asaf
Noam Atzmon wrote: > Hi; > > We are developing a "Face detection over video" blade. > Our Compare() function should do a massive IO and CPU work. > For this reason we wish to make it work parallel. (Obviously) > > The problem is that BLOB (smart blob which holds the UDT) > can't be fragmented over few SBSpaces (as we understand from > the Ans. OnLine (3878.pdf) "1-248 Informix Guide to SQL: Syntax") > > How can we reach our goal? Any idea? > > TIA, Noam & Asaf It really depend on how you are going about your problem. One way will be to break down your initial blobs into smaller components, ie a screen into2, 3, 4, 6, 8, 9 etc components and do the same with the input, and compare the fragments. -- Compliments of QueriX -------------------------------------------------------------------------------------------------- QueriX 4GL Compilers are Informix 4GL Compatible, and Connection to other RDBMS such as Oracle. Hydra 4GL Compiler (Compatible with I4GL) Compile once, run everywhere Phoenix Windows GUI. (Front End to 4GL) Chimera Java GUI The only GUI you will ever need... (Front End to 4GL) Arachne Web Technology (Front End to 4GL on the Web) For more details visit: http://www.querix.com/ ---------------------------------------------------------------------------------------------------
Another issue might be that 9.1 does not support parallel UDRs for smart blobs (your doc reference suggests that you might be on 9.1). 9.2 supports round-robin fragmentation for smart blobs. Also, the mi_lo_* DataBlade API functions are parallelizable in 9.2. That much said, it supports fragmenting a table that contains a smart blob column (i.e., not a single instance of a smart blob, but I may be misunderstanding this part of your post....). This tech note on the IDN DataBlade (*) corner discusses parallel UDRs: http://www.informix.com/idn-secure/DataBlade/Library/Parallel_UDR.htm regards, -jean (*) To access the DataBlade Corner, anybody can register for IDN at http://www.informix.com/idn Noam Atzmon wrote: > Hi; > > We are developing a "Face detection over video" blade. > Our Compare() function should do a massive IO and CPU work. > For this reason we wish to make it work parallel. (Obviously) > > The problem is that BLOB (smart blob which holds the UDT) > can't be fragmented over few SBSpaces (as we understand from > the Ans. OnLine (3878.pdf) "1-248 Informix Guide to SQL: Syntax") > > How can we reach our goal? Any idea? > > TIA, Noam & Asaf
There were more questions on specifically how to fragment a smart
blob column, so I did a quickie test.
First of all, the IDN DataBlade Corner (*) includes a "Smart Blob Info"
DataBlade module that verifies the test -- it includes a SblobSbspace()
routine that outputs the location of a smart blob. (*Anybody can register
for IDN at http://www.informix.com/idn. Once in the DataBlade
Corner, select the "Complete List of Downloadable Demos" link).
So, here's what I did....
1) created a database
dbaccess - -
> create database jtatest with log;
2) installed and registered the Smart Blob Info blade (follow the
instructions in the software for installation):
blademgr
> register sblob_info.1.2 jtatest
3) created and populated the table
create table frag_test
( ID serial,
blob_col blob
)
put blob_col in (sbspace, sbspace2);
insert into frag_test values
(0, FileToBlob('/usr/include/stdio.h', 'server'));
insert into frag_test values
(0, FileToBlob('/usr/include/stdlib.h', 'server'));
Now I can verify with the SblobSbspace() routine
where each smart blob actually got put:
> select id, SblobSbspace(blob_col) from frag_test;
id 1
(expression) sbspace
id 2
(expression) sbspace2
2 row(s) retrieved.
I hope this helps.
-jean
Jean Anderson wrote:
> Another issue might be that 9.1 does not support parallel UDRs
> for smart blobs (your doc reference suggests that you might be
> on 9.1).
>
> 9.2 supports round-robin fragmentation for smart blobs. Also,
> the mi_lo_* DataBlade API functions are parallelizable in 9.2.
> That much said, it supports fragmenting a table that contains a
> smart blob column (i.e., not a single instance of a smart blob,
> but I may be misunderstanding this part of your post....).
>
> This tech note on the IDN DataBlade (*) corner discusses
> parallel UDRs:
>
> http://www.informix.com/idn-secure/DataBlade/Library/Parallel_UDR.htm
>
> regards,
>
> -jean
>
> (*) To access the DataBlade Corner, anybody can register for IDN at
> http://www.informix.com/idn
>
> Noam Atzmon wrote:
>
> > Hi;
> >
> > We are developing a "Face detection over video" blade.
> > Our Compare() function should do a massive IO and CPU work.
> > For this reason we wish to make it work parallel. (Obviously)
> >
> > The problem is that BLOB (smart blob which holds the UDT)
> > can't be fragmented over few SBSpaces (as we understand from
> > the Ans. OnLine (3878.pdf) "1-248 Informix Guide to SQL: Syntax")
> >
> > How can we reach our goal? Any idea?
> >
> > TIA, Noam & Asaf
Noam Atzmon wrote: > We are developing a "Face detection over video" blade. > Our Compare() function should do a massive IO and CPU work. > For this reason we wish to make it work parallel. (Obviously) Actually, it's not as obvious as you might think. The ORDBMS's parallelism is designed with an assumption that each UDR being parallelized is fairly small, and that when you have parallelism enabled, the total amount of memory that all simultaneously executing (parallelized) UDRs consume is less than core memory (usually a lot less). But with this approach, you might find yourelf in the following situation: say each object is 10M in size, and you can parallelize it 10 ways. So you will require 110Meg of memory for the duration of the scan. (This is assuming a scan. It might only be 200Meg for a join). Now, suppose you have 10 concurrently connected users. You're up to a Gig of memory just for the objects involved. And the problem of scaling CPU usage comes up too. That is, 10 way parallelism on a 16 way box means you won't even get to 2 concurrent users before you max out. It isn't that parallelizing these ops is a bad idea. But you might want to use come caution with the degree of parallelism you apply to the problem, particularly in situations where you have a big concurrent user load. There are other ways to do this, involving more complex indexing approaches.