Re: Universe and Windows 2000 Server.
Posted in 2000
Topics: Performance & Tuning, Storage & Space Management, Stored Procedures & SPL, Data Types & Schema Design, Migration, Import/Export & Data Conversion, Java & JDBC Development
From: "Matti Lamprhey" <matti@polka.bikini>
>
>"Obnoxio The Clown" <obnoxio@hotmail.com> wrote...
> > From: "Matti Lamprhey" <matti@polka.bikini>
> > >
> > >I think the use of the word "legacy" is somewhat provocative in the
> > >circumstances, though.
> >
> > Yeah, us dinosaurs using ORDBMS's and datablades take offense
>easily. :-)
>
>Obnoxio, tell me what a datablade is, please. I asked an Informixer
>about this a few weeks ago and the answer was vague in the extreme.
>But I'm sure it's extremely clever, whatever it is.
It *is* very clever. It's a really good idea, conceptually quite simple and
it's really sad that more people haven't adopted it, but I blame that
entirely upon Informix who have marketed it completely wrong, mostly, I
suspect, because they didn't understand it fully themselves.
It is a difficult concept to describe without actually trying it out. I have
CC'd Paul (who's examples I've swiped), because if I screw up the
description or miss anything out, he can correct me (if he doesn't mind!)
The database can be extended by writing new data types and additional
functionality. Whereas most database platforms offer some sort of stored
procedure language, adding new data types is largely beyond them.
Furthermore, the access to writing server-side code is greatly extended by
the ability to make use of existing C or Java code.
Code which has been mangled to make use of the APIs that allow native access
to the server is described as a DataBlade.
As an example: Excalibur had C libraries that allowed you to (amongst other
things) search through text files looking for a set of words in any order
that were within N words of each other.
Let's say that function was called c_near_words(). Now, in a conventional
database, you would have to select your text out of the database in a C
program, then pass the text to the library function in some way (usually by
dumping the text to a file and then invoking the library on the file). This
also means that you have to write some C code for the simplest thing.
By mangling the code to match the Informix API ("mangling" is probably quite
a harsh word for it :-) you have the ability to invoke the text search
directly on the server, greatly boosting performance by reducing network
traffic, file creation, and so on. Also, querying and development becomes
much simpler. You would typically just do something like:
SELECT etx_near_words(fieldname, 5, word, otherword, ...)
FROM table_with_text
WHERE other criteria...
(NOTE: This is just a made up example! :-)
This becomes available in any tool where you can write your own SQL, which
is completely different from the case where you were working with C
libraries.
These are generic datablades, which are typically derived from existing
commercial products that meet some need, like text retrieval, image
scanning, geospatial processing, etc.
However, the real fun starts when you have some business application logic
like a complex proprietary algorithm for (say) analysing credit risk. This
is typically impossible to implement in a stored procedure language due to
complexity and usually written in something like C (or maybe Java), or can
easily be wrapped in C. If you want to do any kind of search for people who
have a rating between X and Y, you would have to write a specific C program
to do this. The number of programs required is massive, and maintenance
becomes a nightmare.
By wrapping the function into a datablade, you can say
SELECT *
FROM customer
WHERE credit_rating(param, param, param) BETWEEN 5 and 7
-- or whatever
from within the standard Informix query tools, or anything that allows you
to write your own SQL. Notice that this is different from pre-calculating
and storing the credit rating, because you can tweak the parameters on the
fly and always get up to the second answers. Also, you don't have to
recalculate and re-write it if an underlying parameter changes.
(An Informix user actually did this exact thing, apparently.)
The other thing, which is even more scary to me, is creating totally new
data types and being able to use (more or less) standard SQL to query them.
I say more or less, because if you create a new data type, you also have to
map "standard" SQL functionality into appropriate routines to perform the
equivalent operation.
A (very) simple example (which can be found at www.informix.com on the IDN
pages) is an opaque data type for long varchars. If the data is <= 255 chars
long, it gets stored in a varchar. If it's longer, it gets stored in a CLOB
type. You as a user don't have to worry about the distinction. However, the
developer has to write functions to allow the management of varchar vs CLOB,
MATCHES, LIKE, UPPER, LOWER, etc, etc, etc.
You can create UDTs that can "understand" things like GIFs or video and
search for certain properties like colours or patterns, or UDTs for your
own, very specific requirements.
While the "commercial" blades are all very nice, it is (IMHO) the ability to
move your existing proprietary, competitive business logic, whether in C or
Java, into the server to consolidate development and present a consistent,
query-oriented interface to that functionality that is truly useful. As
usual, Informix marketing failed to deliver and over emphasised the value of
the commercial datablades and completely forgot to mention the real power.
Hopefully, this will change soon.
With apologies to Sun, I'd say "the database is the development platform".
With a bit of judicious planning, you can extend the database server with
all your clever logic, which means that not only can your clients become
"thinner", but also that even relatively simple tools can get "cleverer".
For in-depth coverage, look into Paul Brown's book on "Developing
Object-Relational Database Applications", due out Real Soon Now (tm) --
allegedly. :)
Oh, and while I'm swiping stuff off Paul, check this out:
http://biz.yahoo.com/prnews/000818/decode_gen.html
Apparently they tried several other database technologies and couldn't even
get close. However, the problem with an example like this is that most
people aren't doing stuff as specific as gene-based Alzheimer's research, so
they think, "Interesting, but too flash and way out for me".
It's not. I'm working on a data warehouse at the moment. One of the major
issues here is large chunks of text which are vital to the use of the
warehouse. We are using Excalibur to search for complex medical data. We are
using an opaque data type to simplify loading and managing the text. This is
a relatively trivial use of a datablade, but has saved us tons of
development time and loading time (and disk space!)
There are lots of "real-world" applications, quite simple to implement, that
can make use of either commercial or "home-grown" datablades. I hope that
people who have made the migration from 7 to 9 try them out.
HTH.
_________________________________________________________________________
Get Your Private, Free E-mail from MSN Hotmail at http://
"Obnoxio The Clown" <obnoxio@hotmail.com> wrote... > From: "Matti Lamprhey" <matti@polka.bikini> > > > >Obnoxio, tell me what a datablade is, please. I asked an Informixer > >about this a few weeks ago and the answer was vague in the extreme. > >But I'm sure it's extremely clever, whatever it is. > > It *is* very clever. It's a really good idea, conceptually quite simple and > it's really sad that more people haven't adopted it, but I blame that > entirely upon Informix who have marketed it completely wrong, mostly, I > suspect, because they didn't understand it fully themselves. > [Excellent explanation snipped] Bravo. Thanks for the enlightenment. It certainly sounds powerful. Matti