Re: storing pdf's in database
Posted in 2008
Topics: Performance & Tuning, Storage & Space Management, Connectivity: ODBC / JDBC / .NET, Connectivity: ESQL/C, 4GL & Embedded SQL, Security, Permissions & Auditing, Platform-Specific Issues, Java & JDBC Development
Sorry for not understanding all of this. What do you guys mean when you mention "indexing" the data. I was originally thinking it meant keeping some sort of pointer, or filedescriptor in the database ( that informix knows how to get the file from the OS ). But now I am thinking you mean that Informix can understand the guts of the blob image ( pdf, doc file whatever ), and have indexes to different parts of the it. How does that work ? Thank you, Floyd ----- Original Message ----- Subject: Re: storing pdf's in database Date: Fri, April 11, 2008 13:49 From: "Paul Watson (Oninit)" <paul@oninit.com> If you are not indexing the data in the PDF then your database will not > be corrupt - you just get 'file not found' type error. If it is indexed > the you are right you can get a mismatch between the data the DB thinks > it has and the file on the file system. Doesn't matter where the data > is held no one is safe from an idiot with a root password. > > Another thing to consider is the backup - if you have 10+ M images at > 100K each or a few K of PDFs at 10M each and they are essentially static > data then they might as well just be on the FS and backed up once in a > blue moon as opposed to every level zero. > > Cheers > Paul > > Floyd Wellershaus wrote: > > Very good point Gumby. Thanks for that. You're right. We don't want my > > app developer coming to me saying that the database is corrupt, because > > someone futzed with a file on the OS. :-) > > > > > > > > > > > > ----- Original Message ----- > > Subject: RE: storing pdf's in database > > From: "Ian Michael Gumby" <im_gumby@hotmail.com> > > Date: Fri, April 11, 2008 10:50 > > > > > > Personally I'd use Python over Perl. (Python is the *new* Perl... ;-) > > Its smokin! and Carsten did the adapter right. (unlike cx_Oracle).... > > > > But the point that you have to consider is that if you store the .pdf > > outside of the database, you tend to lose control over the document. > > That is, even if you create a file system that is owned by Informix, and > > only Informix has r/w permissions, root will still have access. > > (Assuming Linux or Unix) So the documents can be changed, removed, etc > > and you'd have no auditing capabilities. (Ok, so you actually can use a > > third party admin tool to audit the filesystem, but that doesn't change > > the fact that IDS doesn't know of any changes) > > > > Because you can get extremely good performance from IDS in streaming a > > blob from the database, I'd recommend saving the .pdf as a blob. > > > > But then again, I'm a paranoid architect, and it depends on what .pdfs > > you're saving. If you took the time to make them a .pdf, I would imagine > > that you'd want control and be paranoid too. > > > > But what do I know? I'm saving the world from people who blindly follow > > their nav systems and drive in to ponds, lakes, and people's homes. ;-) > > > > -G > > > > ------------------------------------------------------------------------ > > To: informix-list@iiug.org > > Subject: Re: storing pdf's in database > > From: mvakeel@us.ibm.com > > Date: Fri, 11 Apr 2008 09:16:46 -0500 > > > > > > I have helped a customer where we use perl(DBI/DBD) to decipher > > e-mail attachments and upload it to smart blobs in the database, > > and on the front end they use the Web Datablade module to retrieve, > > view and manage the documents. Have had no issues with performance > > so far. I doubt there would be a best recommended method as such, > > > > -Manoj > > > > > > > > *caver <dmcbryde@courts.state.va.us>* > > Sent by: informix-list-bounces@iiug.org 04/11/2008 08:10 AM > > > > To > > informix-list@iiug.org > > cc > > > > Subject > > Re: storing pdf's in database > > > > > > > > > > > > > > > > > > > > On Apr 10, 4:01 pm, "Floyd Wellershaus" <fl...@fwellers.com> wrote: > > > Paul Thank you. > > > So where did the indexes get stored, in a regular table dbspace > > or did you > > > need an sbspace ? > > > > > > Thanks, > > > floyd > > > > > > > > > > > > > > > > > > ----- Original Message ----- > > > Subject: Re: storing pdf's in database > > > Date: Thu, April 10, 2008 15:18 > > > From: "Paul Watson (Oninit)" <p...@oninit.com> > > > Floyd Wellershaus wrote: > > > > > We have IDS10.0FC5 on Aix5.3. > > > > > > > > Is there a best recommended method for storing and retrieving > > pdf files > > > > > to the database, that would be put there through an odbc > > connection > > > > > coming from powerbuilder ? > > > > > > > > I am thinking a clob or blob in an sbspace that would be > > accessed via > > > esqlc. > > > > > > > > As an alternative, is there a datablade that would allow the > > file to be > > > > > stored on the unix OS with the pointer existing in the > > database ? > > > > > > > > Thanks much, > > > > > Floyd > > > > > > > I used to use the ETX blade for this. But the actual PDFs were > > stored > > > > on the filesystem, only the indices where stored in the > > database. The > > > > ETX blade is very good on PDFs > > > > > > > _______________________________________________ > > > > Informix-list mailing list > > > > Informix-l...@iiug.org > > > >http://www.iiug.org/mailman/listinfo/informix-list- Hide quoted > > text - > > > > > > - Show quoted text - > > > > We store images as bytes in a dbspace so that we can replicate the > > images > > to another server via ER. > > We have fairly consistent file sizes, so we did not need smartblobs. > > For 4+ million images the > > retrieve times have never been a problem. We have a keyword index > > stored as > > another table in dbspace. That way we let informix handle all the > > storage and just > > make sure we feed the engine enough raw disk space to keep shoving in > > images. > > We use odbc, but VB (not powerbuilder) as a front end. > > The setup seems to work fairly well. > > > > _______________________________________________ > > Informix-list mailing list > > Informix-list@iiug.org > > http://www.iiug.org/mailman/listinfo/informix-list > > > > > > ------------------------------------------------------------------------ > > Get in touch in an instant. Get Windows Live Messenger now. > > <http://www.windowslive.com/messenger/overview.html?ocid=TXT_TAGLM_WL_Refresh_getintouch_042008> > > ------------------------------------------------------------------------ > > > > _______________________________________________ > > Informix-list mailing list > > Informix-list@iiug.org <javascript:bodyCreateMail('Informix-list%40iiug.org')> > > http://www.iiug.or
If you are using the textblade then the actual contents of the PDF are index'able. The textblade comes with filters that allow you to index the contents of most common file formats. We use it to search, among other things, the IDS manual sets and our website Cheers Paul Floyd Wellershaus wrote: > Sorry for not understanding all of this. What do you guys mean when you > mention "indexing" the data. > I was originally thinking it meant keeping some sort of pointer, or > filedescriptor in the database ( that informix knows how to get the file > from the OS ). But now I am thinking you mean that Informix can understand > the guts of the blob image ( pdf, doc file whatever ), and have indexes to > different parts of the it. > > How does that work ? > > Thank you, > Floyd > > > > ----- Original Message ----- > Subject: Re: storing pdf's in database > Date: Fri, April 11, 2008 13:49 > From: "Paul Watson (Oninit)" <paul@oninit.com> > If you are not indexing the data in the PDF then your database will not >> be corrupt - you just get 'file not found' type error. If it is indexed >> the you are right you can get a mismatch between the data the DB thinks >> it has and the file on the file system. Doesn't matter where the data >> is held no one is safe from an idiot with a root password. >> >> Another thing to consider is the backup - if you have 10+ M images at >> 100K each or a few K of PDFs at 10M each and they are essentially static >> data then they might as well just be on the FS and backed up once in a >> blue moon as opposed to every level zero. >> >> Cheers >> Paul >> >> Floyd Wellershaus wrote: >>> Very good point Gumby. Thanks for that. You're right. We don't want my >>> app developer coming to me saying that the database is corrupt, because >>> someone futzed with a file on the OS. :-) >>> >>> >>> >>> >>> >>> ----- Original Message ----- >>> Subject: RE: storing pdf's in database >>> From: "Ian Michael Gumby" <im_gumby@hotmail.com> >>> Date: Fri, April 11, 2008 10:50 >>> >>> >>> Personally I'd use Python over Perl. (Python is the *new* Perl... ;-) >>> Its smokin! and Carsten did the adapter right. (unlike cx_Oracle).... >>> >>> But the point that you have to consider is that if you store the .pdf >>> outside of the database, you tend to lose control over the document. >>> That is, even if you create a file system that is owned by Informix, and >>> only Informix has r/w permissions, root will still have access. >>> (Assuming Linux or Unix) So the documents can be changed, removed, etc >>> and you'd have no auditing capabilities. (Ok, so you actually can use a >>> third party admin tool to audit the filesystem, but that doesn't change >>> the fact that IDS doesn't know of any changes) >>> >>> Because you can get extremely good performance from IDS in streaming a >>> blob from the database, I'd recommend saving the .pdf as a blob. >>> >>> But then again, I'm a paranoid architect, and it depends on what .pdfs >>> you're saving. If you took the time to make them a .pdf, I would imagine >>> that you'd want control and be paranoid too. >>> >>> But what do I know? I'm saving the world from people who blindly follow >>> their nav systems and drive in to ponds, lakes, and people's homes. ;-) >>> >>> -G >>> >>> ------------------------------------------------------------------------ >>> To: informix-list@iiug.org >>> Subject: Re: storing pdf's in database >>> From: mvakeel@us.ibm.com >>> Date: Fri, 11 Apr 2008 09:16:46 -0500 >>> >>> >>> I have helped a customer where we use perl(DBI/DBD) to decipher >>> e-mail attachments and upload it to smart blobs in the database, >>> and on the front end they use the Web Datablade module to retrieve, >>> view and manage the documents. Have had no issues with performance >>> so far. I doubt there would be a best recommended method as such, >>> >>> -Manoj >>> >>> >>> >>> *caver <dmcbryde@courts.state.va.us>* >>> Sent by: informix-list-bounces@iiug.org 04/11/2008 08:10 AM >>> >>> To >>> informix-list@iiug.org >>> cc >>> >>> Subject >>> Re: storing pdf's in database >>> >>> >>> >>> >>> >>> >>> >>> >>> >>> On Apr 10, 4:01 pm, "Floyd Wellershaus" <fl...@fwellers.com> wrote: >>> > Paul Thank you. >>> > So where did the indexes get stored, in a regular table dbspace >>> or did you >>> > need an sbspace ? >>> > >>> > Thanks, >>> > floyd >>> > >>> > >>> > >>> > >>> > >>> > ----- Original Message ----- >>> > Subject: Re: storing pdf's in database >>> > Date: Thu, April 10, 2008 15:18 >>> > From: "Paul Watson (Oninit)" <p...@oninit.com> >>> > Floyd Wellershaus wrote: >>> > > > We have IDS10.0FC5 on Aix5.3. >>> > >>> > > > Is there a best recommended method for storing and retrieving >>> pdf files >>> > > > to the database, that would be put there through an odbc >>> connection >>> > > > coming from powerbuilder ? >>> > >>> > > > I am thinking a clob or blob in an sbspace that would be >>> accessed via >>> > esqlc. >>> > >>> > > > As an alternative, is there a datablade that would allow the >>> file to be >>> > > > stored on the unix OS with the pointer existing in the >>> database ? >>> > >>> > > > Thanks much, >>> > > > Floyd >>> > >>> > > I used to use the ETX blade for this. But the actual PDFs were >>> stored >>> > > on the filesystem, only the indices where stored in the >>> database. The >>> > > ETX blade is very good on PDFs >>> > >>> > > _______________________________________________ >>> > > Informix-list mailing list >>> > > Informix-l...@iiug.org >>> > >http://www.iiug.org/mailman/listinfo/informix-list- Hide quoted >>> text - >>> > >>> > - Show quoted text - >>> >>> We store images as bytes in a dbspace so that we can replicate the >>> images >>> to another server via ER. >>> We have fairly consistent file sizes, so we did not need smartblobs. >>> For 4+ million images the >>> retrieve times have never been a problem. We have a keyword index >>> stored as >>> another table in dbspace. That way we let informix handle all the >>> storage and just >>> make sure we feed the engine enough raw disk space to keep shoving in >>> images. >>> We use odbc, but VB (not powerbuilder) as a front end. >>> The setup seems to work fairly well. >>> >>> _______________________________________________ >>> Informix-list mailing list >>> Informix-list@iiug.org >>> http://www.iiug.org/mailman/listinfo/informix-list >>> >>> >>> ------------------------------------------------------------------------ >>> Get in touch in an instant. Get Windows Live Messenger now.@
Paul, You're missing the point. For "real" applications that deal with managing archives of .pdfs, you are dealing with data where "file not found" is not a good thing. In today's market, you don't want to store any document outside of the database and just "index" it in a database. You're asking for a lawsuit. (Yeah I know that IBM's DB2 solution does this....) Here's an example ... think of a PACs system. Its a system to store medical images in a digital form rather than film. (X-Rays, MRIs, CTs, Ultrasound, ...) So you can either store them in the database, or you can just use the database to "index" the studies and then store them on the file system. But you have to consider HIPPA compliance. You have to be able to audit who saw what and when. If the studies are not stored in the database, then anyone who knows where the images are located, can then access the image outside of your control. Again, it means having root access to your unix server. Switch industries and consider legal documents. Same problem. Different regulations. Imagine if your law firm does M&A and information gets "leaked". You can bet that the SEC, and a couple of other Fed organizations will be all over your arse.... Then you have the additional headache of litigation. Switch industries to the financial industry. You have derivative contracts. Again you have audit issues. (SIVs,MBOs,Swaps, Swaptions, CDOs, etc ... all are contracts.) So if you're going to be storing them in a digital format, you need to make sure that they are secure. So its always a better idea to have the database contain the actual .pdf than not. We're not talking about WORM drives anymore. You're talking about replication of SAN volumes these days to multiple data centers. And I realize that I'm talking about multi-million dollar IT investments and not a simple shareware app. But the same principal should be applied to either environment. Today's startup could be tomorrow's billion dollar company. But hey! What do I know? Its not like I've looked at PACs or wrote an OBS portfolio management system, or dealt with a bunch of lawyers and auditors... ;-) -G > Date: Fri, 11 Apr 2008 13:03:40 -0500 > From: paul@oninit.com > Subject: Re: storing pdf's in database > To: informix-list@iiug.org > > If you are using the textblade then the actual contents of the PDF are > index'able. The textblade comes with filters that allow you to index the > contents of most common file formats. > > We use it to search, among other things, the IDS manual sets and our website > > Cheers > Paul > > Floyd Wellershaus wrote: > > Sorry for not understanding all of this. What do you guys mean when you > > mention "indexing" the data. > > I was originally thinking it meant keeping some sort of pointer, or > > filedescriptor in the database ( that informix knows how to get the file > > from the OS ). But now I am thinking you mean that Informix can understand > > the guts of the blob image ( pdf, doc file whatever ), and have indexes to > > different parts of the it. > > > > How does that work ? > > > > Thank you, > > Floyd > > > > > > > > ----- Original Message ----- > > Subject: Re: storing pdf's in database > > Date: Fri, April 11, 2008 13:49 > > From: "Paul Watson (Oninit)" <paul@oninit.com> > > If you are not indexing the data in the PDF then your database will not > >> be corrupt - you just get 'file not found' type error. If it is indexed > >> the you are right you can get a mismatch between the data the DB thinks > >> it has and the file on the file system. Doesn't matter where the data > >> is held no one is safe from an idiot with a root password. > >> > >> Another thing to consider is the backup - if you have 10+ M images at > >> 100K each or a few K of PDFs at 10M each and they are essentially static > >> data then they might as well just be on the FS and backed up once in a > >> blue moon as opposed to every level zero. > >> > >> Cheers > >> Paul > >> > >> Floyd Wellershaus wrote: > >>> Very good point Gumby. Thanks for that. You're right. We don't want my > >>> app developer coming to me saying that the database is corrupt, because > >>> someone futzed with a file on the OS. :-) > >>> > >>> > >>> > >>> > >>> > >>> ----- Original Message ----- > >>> Subject: RE: storing pdf's in database > >>> From: "Ian Michael Gumby" <im_gumby@hotmail.com> > >>> Date: Fri, April 11, 2008 10:50 > >>> > >>> > >>> Personally I'd use Python over Perl. (Python is the *new* Perl... ;-) > >>> Its smokin! and Carsten did the adapter right. (unlike cx_Oracle).... > >>> > >>> But the point that you have to consider is that if you store the .pdf > >>> outside of the database, you tend to lose control over the document. > >>> That is, even if you create a file system that is owned by Informix, and > >>> only Informix has r/w permissions, root will still have access. > >>> (Assuming Linux or Unix) So the documents can be changed, removed, etc > >>> and you'd have no auditing capabilities. (Ok, so you actually can use a > >>> third party admin tool to audit the filesystem, but that doesn't change > >>> the fact that IDS doesn't know of any changes) > >>> > >>> Because you can get extremely good performance from IDS in streaming a > >>> blob from the database, I'd recommend saving the .pdf as a blob. > >>> > >>> But then again, I'm a paranoid architect, and it depends on what .pdfs > >>> you're saving. If you took the time to make them a .pdf, I would imagine > >>> that you'd want control and be paranoid too. > >>> > >>> But what do I know? I'm saving the world from people who blindly follow > >>> their nav systems and drive in to ponds, lakes, and people's homes. ;-) > >>> > >>> -G > >>> > >>> ------------------------------------------------------------------------ > >>> To: informix-list@iiug.org > >>> Subject: Re: storing pdf's in database > >>> From: mvakeel@us.ibm.com > >>> Date: Fri, 11 Apr 2008 09:16:46 -0500 > >>> > >>> > >>> I have helped a customer where we use perl(DBI/DBD) to decipher > >>> e-mail attachments and upload it to smart blobs in the database, > >>> and on the front end they use the Web Datablade module to retrieve, > >>> view and manage the documents. Have had no issues with performance > >>> so far. I doubt there would be a best recommended method as such, > >>> > >>> -Manoj > >>> > >>> > >>> > >>> *caver <dmcbryde@courts.state.va.us>* > >>> Sent by: informix-list-bounces@iiug.org 04/11/2008 08:10 AM > >>> > >>> To > >>> informix-list@iiug.org > >>> cc > >>> > >>> Subject > >>> Re: storing pdf's in database > >>> > >>> > >>> > >>> > >>> > >>> > >>> > >>> > >>> > >>> On Apr 10, 4:01 pm, "Floyd Wellershaus" <fl...@fwellers.com> wrote: > >>> > Paul Thank you. > >>> > So where did the indexes get stored, in a regular table dbspace > >>> or did you > >>> > need an sbspace ? > >>> > > >>> > Thanks, > >>> > floyd > >>>