Re: timeseries is the solution for this?
Posted in 2013
Topics: Server Administration, Licensing & Editions
On 8 Apr 2013, at 11:14, JJ <JonRitson@sky.com> wrote: > If this is "capturing continuously" netflow traffic then that would equate to: > > 1,400,000,000 / (24 * 60 * 60) entries a second = 16,204 a second. > > Does the "480" bytes include the timestamp? > > I can't see a storage limitation under Innovator C, the main issue would be whether a single CPU VP could handle the traffic, and potentially with some of the "bulk loader" stuff now available in 12.10 I would say it is plausible. > > On Sunday, April 7, 2013 1:50:51 AM UTC+1, Art S. Kagel wrote: >> Well, not Innovator-C, the data volume is prohibative. Advanced Enterprise Edition with Timeseries and Informix Warehouse Accelerator could handle the data volume and the query time requirements but you would need lots of hardware. Memory, processors, and disk. However, FAR less than you would need with any other system. >> >> On Fri, Apr 5, 2013 at 8:17 PM, Cesar Inacio Martins <cesar.inac...@gmail.com> wrote: >> >> This question kicks my curiosity , timeseries is a good solution for this kind of situation? Able to "beat" the opensource solutions? >> >> Quoting part of the question: >> >> "Which database could handle storage of billions/trillions of records?" >> "We are looking at developing a tool to capture and analyze netflow data, of which we gather tremendous amounts of. Each day we capture about ~1.4 billion flow records..." >> "We would like to be able to do fast searches (less than 10 seconds) on the data set..." >> "The idea is to keep approximately one month of data, which would be ~43.2 billion records. A rough estimate that each record would contain about 480 bytes of data, would equate to ~18.7 terabytes of data in a month, and maybe three times that with indexes. Eventually we would like to grow the capacity of this system to store trillions of records." >> The complete question and details, follow the link. >> >> http://dba.stackexchange.com/q/38793/16135 >> And now a question mine. >> Innovator-C + timeseries . Anyone have experience using them on real situation? able to saving million or billions of records? >> And I dare to ask :) able to return fast searchs? >> >> Just saying...I don't have experience with timeseries (waiting for a bootcamp on Brazil), and at this moment I don know how exemplify any real situation where apply this questions.... The 12.10 License appears to introduce a restriction of 8GB of storage. So, this falls at the first hurdle. Apart from that, ingesting the data on good commodity hardware, I would expect that you would not get more than about 7,500-8,000 rows per second in, although if you opted for SSD that might be different. Finally, a lot depends on how you want to query the data. Aggregating multiple TS into a single virtual TS is not trivial.
If you use kdb you have which I have heard is much faster you also get compression which uses a lot less storage. When I mentioned this to Jerry Keesee Director of Development for IDS he said so go with a specialised timeseries database like kdb then. So why is IBM in the 12.1 webcast mentioning timeseries as a differentiating factor for Informix when other specialist databases do it better? David. On 10 April 2013 at 18:37 Spokey Wheeler Gmail <spokey.wheeler@gmail.com> wrote: > On 8 Apr 2013, at 11:14, JJ <JonRitson@sky.com> wrote: > > > If this is "capturing continuously" netflow traffic then that would equate > to: > > > > 1,400,000,000 / (24 * 60 * 60) entries a second = 16,204 a second. > > > > Does the "480" bytes include the timestamp? > > > > I can't see a storage limitation under Innovator C, the main issue would be > whether a single CPU VP could handle the traffic, and potentially with some of > the "bulk loader" stuff now available in 12.10 I would say it is plausible. > > > > On Sunday, April 7, 2013 1:50:51 AM UTC+1, Art S. Kagel wrote: > >> Well, not Innovator-C, the data volume is prohibative. Advanced Enterprise > Edition with Timeseries and Informix Warehouse Accelerator could handle the > data volume and the query time requirements but you would need lots of > hardware. Memory, processors, and disk. However, FAR less than you would need > with any other system. > >> > >> On Fri, Apr 5, 2013 at 8:17 PM, Cesar Inacio Martins > <cesar.inac...@gmail.com> wrote: > >> > >> This question kicks my curiosity , timeseries is a good solution for this > kind of situation? Able to "beat" the opensource solutions? > >> > >> Quoting part of the question: > >> > >> "Which database could handle storage of billions/trillions of records?" > >> "We are looking at developing a tool to capture and analyze netflow data, > of which we gather tremendous amounts of. Each day we capture about ~1.4 > billion flow records..." > >> "We would like to be able to do fast searches (less than 10 seconds) on the > data set..." > >> "The idea is to keep approximately one month of data, which would be ~43.2 > billion records. A rough estimate that each record would contain about 480 > bytes of data, would equate to ~18.7 terabytes of data in a month, and maybe > three times that with indexes. Eventually we would like to grow the capacity > of this system to store trillions of records." > >> The complete question and details, follow the link. > >> > >> http://dba.stackexchange.com/q/38793/16135 > >> And now a question mine. > >> Innovator-C + timeseries . Anyone have experience using them on real > situation? able to saving million or billions of records? > >> And I dare to ask :) able to return fast searchs? > >> > >> Just saying...I don't have experience with timeseries (waiting for a > bootcamp on Brazil), and at this moment I don know how exemplify any real > situation where apply this questions.... > > The 12.10 License appears to introduce a restriction of 8GB of storage. So, > this falls at the first hurdle. > > Apart from that, ingesting the data on good commodity hardware, I would expect > that you would not get more than about 7,500-8,000 rows per second in, > although if you opted for SSD that might be different. > > Finally, a lot depends on how you want to query the data. Aggregating multiple > TS into a single virtual TS is not trivial. > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. >
Hmmm. I don't remember saying that. Our Time Series technology is a differentiator. We've published benchmarks demonstrating how well we can both load as well as perform analytics / reporting. We're the only relational database with native TS built in. Time Series also leverages its own compression techniques which we've shown results in tremendous storage saving. Jerry Keesee World Wide Director, Informix Software IBM Data Management Tel:(913)599-8713 Cell:(913)710-7472 From: "david@smooth1.co.uk" <david@smooth1.co.uk> To: ids@iiug.org Date: 04/10/2013 03:29 PM Subject: Re: timeseries is the solution for this? [30029] Sent by: ids-bounces@iiug.org If you use kdb you have which I have heard is much faster you also get compression which uses a lot less storage. When I mentioned this to Jerry Keesee Director of Development for IDS he said so go with a specialised timeseries database like kdb then. So why is IBM in the 12.1 webcast mentioning timeseries as a differentiating factor for Informix when other specialist databases do it better? David. On 10 April 2013 at 18:37 Spokey Wheeler Gmail <spokey.wheeler@gmail.com> wrote: > On 8 Apr 2013, at 11:14, JJ <JonRitson@sky.com> wrote: > > > If this is "capturing continuously" netflow traffic then that would equate > to: > > > > 1,400,000,000 / (24 * 60 * 60) entries a second = 16,204 a second. > > > > Does the "480" bytes include the timestamp? > > > > I can't see a storage limitation under Innovator C, the main issue would be > whether a single CPU VP could handle the traffic, and potentially with some of > the "bulk loader" stuff now available in 12.10 I would say it is plausible. > > > > On Sunday, April 7, 2013 1:50:51 AM UTC+1, Art S. Kagel wrote: > >> Well, not Innovator-C, the data volume is prohibative. Advanced Enterprise > Edition with Timeseries and Informix Warehouse Accelerator could handle the > data volume and the query time requirements but you would need lots of > hardware. Memory, processors, and disk. However, FAR less than you would need > with any other system. > >> > >> On Fri, Apr 5, 2013 at 8:17 PM, Cesar Inacio Martins > <cesar.inac...@gmail.com> wrote: > >> > >> This question kicks my curiosity , timeseries is a good solution for this > kind of situation? Able to "beat" the opensource solutions? > >> > >> Quoting part of the question: > >> > >> "Which database could handle storage of billions/trillions of records?" > >> "We are looking at developing a tool to capture and analyze netflow data, > of which we gather tremendous amounts of. Each day we capture about ~1.4 > billion flow records..." > >> "We would like to be able to do fast searches (less than 10 seconds) on the > data set..." > >> "The idea is to keep approximately one month of data, which would be ~43.2 > billion records. A rough estimate that each record would contain about 480 > bytes of data, would equate to ~18.7 terabytes of data in a month, and maybe > three times that with indexes. Eventually we would like to grow the capacity > of this system to store trillions of records." > >> The complete question and details, follow the link. > >> > >> http://dba.stackexchange.com/q/38793/16135 > >> And now a question mine. > >> Innovator-C + timeseries . Anyone have experience using them on real > situation? able to saving million or billions of records? > >> And I dare to ask :) able to return fast searchs? > >> > >> Just saying...I don't have experience with timeseries (waiting for a > bootcamp on Brazil), and at this moment I don know how exemplify any real > situation where apply this questions.... > > The 12.10 License appears to introduce a restriction of 8GB of storage. So, > this falls at the first hurdle. > > Apart from that, ingesting the data on good commodity hardware, I would expect > that you would not get more than about 7,500-8,000 rows per second in, > although if you opted for SSD that might be different. > > Finally, a lot depends on how you want to query the data. Aggregating multiple > TS into a single virtual TS is not trivial. > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > ******************************************************************************* Forum Note: Use "Reply" to post a response in the discussion forum.
Hi Spokey , You right... other "nice" surprise from IBM.... just like the Jack the Ripper... removing part-by-part... From Innovator-C license : 1.3 Database Size Limit The Program allows a maximum of eight (8) Gigabytes of database server storage. Probably the next version will become with limit of user session.. 2013/4/10 Spokey Wheeler Gmail <spokey.wheeler@gmail.com> > On 8 Apr 2013, at 11:14, JJ <JonRitson@sky.com> wrote: > > > If this is "capturing continuously" netflow traffic then that would > equate > to: > > > > 1,400,000,000 / (24 * 60 * 60) entries a second = 16,204 a second. > > > > Does the "480" bytes include the timestamp? > > > > I can't see a storage limitation under Innovator C, the main issue would > be > whether a single CPU VP could handle the traffic, and potentially with > some of > the "bulk loader" stuff now available in 12.10 I would say it is plausible. > > > > On Sunday, April 7, 2013 1:50:51 AM UTC+1, Art S. Kagel wrote: > >> Well, not Innovator-C, the data volume is prohibative. Advanced > Enterprise > Edition with Timeseries and Informix Warehouse Accelerator could handle the > data volume and the query time requirements but you would need lots of > hardware. Memory, processors, and disk. However, FAR less than you would > need > with any other system. > >> > >> On Fri, Apr 5, 2013 at 8:17 PM, Cesar Inacio Martins > <cesar.inac...@gmail.com> wrote: > >> > >> This question kicks my curiosity , timeseries is a good solution for > this > kind of situation? Able to "beat" the opensource solutions? > >> > >> Quoting part of the question: > >> > >> "Which database could handle storage of billions/trillions of records?" > >> "We are looking at developing a tool to capture and analyze netflow > data, > of which we gather tremendous amounts of. Each day we capture about ~1.4 > billion flow records..." > >> "We would like to be able to do fast searches (less than 10 seconds) on > the > data set..." > >> "The idea is to keep approximately one month of data, which would be > ~43.2 > billion records. A rough estimate that each record would contain about 480 > bytes of data, would equate to ~18.7 terabytes of data in a month, and > maybe > three times that with indexes. Eventually we would like to grow the > capacity > of this system to store trillions of records." > >> The complete question and details, follow the link. > >> > >> http://dba.stackexchange.com/q/38793/16135 > >> And now a question mine. > >> Innovator-C + timeseries . Anyone have experience using them on real > situation? able to saving million or billions of records? > >> And I dare to ask :) able to return fast searchs? > >> > >> Just saying...I don't have experience with timeseries (waiting for a > bootcamp on Brazil), and at this moment I don know how exemplify any real > situation where apply this questions.... > > The 12.10 License appears to introduce a restriction of 8GB of storage. So, > this falls at the first hurdle. > > Apart from that, ingesting the data on good commodity hardware, I would > expect > that you would not get more than about 7,500-8,000 rows per second in, > although if you opted for SSD that might be different. > > Finally, a lot depends on how you want to query the data. Aggregating > multiple > TS into a single virtual TS is not trivial. > > > > ******************************************************************************* > Forum Note: Use "Reply" to post a response in the discussion forum. > > --14dae9cc9f969d894604da1ff175