Looking for ESQL/C Multithreaded Source Code Example
Posted in 2010
A user wanted a working multi-threaded ESQL/C example (the thread_safe sample in the manual wouldn't compile) to speed up inserts into a table fragmented over 8 dbspaces with indexes and foreign keys, since 8 parallel processes loaded 2M rows in ~10 minutes versus ~1 hour serially. Replies offered advice rather than code: use a standard C producer/consumer (librarian-reader mutex) pattern with one Informix connection per thread; check page size, I/O saturation, logging/indexes, and consider HPL; and Art Kagel argued multiple separate processes fed by a pipe, message queue or shared memory would be as fast or faster and simpler than threads. No sample code or confirmed resolution is recorded.
Auto-generated by DrWatson from the posts below — may be imperfect; read the full thread.
Topics: Storage & Space Management, Connectivity: ESQL/C, 4GL & Embedded SQL, Triggers, Constraints & Referential Integrity, Platform-Specific Issues, Versions, Editions & End-of-Life
Greetings All, I am looking for an example of functional ESQL/C source code that demonstrates multi-threading. I have found the thread_safe example in the ESQL/C programmers guide, but I have not had any luck getting it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc or anywhere on the net. The problem I am trying to solve is an INSERT bottleneck. My table is partitioned (fragmented) across 8 dbspaces with multiple indexes and foreign keys containing ~ 800,000,000 rows. A serial load will INSERT 2,000,000 record in ~ 1 hour while a parallel load of 8 processes takes about 10 minutes. Suggestions in this group seem to point toward having a reader provide the data to a set of INSERT consumers. Given the circumstances, this would probably not be a good case for utilizing HPL as it would be be beneficial for us to keep everything consolidated within ESQL/C. I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. Thanks in advance for you insight,
Huh? It sounds like you already have a multi-threaded reader. You said 2 million rows in 1 hour in a single thread. 8 million rows in 10 mins. Or were you extrapolating? Look here's the skinny. Look at the Librarian/Reader problem. 1 librarian, n readers. (Simple mutex semaphore example that was written back in the 60's by Djikstra This is standard c thread code. Each reader then creates its own connection to IDS. (Meaning that its a single thread and hence thread safe.) And that's it. When there's no more data, then the thread shuts down and dies. Nothing fancy and very simple. So you need to look at threading in C, and then how to write a single threaded connection to Informix. K? Here's a link to get you started http://74.125.95.132/search?q=cache:ewOhC9kYNyAJ:www.codeproject.com/KB/threads/mutexrw.aspx%3Fmsg%3D923563+librarian+reader+semaphore+problem+Djikstra&cd=2&hl=en&ct=clnk&gl=us&client=firefox-a (The page was cached by google, you can find other pages searching on Djikstra Librarian mutex ...) But hey! What do I know? Its not like I've written something like this 100s of times. ;-) -G > From: the_omegamon@yahoo.com > Subject: Looking for ESQL/C Multithreaded Source Code Example > Date: Sat, 13 Feb 2010 13:08:31 -0800 > To: informix-list@iiug.org > > Greetings All, > > I am looking for an example of functional ESQL/C source code that > demonstrates multi-threading. I have found the thread_safe example > in the ESQL/C programmers guide, but I have not had any luck getting > it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc > or anywhere on the net. > > The problem I am trying to solve is an INSERT bottleneck. My table > is > partitioned (fragmented) across 8 dbspaces with multiple indexes and > foreign keys containing ~ 800,000,000 rows. A serial load will > INSERT > 2,000,000 record in ~ 1 hour while a parallel load of 8 processes > takes > about 10 minutes. Suggestions in this group seem to point toward > having > a reader provide the data to a set of INSERT consumers. Given the > circumstances, this would probably not be a good case for utilizing > HPL > as it would be be beneficial for us to keep everything consolidated > within ESQL/C. > > I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. > > Thanks in advance for you insight, > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list _________________________________________________________________ Hotmail: Powerful Free email with security by Microsoft. http://clk.atdmt.com/GBL/go/201469230/direct/01/
Not exactly. I created a representative input file and used John Miller's bloader program to run a single stream load and additional parallel loads at 2/4/8/16/32 writers. The bloader (circa 1994) is forking processes not threads. It was useful supporting the idea that throughput could be increased substantially. Thanks for the link. I was hoping to roll a little faster and smoother by using an already existing wheel, if one was available. Given the tremendous advantage of this approach, I was somewhat surprised by the lack of a template. I'm probably just not looking in the right place and if anyone knows where that place is, they are most likely on this list. On Feb 13, 5:54 pm, Ian Michael Gumby <im_gu...@hotmail.com> wrote: > Huh? > It sounds like you already have a multi-threaded reader. > You said 2 million rows in 1 hour in a single thread. > 8 million rows in 10 mins. > > Or were you extrapolating? > > Look here's the skinny. > > Look at the Librarian/Reader problem. 1 librarian, n readers. (Simple mutex semaphore example that was written back in the 60's by Djikstra > This is standard c thread code. > > Each reader then creates its own connection to IDS. (Meaning that its a single thread and hence thread safe.) > > And that's it. > When there's no more data, then the thread shuts down and dies. > > Nothing fancy and very simple. So you need to look at threading in C, and then how to write a single threaded connection to Informix. > > K? > Here's a link to get you startedhttp://74.125.95.132/search?q=cache:ewOhC9kYNyAJ:www.codeproject.com/... > > (The page was cached by google, you can find other pages searching on Djikstra Librarian mutex ...) > > But hey! What do I know? Its not like I've written something like this 100s of times. ;-) > > -G > > > > > From: the_omega...@yahoo.com > > Subject: Looking for ESQL/C Multithreaded Source Code Example > > Date: Sat, 13 Feb 2010 13:08:31 -0800 > > To: informix-l...@iiug.org > > > Greetings All, > > > I am looking for an example of functional ESQL/C source code that > > demonstrates multi-threading. I have found the thread_safe example > > in the ESQL/C programmers guide, but I have not had any luck getting > > it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc > > or anywhere on the net. > > > The problem I am trying to solve is an INSERT bottleneck. My table > > is > > partitioned (fragmented) across 8 dbspaces with multiple indexes and > > foreign keys containing ~ 800,000,000 rows. A serial load will > > INSERT > > 2,000,000 record in ~ 1 hour while a parallel load of 8 processes > > takes > > about 10 minutes. Suggestions in this group seem to point toward > > having > > a reader provide the data to a set of INSERT consumers. Given the > > circumstances, this would probably not be a good case for utilizing > > HPL > > as it would be be beneficial for us to keep everything consolidated > > within ESQL/C. > > > I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. > > > Thanks in advance for you insight, > > _______________________________________________ > > Informix-list mailing list > > Informix-l...@iiug.org > >http://www.iiug.org/mailman/listinfo/informix-list > > _________________________________________________________________ > Hotmail: Powerful Free email with security by Microsoft.http://clk.atdmt.com/GBL/go/201469230/direct/01/
the_omegamon@yahoo.com schrieb: > Greetings All, > > I am looking for an example of functional ESQL/C source code that > demonstrates multi-threading. I have found the thread_safe example > in the ESQL/C programmers guide, but I have not had any luck getting > it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc > or anywhere on the net. > > The problem I am trying to solve is an INSERT bottleneck. My table > is > partitioned (fragmented) across 8 dbspaces with multiple indexes and > foreign keys containing ~ 800,000,000 rows. A serial load will > INSERT > 2,000,000 record in ~ 1 hour while a parallel load of 8 processes > takes > about 10 minutes. Suggestions in this group seem to point toward > having > a reader provide the data to a set of INSERT consumers. Given the > circumstances, this would probably not be a good case for utilizing > HPL > as it would be be beneficial for us to keep everything consolidated > within ESQL/C. > > I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. > > Thanks in advance for you insight, not much information for a not_to_small thingy. What is your page size and how many rows per page are in this table? What is your disk hardware? Subsystem (NAS / MAS & iSCCI / FC connected something / DAS)? Do you have logging and indexes enables during load? In my book you need as many parallelism on input AND output, as until your I/O is saturated. This saturation point should be much higher than ~3300 recs/sec. The values you give also shows, that what you did does not scale too well: you went 8 way parallel, but only got 6 fold improvement. One way to go there, and a very cheap one, might be to find a way to split input into 8 parts (like using mod or somesuch) and use 8 prozesses started at the same time. Even Slowaris 10 should be able to read only once physically and 7 of the 8 brothers should read from memory buffers, which normally on plastic hardware is a lil faster, like 200 times if the physical read is not from SSD and even on big, expensive *nix harware it is like 120 times faster. HPL is always worth a try, as using HPL it is possible to have parallel writes together with parallel reading. I believe you must know what you hardware + OS can do and then using IDS over all the other database systems out there makes me always using 95% of this possible performance. dic_k -- Richard Kofler SOLID STATE EDV Dienstleistungen GmbH Vienna/Austria/Europe
On Feb 14, 4:57 am, Richard Kofler <richard.kof...@chello.at> wrote: > the_omega...@yahoo.com schrieb: > > > > > > > Greetings All, > > > I am looking for an example of functional ESQL/C source code that > > demonstrates multi-threading. I have found the thread_safe example > > in the ESQL/C programmers guide, but I have not had any luck getting > > it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc > > or anywhere on the net. > > > The problem I am trying to solve is an INSERT bottleneck. My table > > is > > partitioned (fragmented) across 8 dbspaces with multiple indexes and > > foreign keys containing ~ 800,000,000 rows. A serial load will > > INSERT > > 2,000,000 record in ~ 1 hour while a parallel load of 8 processes > > takes > > about 10 minutes. Suggestions in this group seem to point toward > > having > > a reader provide the data to a set of INSERT consumers. Given the > > circumstances, this would probably not be a good case for utilizing > > HPL > > as it would be be beneficial for us to keep everything consolidated > > within ESQL/C. > > > I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. > > > Thanks in advance for you insight, > > not much information for a not_to_small thingy. > What is your page size and how many rows per page are in this table? > What is your disk hardware? Subsystem (NAS / MAS & iSCCI / FC connected > something / DAS)? > > Do you have logging and indexes enables during load? > > In my book you need as many parallelism on input AND output, as > until your I/O is saturated. This saturation point should be > much higher than ~3300 recs/sec. The values you give also shows, > that what you did does not scale too well: you went 8 way parallel, > but only got 6 fold improvement. > > One way to go there, and a very cheap one, might be to find a way to > split input into 8 parts (like using mod or somesuch) and use 8 > prozesses started at the same time. Even Slowaris 10 should be able to > read only once physically and 7 of the 8 brothers should read from > memory buffers, which normally on plastic hardware is a lil faster, > like 200 times if the physical read is not from SSD and even on > big, expensive *nix harware it is like 120 times faster. > > HPL is always worth a try, as using HPL it is possible to have > parallel writes together with parallel reading. > > I believe you must know what you hardware + OS can do and then > using IDS over all the other database systems out there makes me > always using 95% of this possible performance. > > dic_k > -- > Richard Kofler > SOLID STATE EDV > Dienstleistungen GmbH > Vienna/Austria/Europe Thank you for your response. My page size is 2K and I have 4 records per page (rowsize = 485). Expanding the page size is certainly an option. The disk subsystem is direct attached Fibre channel, RAID 10, 10K / 146 GB devices. Logging and indexes are enabled at execution since the system is expected to be fully available and is a mixed OLTP/DSS environment. I have not yet tried to compress the data pages and gauge the result, but I will. I also have the disadvantage at the moment that the readers and writers access the same LUNs/Dbspaces. The disks are not being completely saturated, in fact far from it at about 30% busy. Some of this may be related to a complex query that feeds the writer(s).
Honestly, in the data load system you are contemplating, I don't think that using 8 or more separate copies of the same task, written once to read data from a pipe or message queue (queues are a bit faster on some UNIXes including Solaris) or using semaphores and shared memory will be any slower than a single multi-threaded task and may even be faster given the simpler coordination between the processes versus threads which have to share memory. Using message queues may be simplest and fastest as you can have all of the consumers/inserters reading from the same queue so that whichever of them is ready to read will pick up the next message, but you'd have to write your own producer to feed the file's records to the queue. Art Art S. Kagel Advanced DataTools (www.advancedatatools.com) IIUG Board of Directors (art@iiug.org) See you at the 2010 IIUG Informix Conference April 25-28, 2010 Overland Park (Kansas City), KS www.iiug.org/conf Disclaimer: Please keep in mind that my own opinions are my own opinions and do not reflect on my employer, Advanced DataTools, the IIUG, nor any other organization with which I am associated either explicitly, implicitly, or by inference. Neither do those opinions reflect those of other individuals affiliated with any entity with which I am affiliated nor those of the entities themselves. On Sat, Feb 13, 2010 at 4:08 PM, <the_omegamon@yahoo.com> wrote: > Greetings All, > > I am looking for an example of functional ESQL/C source code that > demonstrates multi-threading. I have found the thread_safe example > in the ESQL/C programmers guide, but I have not had any luck getting > it to compile. I did not find an example in $INFORMIXDIR/demo/esqlc > or anywhere on the net. > > The problem I am trying to solve is an INSERT bottleneck. My table > is > partitioned (fragmented) across 8 dbspaces with multiple indexes and > foreign keys containing ~ 800,000,000 rows. A serial load will > INSERT > 2,000,000 record in ~ 1 hour while a parallel load of 8 processes > takes > about 10 minutes. Suggestions in this group seem to point toward > having > a reader provide the data to a set of INSERT consumers. Given the > circumstances, this would probably not be a good case for utilizing > HPL > as it would be be beneficial for us to keep everything consolidated > within ESQL/C. > > I am running IDS 10.00 and 11.50 on Solaris 10 with CSDK 3.50. > > Thanks in advance for you insight, > _______________________________________________ > Informix-list mailing list > Informix-list@iiug.org > http://www.iiug.org/mailman/listinfo/informix-list >
Related threads
- the longer you surf, the MORE $$$ you earn !!
- Store procedure
- emulation for Vt100
- extent size questions again ...