count/insert vs insert
Posted in 2011
Topics: Error Codes & Troubleshooting, Connectivity: ESQL/C, 4GL & Embedded SQL, Logging & Checkpoints, Platform-Specific Issues
Hi,
I already see here some recommendations when execute an insert , if
should check before or not if the record exists or if better check if
occur some constraint error (unique index/PK).
As far I remember , the better way on most of the cases, is try insert
and than check if was occur successfully.
I executed a few tests and have 3 doubts. Please follow the tests bellow.
(ifx 11.50 fc8 aix , 4gl 7.32)
After write a little code using 4GL and executing 5 times :
1) execute just an insert for a record what don't exists on the table.
(the insert occur successful)
sesid 7396
2) execute a "select first 1 1 from table" + if sqlca.sqlcode =
notfound, than execute the insert (the insert occur successful because
the record don't exists)
sesid 7409
3) execute just an insert for a record where the unique key (not PK)
already exists (the program abend with error -239)
sesid 7412
4) execute a "select first 1 1 from table" + if sqlca.sqlcode =
notfound, so the insert wasn't executed because the record already exists
sesid 7421
5) execute just an insert for a record where the PK already exists (the
program abend with error -268)
sesid 7470
(I add a PK over the UK)
The result what I got appear to confirm, using just the insert is better
and check later if occur some error.
And appear be better if use PK instead only UK.
This outputs I gather from sys* , saving at sysdbclose...
sesid lockreqs lockwts logrecs maxlog_kb isreads iswrites isrewrites
isdeletes iscommits seqscans total_sorts cpu_time
7396 928 0 6 0.56 664 51 0 0 1 3 0
0.73029
7409 867 0 6 0.56 636 51 0 0 1 3 0
0.95809
7412 854 0 8 0.61 644 52 0 0 0 3 0
0.61418
7421 888 0 0 0.00 660 51 0 0 0 3 0
0.8001
7470 1042 0 10 0.74 718 50 0 0 0 3 0
0.07048
The fields what I pay attention: logrecs, maxlog_kb, iscommits, cpu_time
sesid logrecs maxlog_kb iscommits cpu_time
7396 6 0.56 1 0.73029
7409 6 0.56 1 0.95809
7412 8 0.61 0 0.61418
7421 0 0.00 0 0.8001
7470 10 0.74 0 0.07048
What caught my attention is the use of logical log records and the
little cpu_time when the table have a PK instead only the UK.
My Doubts
1)On the session 7412 , where run the insert and check for errors
after(UK violated). I strange the grown of logical log record.
Checking the logs (with onlog) I see the engine adding the 2 others
index (ADDITEM) before add the UK index and then abend with error.
Should be better the engine always insert first the UK constraints , to
avoid the ADDITEM and the rollback of each one if occur a violation on
the constraint?
2) On this situation (sesid 7412), where the CPU consumption is instead
check before if the record exists (sesid 7421) and having more logical
log records... this way still better way to work ? considering the using
of logical log x cpu_time (looking at this isolate context, test-by-test)
3) On sesid 7470, where I add a PK to table, the engine appear have the
same behave of UK , adding each index of the table before the UK/PK
(onlog/ADDITEM) and then abending with -268 error. How could be the
engine consume 10 times less CPU of the database!???
Cesar
If you will be updating the row if it exists and inserting it if it does
not, the best, ie cheapest, way to approach it is to try the update first.
If zero rows are updated (check the # of rows affected value in
sqlca.sqlerrd[3] in 4GL or sqlca.sqlerrd[2] in ESQL/C since it is not an
error to update zero rows) then perform the insert. Failed updates are
expensive because, as you have surmised, the uniqueness constraints (UK, &
PK as well as unique indexes) are checked only after the row has been
inserted on a page and all of the indexes have been updated. At that point
a violation will cause all of that to be rolled back.
The rule of thumb I use with much success is:
- if fewer than 70% of the attempted rows will end up being inserts than
trying the update first will produce the most efficient code.
- If more than 70% of the attempts will result in an insert then try the
insert first because the lower incidence of the higher cost balances out.
Trying a select first to see if the row is there is the least efficient way
to do this in any case.
Art
Art S. Kagel
Advanced DataTools (www.advancedatatools.com)
Blog: http://informix-myview.blogspot.com/
Disclaimer: Please keep in mind that my own opinions are my own opinions and
do not reflect on my employer, Advanced DataTools, the IIUG, nor any other
organization with which I am associated either explicitly, implicitly, or by
inference. Neither do those opinions reflect those of other individuals
affiliated with any entity with which I am affiliated nor those of the
entities themselves.
On Wed, Sep 21, 2011 at 4:49 PM, Cesar Inacio Martins <
cesar_inacio_martins@yahoo.com.br> wrote:
> Hi,
>
> I already see here some recommendations when execute an insert , if
> should check before or not if the record exists or if better check if
> occur some constraint error (unique index/PK).
>
> As far I remember , the better way on most of the cases, is try insert
> and than check if was occur successfully.
> I executed a few tests and have 3 doubts. Please follow the tests bellow.
> (ifx 11.50 fc8 aix , 4gl 7.32)
>
> After write a little code using 4GL and executing 5 times :
> 1) execute just an insert for a record what don't exists on the table.
> (the insert occur successful)
>
> sesid 7396
> 2) execute a "select first 1 1 from table" + if sqlca.sqlcode =
> notfound, than execute the insert (the insert occur successful because
> the record don't exists)
>
> sesid 7409
> 3) execute just an insert for a record where the unique key (not PK)
> already exists (the program abend with error -239)
>
> sesid 7412
> 4) execute a "select first 1 1 from table" + if sqlca.sqlcode =
> notfound, so the insert wasn't executed because the record already exists
>
> sesid 7421
> 5) execute just an insert for a record where the PK already exists (the
> program abend with error -268)
>
> sesid 7470
>
> (I add a PK over the UK)
>
> The result what I got appear to confirm, using just the insert is better
> and check later if occur some error.
> And appear be better if use PK instead only UK.
> This outputs I gather from sys* , saving at sysdbclose...
>
> sesid lockreqs lockwts logrecs maxlog_kb isreads iswrites isrewrites
> isdeletes iscommits seqscans total_sorts cpu_time
> 7396 928 0 6 0.56 664 51 0 0 1 3 0
> 0.73029
> 7409 867 0 6 0.56 636 51 0 0 1 3 0
> 0.95809
> 7412 854 0 8 0.61 644 52 0 0 0 3 0
> 0.61418
> 7421 888 0 0 0.00 660 51 0 0 0 3 0
> 0.8001
> 7470 1042 0 10 0.74 718 50 0 0 0 3 0
> 0.07048
>
> The fields what I pay attention: logrecs, maxlog_kb, iscommits, cpu_time
>
> sesid logrecs maxlog_kb iscommits cpu_time
> 7396 6 0.56 1 0.73029
> 7409 6 0.56 1 0.95809
> 7412 8 0.61 0 0.61418
> 7421 0 0.00 0 0.8001
> 7470 10 0.74 0 0.07048
>
> What caught my attention is the use of logical log records and the
> little cpu_time when the table have a PK instead only the UK.
>
> My Doubts
> 1)On the session 7412 , where run the insert and check for errors
> after(UK violated). I strange the grown of logical log record.
> Checking the logs (with onlog) I see the engine adding the 2 others
> index (ADDITEM) before add the UK index and then abend with error.
> Should be better the engine always insert first the UK constraints , to
> avoid the ADDITEM and the rollback of each one if occur a violation on
> the constraint?
>
> 2) On this situation (sesid 7412), where the CPU consumption is instead
> check before if the record exists (sesid 7421) and having more logical
> log records... this way still better way to work ? considering the using
> of logical log x cpu_time (looking at this isolate context, test-by-test)
>
> 3) On sesid 7470, where I add a PK to table, the engine appear have the
> same behave of UK , adding each index of the table before the UK/PK
> (onlog/ADDITEM) and then abending with -268 error. How could be the
> engine consume 10 times less CPU of the database!???
>
> Cesar
>
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
>
--90e6ba1efeaeabd72a04ad7b61b8
If the rows are already in a table then you can
consider using the MERGE statement.
The use of the MERE statement does not
accept host variables, but you can
EXECUTE IMMEDIATE with constants.
See example below.
create table t1( c1 char(20), c2 char(20), c3 char(20) );MERGE
INTO t1 AS t
USING ( select "JOHN", "miller" ,"III" FROM sysmaster:sysdual) as s
(col1,col2,col3)
ON t.c1 = s.col1
WHEN MATCHED THEN UPDATE
SET t.c2 = s.col2
WHEN NOT MATCHED THEN INSERT
(t.c1, t.c2, t.c3)
VALUES
(s.col1, s.col2,s.col3);
John F. Miller III
STSM, Embedability Architect
miller3@us.ibm.com
503-578-5645
IBM Informix Dynamic Server (IDS)
ids-bounces@iiug.org wrote on 09/21/2011 03:49:01 PM:
> From: "Art Kagel" <art.kagel@gmail.com>
> To: ids@iiug.org
> Date: 09/21/2011 03:52 PM
> Subject: Re: count/insert vs insert [24990]
> Sent by: ids-bounces@iiug.org
>
> If you will be updating the row if it exists and inserting it if it does
> not, the best, ie cheapest, way to approach it is to try the update
first.
> If zero rows are updated (check the # of rows affected value in
> sqlca.sqlerrd[3] in 4GL or sqlca.sqlerrd[2] in ESQL/C since it is not an
> error to update zero rows) then perform the insert. Failed updates are
> expensive because, as you have surmised, the uniqueness constraints (UK,
&
> PK as well as unique indexes) are checked only after the row has been
> inserted on a page and all of the indexes have been updated. At that
point
> a violation will cause all of that to be rolled back.
>
> The rule of thumb I use with much success is:
>
> - if fewer than 70% of the attempted rows will end up being inserts than
>
> trying the update first will produce the most efficient code.
>
> - If more than 70% of the attempts will result in an insert then try the
>
> insert first because the lower incidence of the higher cost balances out.
>
> Trying a select first to see if the row is there is the least efficient
way
> to do this in any case.
>
> Art
>
> Art S. Kagel
> Advanced DataTools (www.advancedatatools.com)
> Blog: http://informix-myview.blogspot.com/
>
> Disclaimer: Please keep in mind that my own opinions are my own opinions
and
> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
other
> organization with which I am associated either explicitly, implicitly, or
by
> inference. Neither do those opinions reflect those of other individuals
> affiliated with any entity with which I am affiliated nor those of the
> entities themselves.
>
> On Wed, Sep 21, 2011 at 4:49 PM, Cesar Inacio Martins <
> cesar_inacio_martins@yahoo.com.br> wrote:
>
> > Hi,
> >
> > I already see here some recommendations when execute an insert , if
> > should check before or not if the record exists or if better check if
> > occur some constraint error (unique index/PK).
> >
> > As far I remember , the better way on most of the cases, is try insert
> > and than check if was occur successfully.
> > I executed a few tests and have 3 doubts. Please follow the tests
bellow.
> > (ifx 11.50 fc8 aix , 4gl 7.32)
> >
> > After write a little code using 4GL and executing 5 times :
> > 1) execute just an insert for a record what don't exists on the table.
> > (the insert occur successful)
> >
> > sesid 7396
> > 2) execute a "select first 1 1 from table" + if sqlca.sqlcode =
> > notfound, than execute the insert (the insert occur successful because
> > the record don't exists)
> >
> > sesid 7409
> > 3) execute just an insert for a record where the unique key (not PK)
> > already exists (the program abend with error -239)
> >
> > sesid 7412
> > 4) execute a "select first 1 1 from table" + if sqlca.sqlcode =
> > notfound, so the insert wasn't executed because the record already
exists
> >
> > sesid 7421
> > 5) execute just an insert for a record where the PK already exists (the
> > program abend with error -268)
> >
> > sesid 7470
> >
> > (I add a PK over the UK)
> >
> > The result what I got appear to confirm, using just the insert is
better
> > and check later if occur some error.
> > And appear be better if use PK instead only UK.
> > This outputs I gather from sys* , saving at sysdbclose...
> >
> > sesid lockreqs lockwts logrecs maxlog_kb isreads iswrites isrewrites
> > isdeletes iscommits seqscans total_sorts cpu_time
> > 7396 928 0 6 0.56 664 51 0 0 1 3 0
> > 0.73029
> > 7409 867 0 6 0.56 636 51 0 0 1 3 0
> > 0.95809
> > 7412 854 0 8 0.61 644 52 0 0 0 3 0
> > 0.61418
> > 7421 888 0 0 0.00 660 51 0 0 0 3 0
> > 0.8001
> > 7470 1042 0 10 0.74 718 50 0 0 0 3 0
> > 0.07048
> >
> > The fields what I pay attention: logrecs, maxlog_kb, iscommits,
cpu_time
> >
> > sesid logrecs maxlog_kb iscommits cpu_time
> > 7396 6 0.56 1 0.73029
> > 7409 6 0.56 1 0.95809
> > 7412 8 0.61 0 0.61418
> > 7421 0 0.00 0 0.8001
> > 7470 10 0.74 0 0.07048
> >
> > What caught my attention is the use of logical log records and the
> > little cpu_time when the table have a PK instead only the UK.
> >
> > My Doubts
> > 1)On the session 7412 , where run the insert and check for errors
> > after(UK violated). I strange the grown of logical log record.
> > Checking the logs (with onlog) I see the engine adding the 2 others
> > index (ADDITEM) before add the UK index and then abend with error.
> > Should be better the engine always insert first the UK constraints , to
> > avoid the ADDITEM and the rollback of each one if occur a violation on
> > the constraint?
> >
> > 2) On this situation (sesid 7412), where the CPU consumption is instead
> > check before if the record exists (sesid 7421) and having more logical
> > log records... this way still better way to work ? considering the
using
> > of logical log x cpu_time (looking at this isolate context,
test-by-test)
> >
> > 3) On sesid 7470, where I add a PK to table, the engine appear have the
> > same behave of UK , adding each index of the table before the UK/PK
> > (onlog/ADDITEM) and then abending with -268 error. How could be the
> > engine consume 10 times less CPU of the database!???
> >
> > Cesar
> >
> >
> >
> >
>
*******************************************************************************
> > Forum Note: Use "Reply" to post a response in the discussion forum.
> >
> >
>
> --90e6ba1efeaeabd72a04ad7b61b8
>
>
>
*******************************************************************************
> Forum Note: Use "Reply" to post a response in the discussion forum.
>
Hi John, Hi Art,
At this moment I considering only insert , not update over the data...
The situation is : if not exists... insert , any else, go to next case...
On 21/9/2011 20:19, John Miller iii wrote:
> If the rows are already in a table then you can
> consider using the MERGE statement.
>
> The use of the MERE statement does not
> accept host variables, but you can
> EXECUTE IMMEDIATE with constants.
> See example below.
>
> create table t1( c1 char(20), c2 char(20), c3 char(20) );> MERGE
> INTO t1 AS t
> USING ( select "JOHN", "miller" ,"III" FROM sysmaster:sysdual) as s
> (col1,col2,col3)
> ON t.c1 = s.col1
> WHEN MATCHED THEN UPDATE
>
> SET t.c2 = s.col2
> WHEN NOT MATCHED THEN INSERT
>
> (t.c1, t.c2, t.c3)
>
> VALUES
>
> (s.col1, s.col2,s.col3);
>
> John F. Miller III
> STSM, Embedability Architect
> miller3@us.ibm.com
> 503-578-5645
> IBM Informix Dynamic Server (IDS)
>
> ids-bounces@iiug.org wrote on 09/21/2011 03:49:01 PM:
>
>> From: "Art Kagel"<art.kagel@gmail.com>
>> To: ids@iiug.org
>> Date: 09/21/2011 03:52 PM
>> Subject: Re: count/insert vs insert [24990]
>> Sent by: ids-bounces@iiug.org
>>
>> If you will be updating the row if it exists and inserting it if it does
>> not, the best, ie cheapest, way to approach it is to try the update
> first.
>> If zero rows are updated (check the # of rows affected value in
>> sqlca.sqlerrd[3] in 4GL or sqlca.sqlerrd[2] in ESQL/C since it is not an
>> error to update zero rows) then perform the insert. Failed updates are
>> expensive because, as you have surmised, the uniqueness constraints (UK,
> &
>> PK as well as unique indexes) are checked only after the row has been
>> inserted on a page and all of the indexes have been updated. At that
> point
>> a violation will cause all of that to be rolled back.
>>
>> The rule of thumb I use with much success is:
>>
>> - if fewer than 70% of the attempted rows will end up being inserts than
>>
>> trying the update first will produce the most efficient code.
>>
>> - If more than 70% of the attempts will result in an insert then try the
>>
>> insert first because the lower incidence of the higher cost balances out.
>> Trying a select first to see if the row is there is the least efficient
> way
>> to do this in any case.
>>
>> Art
>>
>> Art S. Kagel
>> Advanced DataTools (www.advancedatatools.com)
>> Blog: http://informix-myview.blogspot.com/
>>
>> Disclaimer: Please keep in mind that my own opinions are my own opinions
> and
>> do not reflect on my employer, Advanced DataTools, the IIUG, nor any
> other
>> organization with which I am associated either explicitly, implicitly, or
> by
>> inference. Neither do those opinions reflect those of other individuals
>> affiliated with any entity with which I am affiliated nor those of the
>> entities themselves.
>>
>> On Wed, Sep 21, 2011 at 4:49 PM, Cesar Inacio Martins<
>> cesar_inacio_martins@yahoo.com.br> wrote:
>>
>>> Hi,
>>>
>>> I already see here some recommendations when execute an insert , if
>>> should check before or not if the record exists or if better check if
>>> occur some constraint error (unique index/PK).
>>>
>>> As far I remember , the better way on most of the cases, is try insert
>>> and than check if was occur successfully.
>>> I executed a few tests and have 3 doubts. Please follow the tests
> bellow.
>>> (ifx 11.50 fc8 aix , 4gl 7.32)
>>>
>>> After write a little code using 4GL and executing 5 times :
>>> 1) execute just an insert for a record what don't exists on the table.
>>> (the insert occur successful)
>>>
>>> sesid 7396
>>> 2) execute a "select first 1 1 from table" + if sqlca.sqlcode =
>>> notfound, than execute the insert (the insert occur successful because
>>> the record don't exists)
>>>
>>> sesid 7409
>>> 3) execute just an insert for a record where the unique key (not PK)
>>> already exists (the program abend with error -239)
>>>
>>> sesid 7412
>>> 4) execute a "select first 1 1 from table" + if sqlca.sqlcode =
>>> notfound, so the insert wasn't executed because the record already
> exists
>>> sesid 7421
>>> 5) execute just an insert for a record where the PK already exists (the
>>> program abend with error -268)
>>>
>>> sesid 7470
>>>
>>> (I add a PK over the UK)
>>>
>>> The result what I got appear to confirm, using just the insert is
> better
>>> and check later if occur some error.
>>> And appear be better if use PK instead only UK.
>>> This outputs I gather from sys* , saving at sysdbclose...
>>>
>>> sesid lockreqs lockwts logrecs maxlog_kb isreads iswrites isrewrites
>>> isdeletes iscommits seqscans total_sorts cpu_time
>>> 7396 928 0 6 0.56 664 51 0 0 1 3 0
>>> 0.73029
>>> 7409 867 0 6 0.56 636 51 0 0 1 3 0
>>> 0.95809
>>> 7412 854 0 8 0.61 644 52 0 0 0 3 0
>>> 0.61418
>>> 7421 888 0 0 0.00 660 51 0 0 0 3 0
>>> 0.8001
>>> 7470 1042 0 10 0.74 718 50 0 0 0 3 0
>>> 0.07048
>>>
>>> The fields what I pay attention: logrecs, maxlog_kb, iscommits,
> cpu_time
>>> sesid logrecs maxlog_kb iscommits cpu_time
>>> 7396 6 0.56 1 0.73029
>>> 7409 6 0.56 1 0.95809
>>> 7412 8 0.61 0 0.61418
>>> 7421 0 0.00 0 0.8001
>>> 7470 10 0.74 0 0.07048
>>>
>>> What caught my attention is the use of logical log records and the
>>> little cpu_time when the table have a PK instead only the UK.
>>>
>>> My Doubts
>>> 1)On the session 7412 , where run the insert and check for errors
>>> after(UK violated). I strange the grown of logical log record.
>>> Checking the logs (with onlog) I see the engine adding the 2 others
>>> index (ADDITEM) before add the UK index and then abend with error.
>>> Should be better the engine always insert first the UK constraints , to
>>> avoid the ADDITEM and the rollback of each one if occur a violation on
>>> the constraint?
>>>
>>> 2) On this situation (sesid 7412), where the CPU consumption is instead
>>> check before if the record exists (sesid 7421) and having more logical
>>> log records... this way still better way to work ? considering the
> using
>>> of logical log x cpu_time (looking at this isolate context,
> test-by-test)
>>> 3) On sesid 7470, where I add a PK to table, the engine appear have the
>>> same behave of UK , adding each index of the table before the UK/PK
>>> (onlog/ADDITEM) and then abending with -268 error. How could be the
>>> engine consume 10 times less CPU of the database!???
>>>
>>> Cesar
>>>
>>>
>>>
>>>
>
*******************************************************************************
>
>>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>>
>>>
>> --90e6ba1efeaeabd72a04ad7b61b8
>>
>>
>>
>
*******************************************************************************
>
>> Forum Note: Use "Reply" to post a response in the discussion forum.
>>
>
>
*******************************************************************************
> Forum Note: Use "Reply@@
Related threads
- Some SQL errors...
- Is there any way to log failed insert row because of unique index violation - Informix
- dbexport miracle