Re: Help -- OnLine server refuses connections
Posted in 1997
In article <336FAE78.7985@surf.com>, David Anderson <anderson@surf.com>
writes
>I'm running INFORMIX-OnLine Version 7.20.UC3
>
>Once or twice a day (possible correlated with heavy load)
>the server abruptly stops accepting connections.
>More details follow.
>Any tips or suggestions would be greatly appreciated.
>Please reply by email. Thanks!
>
>David Anderson
>anderson@tunes.com
>------------------
>When the server is in this state,
>an ESQL client hangs with the following stack:
>
>Program received signal SIGINT, Interrupt.
>0x3ff83058a18 in usleep_thread ()
>(gdb) where
>#0 0x3ff83058a18 in usleep_thread ()
>#1 0x3ff830aadf4 in __sleep ()
>#2 0x1200ea9c8 in ifxOS_sleep () at osmutex.c:705
>#3 0x1200d2210 in connshm () at shm_fe.c:719
This is connecting via shared memory (connshm)..
>#4 0x1200cde74 in tlConnect () at tl.c:152
>#5 0x1200d9168 in slSQIreq () at asfslsqi.c:857
>#6 0x1200cca38 in pfConReq () at asfpfsqi.c:2413
>#7 0x1200c83b8 in cmReqSync () at cm.c:2222
>#8 0x1200c6d28 in cmConReq () at cm.c:541
>#9 0x1200bfab8 in ascRequest () at al.c:419
>#10 0x1200bd55c in ASF_Call () at asfapi.c:617
>#11 0x1200a9828 in asf_connect () at iqconnct.c:907
>#12 0x1200aa584 in _iqconnect () at iqconnct.c:1364
>#13 0x1200ac1f4 in _sqs_ () at iqconnct.c:2508
>#14 0x12009d890 in _iqdbase () at iqsimple.c:482
>-----------
>And here's a typical output of onstat when the server is in this state:
>
>Userthreads
>address flags sessid user tty wait
>tout locks nreads nwrites
>20033e028 ---P--D 1 informix - 0
>0 0 69 384
^
+---- normal process a few reads
>20033e658 ---P--F 0 informix - 0
>0 0 0 3532
^
---------normal process doing query which returns a lot of
rows (e.g. ~1000 via an index)
>20033ec88 ---P--B 8 informix - 0
>0 0 0 0
>20033f2b8 ---P--D 66 informix - 0
>0 0 0 0
^
-------'quit' processes.
>20033ff18 Y--P--- 155 anderson ttyp6 2006f2340
>0 1 19653 96
^
+======== **** WHAT **** LOOKS LIKE A QUERY WHERE IT
SCANS ONE TABLE BECAUSE
a) an index is missing
b) someone forgot to run update staistics.
>200340b78 Y--P--- 176 anderson ttyp8 2006a31a0
>0 1 229609 0
^
+======== **** WHOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOA ****
LOOKS LIKE A QUERY WHERE IT SCANS ONE MASSIVE TABLE
OR A SMALL TABLE >50 TIMES.
IS A COORELATED SUBQUERY OR HAS A JOIN MISSING
I.E. a CROSS PRODUCT OF 2 TABLES.
Either
a) an index is missing
b) someone forgot to run update staistics.
c) some idiot missed a join condition off the query
e.g. select col1
from tab1,tab2,tab3
where tab1.col1 = tab2.col
NOTE join from tab2 to tab3 missing!!
I once did this and the query ran for 3 hours
and did +200,000 to 300,000 reads. (It scans a
40,000 row table over and over and over...
after adding the missing join the query took
<2 seconds!!
Get the programmer to add
- change to users home directory
- SET EXPLAIN ON
the application code.
Then wait for the same situation to happen and look in the
sqexplain.out file in offending users home directory.
This should give the programmers more info.
> 6 active, 128 total, 38 maximum concurrent
>
>Profile
>dskreads pagreads bufreads %cached dskwrits pagwrits bufwrits %cached
>163276 436027 3919489 95.83 5562 31648 8938 37.77
>
>isamtot open start read write rewrite delete commit
>rollbk
>3341678 236085 456371 1177677 757 98 1 1042
>0
>
>ovlock ovuserthread ovbuff usercpu syscpu numckpts flushes
>0 0 0 1441.10 96.92 17 34
>
>bufwaits lokwaits lockreqs deadlks dltouts ckpwaits compress seqscans
>50869 2 2027917 0 0 8 582 226
>
>ixda-RA idx-RA da-RA RA-pgsused lchwaits
>85118 1268 3808 86863 8
--
David Williams