Performance problem - Informix,cgi-webdriver,HP-UX
Posted in 1999
Topics: Performance & Tuning, Server Administration, Platform-Specific Issues, Third-Party Tools & Monitoring
Hi gurus,
Sorry for the long mail.
We are running a database backed intranet website on a HP server running HP-UX
10.20,
Netscape Enterprize 3.0, Informix 7.3 & WebBlade 3.0 using cgi implementation of
webdriver.
The server has 2 X 180MHz cpus and 3 GB of RAM.
Users access the website remotely with web browsers running on their PCs.
Users are frequently experiencing problems with the website; response is always
slow
and in the worst case they cannot access the webdriver driven part of the site
ie.
http://IP address/sitename/cgi-bin/webdriver at all.
But static content on the same web server eg. http://IP address/static.html can
be accessed.
At the times when this problem is experienced all relevant services are still
running:
- HP-UX is running (telnet sessions OK, responds to unix commands OK)
- Netscape Enterprize Server is running (users can access static HTML
content)
- Informix is online
- WebBlade webdriver is running.
But users cannot access dynamic content via Informix WebBlade cgi webdriver.
They cannot
access/ query the Informix database.
One of the symptoms of the problem is that the no. of instances of webdriver
jumps very
suddenly from < 5 to > 20 and the total no. of processes jumps from < 220 to >
260 and
both cpus hit approx. 100%.
This happens at least once every day !
My quick and dirty fix is to temporarily rename webdriver, kill the oldest
webdriver processes,
wait till CPU load drops and then restore webdriver.
I am just a humble systems admin guy not a DBA and I can't figure out what I can
do to fix this
problem. Maybe I could resign ;-)
Any advice on (a) how to analyze the problem and
(b) how to fix it/ improve performance etc.
would be greatly appreciated.
Best regards,
Michael Martin
michael_martin@komatsu.co.jp
--------------------------------------------------------------------
The following is a snapshot from GlancePlus (a HP monitoring tool) at a
time when the website had frozen.
My reading of this is:
- two informix processes running at relatively low pri of 233 blocked on PRI
but munching up cpu (these are big.. RSS includes everything related to a
process that uses RAM eg. data, stack, text and shared memory segments
- lots of webdriver processes blocked on SOCKT (waiting for socket operations to
complete) and all of them have the same PPID
- ns-httpd blocked on SLEEP ie. just waiting
My thought at the time was that Informix wasn't replying fast enough to
webdriver
and that may be part of the problem but I guess I need to figure out exactly
what
socket operations the webdriver processes are waiting on. And maybe more
importantly how does webdriver respond to being blocked on SOCKT... does it
keep madly spawning new processes ?
Could this socket stuff mean that their is a communications problem on top
of a possible database design/SQL efficiency problem ?
B3692A GlancePlus C.02.15.00 15:54:06 OYAPSV01 9000/889 Current Avg High
--------------------------------------------------------------------------------
Cpu Util S SARU U |100% 39% 100%
Disk Util F F | 6% 4% 100%
Mem Util S SU UB B | 48% 43% 48%
Swap Util U UR R | 21% 18% 21%
--------------------------------------------------------------------------------
PROCESS LIST Users= 3
User CPU Util Cum Disk Block
Process Name PID PPID Pri Name ( 200% max) CPU IO Rate RSS On
--------------------------------------------------------------------------------
oninit 1974 1972 233 informix 95.4/ 7.3 4668.8 0.0/ 0.7 619.5mb PRI
oninit 1972 1971 233 informix 89.9/12.3 7901.4 0.0/ 1.1 619.5mb PRI
ns-httpd 1397 1396 154 www 2.7/ 3.7 12264.6 0.0/ 0.0 27.8mb SLEEP
scopeux 1414 1349 127 root 2.5/ 0.2 627.9 6.9/ 0.5 8.2mb SLEEP
ns-httpd 1748 1747 154 www 2.1/ 2.0 1263.9 0.0/ 0.0 26.3mb SLEEP
oninit 1752 1726 156 infgus 0.0/ 0.0 5.6 0.0/ 0.0 119.4mb SEM
webdriver 10326 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
oninit 1726 1710 168 infgus 0.0/ 0.0 17.9 0.0/ 0.0 119.4mb SLEEP
oninit 1727 1710 154 infgus 0.0/ 0.0 29.7 0.0/ 0.0 122.2mb SLEEP
oninit 1998 1973 156 informix 0.0/ 0.0 12.6 0.0/ 0.0 616.5mb SEM
webdriver 17255 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.6mb SOCKT
webdriver 17730 1397 154 www 0.0/ 0.3 0.3 0.0/ 0.0 2.6mb SOCKT
oninit 1985 1973 156 informix 0.0/ 0.0 7.5 0.0/ 0.0 616.4mb SEM
webdriver 17254 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.7mb SOCKT
oninit 1999 1973 154 informix 0.0/ 0.1 53.7 0.0/ 0.0 616.5mb SLEEP
webdriver 17510 1397 154 www 0.0/ 0.3 0.3 0.0/ 0.0 2.6mb SOCKT
webdriver 10327 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
webdriver 17253 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.7mb SOCKT
webdriver 17769 1397 154 www 0.0/ 0.4 0.3 0.0/ 0.0 2.6mb SOCKT
webdriver 10344 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
oninit 1987 1973 156 informix 0.0/ 0.0 15.4 0.0/ 0.0 616.5mb SEM
oninit 2000 1973 154 informix 0.0/ 0.1 35.9 0.0/ 0.0 616.5mb SLEEP
webdriver 18244 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 3.0mb SOCKT
oninit 1988 1973 156 informix 0.0/ 0.0 9.0 0.0/ 0.0 616.5mb SEM
uxwdog 1396 1395 168 www 0.0/ 0.0 0.0 0.0/ 0.0 3.9mb SLEEP
webdriver 17939 1397 154 www 0.0/ 1.7 0.3 0.0/ 0.0 2.6mb SOCKT
OYAPSV01#>sar -v 4 2
HP-UX OYAPSV01 B.10.20 E 9000/889 10/28/99
15:53:05 text-sz ov proc-sz ov inod-sz ov file-sz ov
15:53:09 N/A N/A 239/1012 0 1212/1212 0 852/9380 0
15:53:13 N/A N/A 240/1012 0 1212/1212 0 860/9380 0
OYAPSV01#>ps -ef | grep webd | sort -nr -k 2
www 18950 1397 0 18:53:30 ? 0:00 webdriver
www 18245 1397 0 18:50:55 ? 0:00 webdriver
* www 18244 1397 0 18:50:54 ? 0:00 webdriver
www 17978 1397 0 18:49:46 ? 0:00 webdriver
root 17868 28714 0 15:53:38 ttyp3 0:00 grep webd
www 17787 1397 0 15:53:15 ? 0:00 webdriver
www 17773 1397 0 15:53:05 ? 0:00 webdriver
www 17772 1397 0 15:53:05 ? 0:00 webdriver
www 17771 1397 0 15:53:05 ? 0:00 webdriver
* www 17769 1397 0 15:53:05 ? 0:00 webdriver
* www 17730 1397 0 15:52:58 ? 0:00 webdriver
www 17513 1397 0 15:52:32 ? 0:00 webdriver
www 17512 1397 0 15:52:32 ? 0:00 webdriver
* www 17510 1397 0 15:52:29 ? 0:00 webdriver
www 17256 1397 0 15:51:38 ? 0:00 webdriver
* www 17255 1397 0 15:51:38 ? 0:00 webdriver
* www 17254 1397 0 15:51:38 ?
In article <7vbvls$6j3$1@news.xmission.com>, "michael martin" <michael_martin@komatsu.co.jp> wrote: > > Hi gurus, > > Sorry for the long mail. > > We are running a database backed intranet website on a HP server running HP-UX > 10.20, > Netscape Enterprize 3.0, Informix 7.3 & WebBlade 3.0 using cgi implementation of > webdriver. > > Hmm... sounds familiar. My 2 cents: Check to make sure that you do not have the webexplode log turned on. This slows down the system quite a bit. Also check if the "set explain on" statement is set in the configuration file for the web server. Find out where the netscape is s/w is installed and there you should have the config files under the congig directory. The "set explain on" statement could be turned on with the statement as shown below: MI_WEBINITIALSQL SET LOCK MODE TO WAIT 60; SET EXPLAIN ON; We have a NSAPI based web server 3.31 running on HP-UX 10.20 with informix 9.13. Just worth checking it.... thanks Ram S Sent via Deja.com http://www.deja.com/ Before you buy.
Sounds like the engine needs serious tuning and you MAY also need a newer
bigger faster server box. Post the following and we will try to help out
(BTW you have IDS 9.14 not 7.3):
ONCONFIG file
sqlhosts file
onstat -g glo
onstat -g iov
onstat -g ioq
onstat -d
OS version
Amount of memory & swap
some description of disk farm layout
We can try to squeeze out as much performance from this system as you can
but bottom line it looks over burdened. Look into upgrading the CPUs to
newer faster HP-PA/RISC CPUs which are now available faster than 400MHZ.
Art S. Kagel
michael martin wrote:
>
> Hi gurus,
>
> Sorry for the long mail.
>
> We are running a database backed intranet website on a HP server running HP-UX
> 10.20,
> Netscape Enterprize 3.0, Informix 7.3 & WebBlade 3.0 using cgi implementation of
> webdriver.
>
> The server has 2 X 180MHz cpus and 3 GB of RAM.
>
> Users access the website remotely with web browsers running on their PCs.
>
> Users are frequently experiencing problems with the website; response is always
> slow
> and in the worst case they cannot access the webdriver driven part of the site
> ie.
> http://IP address/sitename/cgi-bin/webdriver at all.
>
> But static content on the same web server eg. http://IP address/static.html can
> be accessed.
>
> At the times when this problem is experienced all relevant services are still
> running:
> - HP-UX is running (telnet sessions OK, responds to unix commands OK)
> - Netscape Enterprize Server is running (users can access static HTML
> content)
> - Informix is online
> - WebBlade webdriver is running.
>
> But users cannot access dynamic content via Informix WebBlade cgi webdriver.
> They cannot
> access/ query the Informix database.
>
> One of the symptoms of the problem is that the no. of instances of webdriver
> jumps very
> suddenly from < 5 to > 20 and the total no. of processes jumps from < 220 to >
> 260 and
> both cpus hit approx. 100%.
>
> This happens at least once every day !
>
> My quick and dirty fix is to temporarily rename webdriver, kill the oldest
> webdriver processes,
> wait till CPU load drops and then restore webdriver.
>
> I am just a humble systems admin guy not a DBA and I can't figure out what I can
> do to fix this
> problem. Maybe I could resign ;-)
>
> Any advice on (a) how to analyze the problem and
> (b) how to fix it/ improve performance etc.
> would be greatly appreciated.
>
> Best regards,
>
> Michael Martin
> michael_martin@komatsu.co.jp
>
> --------------------------------------------------------------------
> The following is a snapshot from GlancePlus (a HP monitoring tool) at a
> time when the website had frozen.
>
> My reading of this is:
>
> - two informix processes running at relatively low pri of 233 blocked on PRI
> but munching up cpu (these are big.. RSS includes everything related to a
> process that uses RAM eg. data, stack, text and shared memory segments
> - lots of webdriver processes blocked on SOCKT (waiting for socket operations to
> complete) and all of them have the same PPID
> - ns-httpd blocked on SLEEP ie. just waiting
>
> My thought at the time was that Informix wasn't replying fast enough to
> webdriver
> and that may be part of the problem but I guess I need to figure out exactly
> what
> socket operations the webdriver processes are waiting on. And maybe more
> importantly how does webdriver respond to being blocked on SOCKT... does it
> keep madly spawning new processes ?
>
> Could this socket stuff mean that their is a communications problem on top
> of a possible database design/SQL efficiency problem ?
>
> B3692A GlancePlus C.02.15.00 15:54:06 OYAPSV01 9000/889 Current Avg High
> --------------------------------------------------------------------------------
> Cpu Util S SARU U |100% 39% 100%
> Disk Util F F | 6% 4% 100%
> Mem Util S SU UB B | 48% 43% 48%
> Swap Util U UR R | 21% 18% 21%
> --------------------------------------------------------------------------------
> PROCESS LIST Users= 3
> User CPU Util Cum Disk Block
> Process Name PID PPID Pri Name ( 200% max) CPU IO Rate RSS On
> --------------------------------------------------------------------------------
> oninit 1974 1972 233 informix 95.4/ 7.3 4668.8 0.0/ 0.7 619.5mb PRI
> oninit 1972 1971 233 informix 89.9/12.3 7901.4 0.0/ 1.1 619.5mb PRI
> ns-httpd 1397 1396 154 www 2.7/ 3.7 12264.6 0.0/ 0.0 27.8mb SLEEP
> scopeux 1414 1349 127 root 2.5/ 0.2 627.9 6.9/ 0.5 8.2mb SLEEP
> ns-httpd 1748 1747 154 www 2.1/ 2.0 1263.9 0.0/ 0.0 26.3mb SLEEP
> oninit 1752 1726 156 infgus 0.0/ 0.0 5.6 0.0/ 0.0 119.4mb SEM
> webdriver 10326 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
> oninit 1726 1710 168 infgus 0.0/ 0.0 17.9 0.0/ 0.0 119.4mb SLEEP
> oninit 1727 1710 154 infgus 0.0/ 0.0 29.7 0.0/ 0.0 122.2mb SLEEP
> oninit 1998 1973 156 informix 0.0/ 0.0 12.6 0.0/ 0.0 616.5mb SEM
> webdriver 17255 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.6mb SOCKT
> webdriver 17730 1397 154 www 0.0/ 0.3 0.3 0.0/ 0.0 2.6mb SOCKT
> oninit 1985 1973 156 informix 0.0/ 0.0 7.5 0.0/ 0.0 616.4mb SEM
> webdriver 17254 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.7mb SOCKT
> oninit 1999 1973 154 informix 0.0/ 0.1 53.7 0.0/ 0.0 616.5mb SLEEP
> webdriver 17510 1397 154 www 0.0/ 0.3 0.3 0.0/ 0.0 2.6mb SOCKT
> webdriver 10327 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
> webdriver 17253 1397 154 www 0.0/ 0.2 0.3 0.0/ 0.0 2.7mb SOCKT
> webdriver 17769 1397 154 www 0.0/ 0.4 0.3 0.0/ 0.0 2.6mb SOCKT
> webdriver 10344 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 2.7mb SOCKT
> oninit 1987 1973 156 informix 0.0/ 0.0 15.4 0.0/ 0.0 616.5mb SEM
> oninit 2000 1973 154 informix 0.0/ 0.1 35.9 0.0/ 0.0 616.5mb SLEEP
> webdriver 18244 1397 154 www 0.0/ 0.0 0.3 0.0/ 0.0 3.0mb SOCKT
> oninit 1988 1973 156 informix 0.0/ 0.0 9.0 0.0/ 0.0 616.5mb SEM
> uxwdog 1396 1395 168 www 0.0/ 0.0 0.0 0.0/ 0.0 3.9mb SLEEP
> webdriver 17939 1397 154 www 0.0/ 1.7 0.3 0.0/ 0.0 2.6mb SOCKT
>
> OYAPSV01#>sar -v 4 2
>
> HP-UX OYAPSV01 B.10.20 E 9000/889 10/28/99
>
> 15:53:05 text-sz ov proc-sz ov inod-sz ov file-sz ov
> 15:53:09 N/A N/A 239/1012 0 1212/1212 0 852/9380 0
> 15:53:13 N/A N/A 240/1012 0 1212/1212 0 860/9380 0>
> OYAPSV01#>ps -ef | grep webd | sort -nr -k 2
> www 18950 1397 0 18:53:30 ? 0:00 webdriver
> www 18245 1397 0 18:50:55 ?
A two CPU box running a parallel DB instance and a web server is over loading the machine. get a 4CPu host for the DB server. Leave the webserver on the 2CPu host. -- --------------------------------------------------------- Steven Hauser email: hause011@tc.umn.edu URL: http://www.tc.umn.edu/~hause011 ---------------------------------------------------------
A couple of small things that might help (or might not): 1)Make sure that in the web.cnf file the INFORMIXSERVER variable is set to the TCP/IP connection to your DB (or whatever it is on HP, not a shared memory type connection, or whatever it is on HP). 2)Do your pages have anything that is stored in the DB as blobs (like images or documents? If so, turn on large object cacheing (MI_WEBCACHEDIR, MI_WEBCACHESUB and MI_WEBCACHEMAXLO - these are the variables that work [MI_CACHECRON and MI_WEBCACHELIFE don't work], so look them up in the manual and set them). This will stop any queries to the DB for blobs (after the first). The web server will cache them. Some versions of the blade will allow this for web(app) pages as well. Check your release notes. If you have some static pages, you could cache them too. 3)Run update statistics on your DB. Regularly. 4)Look at setting MI_WEBDRVLEVEL to at least contain the "512" value, then look at the release notes ($INFORMIXDIR/extend/web*/doc/*.txt) to see how to read it. This will slow the system down, but it might help you debug the problem. 5)Put the line: .OPTOFC 1 into your web.cnf file. 6)Think about using the NSAPI interface instead? (although I can't remember ever trying in on the HP platform). 7)If your web application DOES NOT UPDATE the DB in any way (READ ONLY) put the following into your web.cnf file: MI_WEBINITIALSQL SET ISOLATION TO DIRTY READ; That's all I could think of right now. Good luck, tony In article <7vbvls$6j3$1@news.xmission.com>, "michael martin" <michael_martin@komatsu.co.jp> wrote: > > Hi gurus, > > Sorry for the long mail. > > We are running a database backed intranet website on a HP server running HP-UX > 10.20, > Netscape Enterprize 3.0, Informix 7.3 & WebBlade 3.0 using cgi implementation of > webdriver. > > The server has 2 X 180MHz cpus and 3 GB of RAM. > > Users access the website remotely with web browsers running on their PCs. > > Users are frequently experiencing problems with the website; response is always > slow > and in the worst case they cannot access the webdriver driven part of the site > ie. > http://IP address/sitename/cgi-bin/webdriver at all. > > But static content on the same web server eg. http://IP address/static.html can > be accessed. > > At the times when this problem is experienced all relevant services are still > running: > - HP-UX is running (telnet sessions OK, responds to unix commands OK) > - Netscape Enterprize Server is running (users can access static HTML > content) > - Informix is online > - WebBlade webdriver is running. > > But users cannot access dynamic content via Informix WebBlade cgi webdriver. > They cannot > access/ query the Informix database. > > One of the symptoms of the problem is that the no. of instances of webdriver > jumps very > suddenly from < 5 to > 20 and the total no. of processes jumps from < 220 to > > 260 and > both cpus hit approx. 100%. > > This happens at least once every day ! > > My quick and dirty fix is to temporarily rename webdriver, kill the oldest > webdriver processes, > wait till CPU load drops and then restore webdriver. > > I am just a humble systems admin guy not a DBA and I can't figure out what I can > do to fix this > problem. Maybe I could resign ;-) > > Any advice on (a) how to analyze the problem and > (b) how to fix it/ improve performance etc. > would be greatly appreciated. > > Best regards, > > Michael Martin > michael_martin@komatsu.co.jp > > -------------------------------------------------------------------- > The following is a snapshot from GlancePlus (a HP monitoring tool) at a > time when the website had frozen. > > My reading of this is: > > - two informix processes running at relatively low pri of 233 blocked on PRI > but munching up cpu (these are big.. RSS includes everything related to a > process that uses RAM eg. data, stack, text and shared memory segments > - lots of webdriver processes blocked on SOCKT (waiting for socket operations to > complete) and all of them have the same PPID > - ns-httpd blocked on SLEEP ie. just waiting > > My thought at the time was that Informix wasn't replying fast enough to > webdriver > and that may be part of the problem but I guess I need to figure out exactly > what > socket operations the webdriver processes are waiting on. And maybe more > importantly how does webdriver respond to being blocked on SOCKT... does it > keep madly spawning new processes ? > > Could this socket stuff mean that their is a communications problem on top > of a possible database design/SQL efficiency problem ? > -- -- The opinions expressed here are mine and only mine. asalerno@my-deja.com www.summitdata.com Sent via Deja.com http://www.deja.com/ Before you buy.
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g