Re: On-Line Perform. on RS6000
Posted in 1994
LMARKS@HANYS.ORG writes: |> We have been running INFORMIX On-Line (now vers. 5.01) on an IBM |> RS6000 Model 930 for almost 4 years. We have been disappointed with |> the performance from the beginning and have sought the advice of both |> IBM and INFORMIX supplied consultants. None have been able to identify |> and resolve our problems. We run OnLine 5.00 on an RS/6000 Model 980 with a 10.7GB database and a daytime average of 180+ users (using I-Star/I-Net from clients). Although we ~have~ had problems with Informix, none is along the lines of your description. It's been a while since we've had a database server with an x30 processor (and I'm sure IBM's told you to upgrade), but that shouldn't cause your performance to degrade. There are a few things I'd suggest as initial investigative steps, but I'd hope that you have used them already (or IBM has). In any case, I'll list a few in case one helps: * Use "ps -ef | sort -n +3 | tail". This will tell you the top ten 'short term' CPU hogs (the 4th column). A 'large' number could show a culprit process. This must be used in conjunction with other CPU measures (from "vmstat" or "sar -q"), but it ~could~ show a problem. If you see anything that's system-related there all the time on a loaded system, I'd be surprised. Our AIX O/S doesn't impose itself that much. * Use "iostat" to check for high disk activity (the default will also show CPU activity at a high level). In our case, run queues above 1-2 usually mean no I/O wait. On the other hand, with a run queue of 0, you can't help but have I/O wait in most cases. Anyway, look for frequent "tm_act" numbers above 35% (my rule of thumb for I/O concern). One special case where you may have problems is when your root dbspace shares disk with a high-use table. Root dbspace disks with 100% I/O activity rates are bad news for interactive users. Over time, we have striped our root dbspace across 5 disks using the LVM and tried to keep high-use tables (some with millions of rows) away from each other. It's not pretty, but you can use "tbstat -t" and the "systables" (to translate the hex) to find out which tables are most active (if you don't know). Of course, with a smaller database you may have fewer disks across which to spread your I/O; however, with the low cost of disk, a few 1GB drives could be a big help. * IBM has two non-released tools (friendly SE's ~used~ to be a source) which were handy for monitoring the system: xmconsole (X Windows) and smon (ASCII). These can show you if you're running into memory problems. Of course, if you page a lot, you have to worry about that I/O, too. We keep our O/S and user disks separate from our Informix disks. * I'm sure Informix mentioned this, but _8_ is supposed to be a magic number for table extents. Many extents can mean poor performance. |> I have noticed that immediately after we reboot the system, and with only |> 1 individual logged in, anything run in INFORMIX completes very quickly. |> As an example, we have a 2.5 million row, 850MB database. If we do a |> select (*) from the db, and dump the results to /dev/null, we can read |> through the file in about 40 minutes. But, after the system has been up |> and running, and with other users logged in, performance degrades to, |> at best, 2-3 hours for the same select, and at worst, can take 6-8 hours. As an example of a *minor* database design error having *major* consequences, we had an unindexed 10,000 row table bring our system to it's knees. This table started out small and grew over two months, so the performance problem snuck up on us. Before it was fixed, the 30-second search, multiplied by 100 users, caused horrendous response times, high run queues, and angry executives. :-( That's not to say that I *assume* you have a poorly designed database or lousy applications - just that, in my experience, minor mistakes can quickly multiply in large database environments with many users. |> I realize that with a single processor machine, load is not handled well, |> and some degradation occurs, but am wondering if there are "hidden or |> invisible" things, like daemons, memory eaters, cpu hogs, processes |> hung and still running, or other things that might be causing our extreme |> degradation. Still, even AIX's UP systems ~are~ designed to multi-task, and we ~do~ support nearly 200 users with a database over 10 times as large (albeit on a faster processor). Of course, at the current, relatively low price of the larger processors, it may be worth upgrading your 9X0 (although it may be hard to justify spending any more after 4 years of bad performance). -- Dave Newell/dbn@alert.com o_ _/ _/ _/_/_/ _/_/_/ _/_/_/ _0 Alert Centre, Inc. | ' _/ _/ _/ _/_/ _/ _/ _/ `\\ \\ Englewood, CO `\\(*)_/_/_/ _/ _/ _/_/ _/ (*)/ ' (303)488-7781_______________(*)___/ _/ _/_/_/ _/_/_/ _/ _/ _/________(*)_