RE: System Freezes
Posted in 2011
Topics: General Discussion
It isn't a virtualized host. It is a dedicated HPUX box.
I'm not seeing substantially more web traffic. I have thousands of
other CGI's and I've never seen anything like this before. What is new
is some very complex SQL.
We're using raw devices.
From: Fernando Nunes [mailto:domusonline@gmail.com]
Sent: Tuesday, January 11, 2011 4:39 PM
To: Rubinstein, James
Cc: informix-list@iiug.org
Subject: Re: System Freezes
I understand that your system freezing will hide any useful stuff from
top or other tools like that. But if you take some snapshots (vmstat?)
it may reveal something.
Also... Is this a virtualized host? Are you sure no other host is
"stealing" your CPU cores? (This would appear in the logs...)
Wild guess... CGIs are nasty... Have you checked your netstat output? Do
you have enough tcp ports configured? CGIs typically connec/disconnect,
and if you have large TIME_WAIT parameters you can end up exhausting
your available ports (it would block new connections, but should not
affect existing ones).
Are you using raw devices or file system? If it's fs, which one?
As you may have noted, I'm shooting into the air...
On Tue, Jan 11, 2011 at 10:26 PM, Rubinstein, James
<JRUBIN@midwestern.edu> wrote:
Nothing of interest in our system log. The problems started Friday when
I turned on my new perl CGI's (which do database operations), were no
existent over the weekend, when the scripts were not being used, and
then started again yesterday, when our university opened for business.
I don't notice any freezes during weekends of evenings (when no one is
working except for me) and they start again when the web traffic and
database activity starts up again.
From: informix-list-bounces@iiug.org
[mailto:informix-list-bounces@iiug.org] On Behalf Of Everett Mills
Sent: Tuesday, January 11, 2011 3:24 PM
Cc: informix-list@iiug.org
Subject: RE: System Freezes
Those symptoms sound like a hardware issue to me. Have you looked for
error messages in /var/adm/syslog/syslog.log?
--EEM
From: informix-list-bounces@iiug.org
[mailto:informix-list-bounces@iiug.org] On Behalf Of Fernando Nunes
Sent: Tuesday, January 11, 2011 4:07 PM
To: Rubinstein, James
Cc: informix-list@iiug.org
Subject: Re: System Freezes
On Tue, Jan 11, 2011 at 9:29 PM, Rubinstein, James
<JRUBIN@midwestern.edu> wrote:
I'm running IDS 11.50.FC6 on HPUX 11.31. We recently rolled out an
internally developed web-based system (Apache perl/mod_perl) and
immediately started noticing that our HPUX system becomes completely
unresponsive, for 20-30 seconds at a time, many times throughout the
day. My first thought was network problems, but we have pretty much
ruled this out since I can connect to a twin HPUX server which seems
fine during the outages. During these system freezes, any connections
to the database fail and the system is completely unresponsive to the
point that I cannot even type any commands at the shell. I have seen
this behavior in the past when the oninit processes use a lot of CPU
resources, but it is usually pretty easy to track these down to some bad
SQL/report writing. In this case, I'm trying to figure out what may be
causing the system freezes. I have top and the dbtop utility from IIUG,
but those don't refresh during or freezes. I am also unable to type any
onstat commands until the system comes back. By that time, everything
looks pretty normal with low load averages and our oninit processes at
normal levels. I'm looking at the various system reports in OAT, but
don't see anything that jumps out as the culprit. I'd appreciate any
troubleshooting ideas.
It may look as I'm defending Informix, but I find it very hard to
believe that the symptoms you describe can be caused by any kind of bad
SQL.
I'd start looking at memory usage etc. Check the OS ratios for
filesystem cache vs program memory. Check you memory usage. Try to keep
a "top" or similar tool open and see if you notice something. Take
frequent snapshots of paging status (so that you can compare
before/after counters.
Also check your system logs.
Last time I saw something similar (not on HP) it was the
filesystem/program memory ratios. If froze the machine whenever a
filesystem intensive operation was run.
By no means I'm insinuating it's everything ok with Informix, but
whatever happens with it should not cause that effect.
Regards.
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/mailman/listinfo/informix-list
--
Fernando Nunes
Portugal
http://informix-technology.blogspot.com
My email works... but I don't check it frequently...
Hello James,
have you set noage or use kaio or use rtprio , in the past using those
may have frozen the os (that is on pa-risc)
post you onconfig maybe someone else sees something obvious??
Are you using cooked or raw??
How much memory is in the machine and how much is informix allowed to
use,
what are the values for
#define DBC_MAX_PCT
#define DBC_MIN_PCT
in your kernel (default 50 % is used for filesystem cache!!!!)
is your swap configured correctly ?? (twice your memory?? )
What does HP say about this.....
Superboer.
On 12 jan, 00:41, "Rubinstein, James" <JRU...@midwestern.edu> wrote:
> It isn't a virtualized host. It is a dedicated HPUX box.
>
> I'm not seeing substantially more web traffic. I have thousands of
> other CGI's and I've never seen anything like this before. What is new
> is some very complex SQL.
>
> We're using raw devices.
>
> From: Fernando Nunes [mailto:domusonl...@gmail.com]
> Sent: Tuesday, January 11, 2011 4:39 PM
> To: Rubinstein, James
> Cc: informix-l...@iiug.org
> Subject: Re: System Freezes
>
> I understand that your system freezing will hide any useful stuff from
> top or other tools like that. But if you take some snapshots (vmstat?)
> it may reveal something.
> Also... Is this a virtualized host? Are you sure no other host is
> "stealing" your CPU cores? (This would appear in the logs...)
>
> Wild guess... CGIs are nasty... Have you checked your netstat output? Do
> you have enough tcp ports configured? CGIs typically connec/disconnect,
> and if you have large TIME_WAIT parameters you can end up exhausting
> your available ports (it would block new connections, but should not
> affect existing ones).
>
> Are you using raw devices or file system? If it's fs, which one?
>
> As you may have noted, I'm shooting into the air...
>
> On Tue, Jan 11, 2011 at 10:26 PM, Rubinstein, James
>
> <JRU...@midwestern.edu> wrote:
>
> Nothing of interest in our system log. The problems started Friday when
> I turned on my new perl CGI's (which do database operations), were no
> existent over the weekend, when the scripts were not being used, and
> then started again yesterday, when our university opened for business.
> I don't notice any freezes during weekends of evenings (when no one is
> working except for me) and they start again when the web traffic and
> database activity starts up again.
>
> From: informix-list-boun...@iiug.org
> [mailto:informix-list-boun...@iiug.org] On Behalf Of Everett Mills
> Sent: Tuesday, January 11, 2011 3:24 PM
>
> Cc: informix-l...@iiug.org
>
> Subject: RE: System Freezes
>
> Those symptoms sound like a hardware issue to me. Have you looked for
> error messages in /var/adm/syslog/syslog.log?
>
> --EEM
>
> From: informix-list-boun...@iiug.org
> [mailto:informix-list-boun...@iiug.org] On Behalf Of Fernando Nunes
> Sent: Tuesday, January 11, 2011 4:07 PM
> To: Rubinstein, James
> Cc: informix-l...@iiug.org
> Subject: Re: System Freezes
>
> On Tue, Jan 11, 2011 at 9:29 PM, Rubinstein, James
>
> <JRU...@midwestern.edu> wrote:
>
> I'm running IDS 11.50.FC6 on HPUX 11.31. We recently rolled out an
> internally developed web-based system (Apache perl/mod_perl) and
> immediately started noticing that our HPUX system becomes completely
> unresponsive, for 20-30 seconds at a time, many times throughout the
> day. My first thought was network problems, but we have pretty much
> ruled this out since I can connect to a twin HPUX server which seems
> fine during the outages. During these system freezes, any connections
> to the database fail and the system is completely unresponsive to the
> point that I cannot even type any commands at the shell. I have seen
> this behavior in the past when the oninit processes use a lot of CPU
> resources, but it is usually pretty easy to track these down to some bad
> SQL/report writing. In this case, I'm trying to figure out what may be
> causing the system freezes. I have top and the dbtop utility from IIUG,
> but those don't refresh during or freezes. I am also unable to type any
> onstat commands until the system comes back. By that time, everything
> looks pretty normal with low load averages and our oninit processes at
> normal levels. I'm looking at the various system reports in OAT, but
> don't see anything that jumps out as the culprit. I'd appreciate any
> troubleshooting ideas.
>
> It may look as I'm defending Informix, but I find it very hard to
> believe that the symptoms you describe can be caused by any kind of bad
> SQL.
> I'd start looking at memory usage etc. Check the OS ratios for
> filesystem cache vs program memory. Check you memory usage. Try to keep
> a "top" or similar tool open and see if you notice something. Take
> frequent snapshots of paging status (so that you can compare
> before/after counters.
> Also check your system logs.
> Last time I saw something similar (not on HP) it was the
> filesystem/program memory ratios. If froze the machine whenever a
> filesystem intensive operation was run.
>
> By no means I'm insinuating it's everything ok with Informix, but
> whatever happens with it should not cause that effect.
>
> Regards.
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
>
> _______________________________________________
> Informix-list mailing list
> Informix-l...@iiug.orghttp://www.iiug.org/mailman/listinfo/informix-list
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
I'm using raw. I have 8GB RAM on the machine. It appears that dbc_max_pct and dbc_min_pct have been replaced by filecache_max and filecahce_min in the hpux 11.31. Here's what I see for those:
# kctune -v filecache_max
Description Maximum amount of physical memory to be used for caching file I/O data
Module fs_bufcache
Current Value 1220798668 [15%]
Value at Next Boot (auto) [15%]
Value at Last Boot 1220798668
Default Value 4069326848 (automatic)
Can Change Immediately (Automatic Tuning Disabled)
# kctune -v filecache_min
Tunable filecache_min
Description Minimum guaranteed physical memory used for caching file I/O data
Module fs_bufcache
Current Value 406929408 [Default]
Value at Next Boot Default (automatic)
Value at Last Boot 406929408
Default Value 406929408 (automatic)
Can Change Automatic Tuning Enabled
I am using noage. I'll post my config in another emal, since I'm exceeding the 50KB limit when I include it here.
-----Original Message-----
From: informix-list-bounces@iiug.org [mailto:informix-list-bounces@iiug.org] On Behalf Of Superboer
Sent: Wednesday, January 12, 2011 2:24 AM
To: informix-list@iiug.org
Subject: Re: System Freezes
Hello James,
have you set noage or use kaio or use rtprio , in the past using those
may have frozen the os (that is on pa-risc)
post you onconfig maybe someone else sees something obvious??
Are you using cooked or raw??
How much memory is in the machine and how much is informix allowed to
use,
what are the values for
#define DBC_MAX_PCT
#define DBC_MIN_PCT
in your kernel (default 50 % is used for filesystem cache!!!!)
is your swap configured correctly ?? (twice your memory?? )
What does HP say about this.....
Superboer.
On 12 jan, 00:41, "Rubinstein, James" <JRU...@midwestern.edu> wrote:
> It isn't a virtualized host. It is a dedicated HPUX box.
>
> I'm not seeing substantially more web traffic. I have thousands of
> other CGI's and I've never seen anything like this before. What is new
> is some very complex SQL.
>
> We're using raw devices.
>
> From: Fernando Nunes [mailto:domusonl...@gmail.com]
> Sent: Tuesday, January 11, 2011 4:39 PM
> To: Rubinstein, James
> Cc: informix-l...@iiug.org
> Subject: Re: System Freezes
>
> I understand that your system freezing will hide any useful stuff from
> top or other tools like that. But if you take some snapshots (vmstat?)
> it may reveal something.
> Also... Is this a virtualized host? Are you sure no other host is
> "stealing" your CPU cores? (This would appear in the logs...)
>
> Wild guess... CGIs are nasty... Have you checked your netstat output? Do
> you have enough tcp ports configured? CGIs typically connec/disconnect,
> and if you have large TIME_WAIT parameters you can end up exhausting
> your available ports (it would block new connections, but should not
> affect existing ones).
>
> Are you using raw devices or file system? If it's fs, which one?
>
> As you may have noted, I'm shooting into the air...
>
> On Tue, Jan 11, 2011 at 10:26 PM, Rubinstein, James
>
> <JRU...@midwestern.edu> wrote:
>
> Nothing of interest in our system log. The problems started Friday when
> I turned on my new perl CGI's (which do database operations), were no
> existent over the weekend, when the scripts were not being used, and
> then started again yesterday, when our university opened for business.
> I don't notice any freezes during weekends of evenings (when no one is
> working except for me) and they start again when the web traffic and
> database activity starts up again.
>
> From: informix-list-boun...@iiug.org
> [mailto:informix-list-boun...@iiug.org] On Behalf Of Everett Mills
> Sent: Tuesday, January 11, 2011 3:24 PM
>
> Cc: informix-l...@iiug.org
>
> Subject: RE: System Freezes
>
> Those symptoms sound like a hardware issue to me. Have you looked for
> error messages in /var/adm/syslog/syslog.log?
>
> --EEM
>
> From: informix-list-boun...@iiug.org
> [mailto:informix-list-boun...@iiug.org] On Behalf Of Fernando Nunes
> Sent: Tuesday, January 11, 2011 4:07 PM
> To: Rubinstein, James
> Cc: informix-l...@iiug.org
> Subject: Re: System Freezes
>
> On Tue, Jan 11, 2011 at 9:29 PM, Rubinstein, James
>
> <JRU...@midwestern.edu> wrote:
>
> I'm running IDS 11.50.FC6 on HPUX 11.31. We recently rolled out an
> internally developed web-based system (Apache perl/mod_perl) and
> immediately started noticing that our HPUX system becomes completely
> unresponsive, for 20-30 seconds at a time, many times throughout the
> day. My first thought was network problems, but we have pretty much
> ruled this out since I can connect to a twin HPUX server which seems
> fine during the outages. During these system freezes, any connections
> to the database fail and the system is completely unresponsive to the
> point that I cannot even type any commands at the shell. I have seen
> this behavior in the past when the oninit processes use a lot of CPU
> resources, but it is usually pretty easy to track these down to some bad
> SQL/report writing. In this case, I'm trying to figure out what may be
> causing the system freezes. I have top and the dbtop utility from IIUG,
> but those don't refresh during or freezes. I am also unable to type any
> onstat commands until the system comes back. By that time, everything
> looks pretty normal with low load averages and our oninit processes at
> normal levels. I'm looking at the various system reports in OAT, but
> don't see anything that jumps out as the culprit. I'd appreciate any
> troubleshooting ideas.
>
> It may look as I'm defending Informix, but I find it very hard to
> believe that the symptoms you describe can be caused by any kind of bad
> SQL.
> I'd start looking at memory usage etc. Check the OS ratios for
> filesystem cache vs program memory. Check you memory usage. Try to keep
> a "top" or similar tool open and see if you notice something. Take
> frequent snapshots of paging status (so that you can compare
> before/after counters.
> Also check your system logs.
> Last time I saw something similar (not on HP) it was the
> filesystem/program memory ratios. If froze the machine whenever a
> filesystem intensive operation was run.
>
> By no means I'm insinuating it's everything ok with Informix, but
> whatever happens with it should not cause that effect.
>
> Regards.
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
>
> _______________________________________________
> Informix-list mailing list
> Informix-l...@iiug.orghttp://www.iiug.org/mailman/listinfo/informix-list
>
> --
> Fernando Nunes
> Portugal
>
> http://informix-technology.blogspot.com
> My email works... but I don't check it frequently...
_______________________________________________
Informix-list mailing list
Informix-list@iiug.org
http://www.iiug.org/m
Here's my config. I've removed comments to get it under the 50KB IIUG
listserv limit.
ROOTNAME root
ROOTPATH /opt/informix/dev/root.1
ROOTOFFSET 0
ROOTSIZE 8388608MIRROR 1
MIRRORPATH
MIRROROFFSET 0
PHYSFILE 2048000
PLOG_OVERFLOW_PATH /opt/informix/tmp
PHYSBUFF 128
LOGFILES 60
LOGSIZE 87040
DYNAMIC_LOGS 0
LOGBUFF 64
LTXHWM 45
LTXEHWM 55
MSGPATH /opt/informix/Logs/cars.log
CONSOLE /dev/console
TBLTBLFIRST 0
TBLTBLNEXT 0
TBLSPACE_STATS 1
DBSPACETEMP temp0
SBSPACETEMP
SBSPACENAME
SYSSBSPACENAME
ONDBSPACEDOWN 1
SERVERNUM 0
DBSERVERNAME hpk360
DBSERVERALIASES istarcars # List of alternate dbservernames
NETTYPE ipcshm,3,250,CPU # Configure poll thread(s) fornettype
NETTYPE soctcp,2,,NET # Configure poll thread(s) for nettype
FASTPOLL 1
LISTEN_TIMEOUT 10
MAX_INCOMPLETE_CONNECTIONS 1024
MULTIPROCESSOR 1
VPCLASS cpu,num=3,max=4,noage
VP_MEMORY_CACHE_KB 0
SINGLE_CPU_VP 0
#VPCLASS aio,num=
CLEANERS 8AUTO_AIOVPS 1
DIRECT_IO 0
LOCKS 512000
DEF_TABLE_LOCKMODE row
RESIDENT 0
SHMBASE 0x0
SHMVIRTSIZE 512000
SHMADD 65536
EXTSHMADD 32768
SHMTOTAL 0
SHMVIRT_ALLOCSEG 0.000000
SHMNOACCESS
CKPTINTVL 900AUTO_CKPTS 1
RTO_SERVER_RESTART 0
BLOCKTIMEOUT 3600
CONVERSION_GUARD 1
RESTORE_POINT_DIR /opt/informix/tmp
TXTIMEOUT 300
DEADLOCK_TIMEOUT 60
HETERO_COMMIT 0
TAPEDEV /dev/rmt/0m
TAPEBLK 32
TAPESIZE 0
LTAPEDEV /dev/rmt/1m
LTAPEBLK 32
LTAPESIZE 0
BAR_ACT_LOG /opt/informix/Logs/bar_act.log
BAR_DEBUG_LOG /opt/informix/Logs/bar_dbug.log
BAR_DEBUG 0
BAR_MAX_BACKUP 0
BAR_RETRY 1
BAR_NB_XPORT_COUNT 20
BAR_XFER_BUF_SIZE 31
RESTARTABLE_RESTORE on
BAR_PROGRESS_FREQ 0
BAR_BSALIB_PATH
BACKUP_FILTER
RESTORE_FILTER
BAR_PERFORMANCE 0
ISM_DATA_POOL ISMData
ISM_LOG_POOL ISMLogs
DD_HASHSIZE 31
DD_HASHMAX 10
DS_HASHSIZE 31
DS_POOLSIZE 127
PC_HASHSIZE 31
PC_POOLSIZE 127
STMT_CACHE 0
STMT_CACHE_HITS 0
STMT_CACHE_SIZE 512
STMT_CACHE_NOLIMIT 0
STMT_CACHE_NUMPOOL 1
USEOSTIME 0
STACKSIZE 128
ALLOW_NEWLINE 0
USELASTCOMMITTED NONE
FILLFACTOR 90
MAX_FILL_DATA_PAGES 0
BTSCANNER num=1,threshold=5000,rangesize=-1,alice=6,compression=default
ONLIDX_MAXMEM 5120
MAX_PDQPRIORITY 100
DS_MAX_QUERIES
DS_TOTAL_MEMORY
DS_MAX_SCANS 1048576
DS_NONPDQ_QUERY_MEM 128
DATASKIP off
OPTCOMPIND 0
DIRECTIVES 1
EXT_DIRECTIVES 0
OPT_GOAL -1
IFX_FOLDVIEW 0AUTO_REPREPARE 1
RA_PAGES
RA_THRESHOLD
BATCHEDREAD_TABLE 0
EXPLAIN_STAT 1
#SQLTRACE level=low,ntraces=1000,size=2,mode=global
DBCREATE_PERMISSION informix
DBCREATE_PERMISSION root
# DB_LIBRARY_PATH /opt/informix/extend
IFX_EXTEND_ROLE 1
SECURITY_LOCALCONNECTION
UNSECURE_ONSTAT 1
ADMIN_USER_MODE_WITH_DBSA
ADMIN_MODE_USERS
SSL_KEYSTORE_LABEL
PLCY_POOLSIZE 127
PLCY_HASHSIZE 31
USRC_POOLSIZE 127
USRC_HASHSIZE 31
STAGEBLOB
OPCACHEMAX 0
ENCRYPT_HDR
ENCRYPT_SMX
ENCRYPT_CDR 0
ENCRYPT_CIPHERS
ENCRYPT_MAC
ENCRYPT_MACFILE
ENCRYPT_SWITCH
CDR_EVALTHREADS 1,2
CDR_DSLOCKWAIT 5
CDR_QUEUEMEM 4096
CDR_NIFCOMPRESS 0
CDR_SERIAL 0,0
CDR_DBSPACE
CDR_QHDR_DBSPACE
CDR_QDATA_SBSPACE
CDR_MAX_DYNAMIC_LOGS 0
CDR_SUPPRESS_ATSRISWARN
DRAUTO 0
DRINTERVAL 30
DRTIMEOUT 30
HA_ALIAS
DRLOSTFOUND /opt/informix/etc/dr.lostfound
DRIDXAUTO 0
LOG_INDEX_BUILDS
SDS_ENABLE
SDS_TIMEOUT 20
SDS_TEMPDBS
SDS_PAGING
UPDATABLE_SECONDARY 0
FAILOVER_CALLBACK
TEMPTAB_NOLOG 0
DELAY_APPLY 0
STOP_APPLY 0
LOG_STAGING_DIR
ON_RECVRY_THREADS 1
OFF_RECVRY_THREADS 10
DUMPDIR /opt/informix/tmp
DUMPSHMEM 1
DUMPGCORE 0
DUMPCORE 0
DUMPCNT 1
ALARMPROGRAM /opt/informix/etc/no_log.sh
ALRM_ALL_EVENTS 0
STORAGE_FULL_ALARM 600,3
SYSALARMPROGRAM /opt/informix/etc/evidence.sh
RAS_PLOG_SPEED 0
RAS_LLOG_SPEED 0
EILSEQ_COMPAT_MODE 0
QSTATS 0
WSTATS 0
jvp,num=1,noage
JVPJAVAHOME /opt/informix/extend/krakatoa/jre
JVPHOME /opt/informix/extend/krakatoa
JVPPROPFILE /opt/informix/extend/krakatoa/.jvpprops
JVPLOGFILE /opt/informix/Logs/jvp.log
#JDKVERSION 1.5
JVPJAVALIB
JVPJAVAVM libjava.so
#JVPARGS -verbose:jni#JVPCLASSPATH
/opt/informix/extend/krakatoa/krakatoa_g.jar:/opt/informix/extend/krakat
oa/jdbc_g.jar
JVPCLASSPATH
/opt/informix/extend/krakatoa/krakatoa.jar:/opt/informix/extend/krakatoa
/jdbc.jar
AUTO_LRU_TUNING 1
BUFFERPOOL
size=2K,buffers=200000,lrus=32,lru_min_dirty=2.000000,lru_max_dirty=4.000000