Re: Huge performance problems Sun/Informix
Posted in 2000
Topics: High Availability & Replication, Backup & Restore, Performance & Tuning, Storage & Space Management, Server Administration, Transactions, Locking & Isolation, Logging & Checkpoints, Platform-Specific Issues, Clustering, Grid & MACH11, Jobs, Consulting & Announcements
Thanks for your reply !
The size of physical dbspace is 50000 KB.
We have made a call to Informix Techsupport a few months ago because we
had long checkpoints (>20sec)
but i/o on the disks for physical and logical logs was very low. So we
both decided
that we may have some problems with small i/o request sizes (as
mentioned in my first
posting) and raised the values for PHYSBUFF and LOGBUFF. Now our
checkpoint
duration is about 1-7 seconds. But we also changed the number of
CLEANERS and LRUsand raised the checkpoint interval up to 900 seconds, because we had
very little time
to find a solution for those long checkpoints (users were not able to do
their work).
So I do not know what action caused better checkpoint times.
Regarding the tunables like RESIDENT or SHMVIRTSIZE or larger stripes I
must say
it doesn't make ANY difference. After two months of operation IDS
allocated
two additional shared mem segments and that sounds not very high for me.
Residency whether enabled or disabled didn't lead to more disk
throughput anyway.
And we tried several types of disk layout: stripes with more disks,
stripes with greater
and smaller unit sizes, modified Veritas config (read from mirror plexes
either by
'preferred plex' method or round robin) and so on.... Nothing had lead
to any
performance enhancement.
In my opinion it must be possible to do a full table scan with more
than 10MB/s stripe
throughput. Or am I wrong ?? How does anybody handle large DSS with a
lot of sequential scans ?
Please help !!!!
Thomas Vogt
Firmengruppe Dr. Gueldener
Steve Romankiw wrote:
>
> Thomas:
>
> I have a very similiar environment (almost identical!). The only difference is I use
> QualixHA+ for failover.
>
> A few that I would recommend:
>
> ** Turn on Residency
> RESIDENT 1 # Forced residency flag (Yes = 1, No =0)>
> ** Do you have a large number of virtual segments created (onstat -g seg)? If so
> increase the initial size
> SHMVIRTSIZE 32768 # initial virtual shared memory segment size>
> ** As for disk layout, I created stripe sets (RAID0+1) with 5 members, The stripe
> unit is set to 16K for OLTP and 64K for DSS. This seemed to get the best all-around
> performance.
>
> ** How big is your Physical dbspace?
>
> ** Why is your PHYSBUFF and LOGBUFF so big?
>
> Steve Romankiw
>
> Thomas Vogt wrote:
>
> > Hello everybody,
> >
> > we have huge performance problems on our 2-node Sun-Cluster with
> > Informix:
> >
> > Environment:
> > Sun E4500, 6 CPU's, 1GB Memory (primary host for database)
> > Sun E3500, 4 CPU's, 1GB Memory (secondary host for takeover)
> > 4 A5000-Diskarrays with 14 9,2GB disks each connected via FCAL
> > Solaris 2.6, Sun Cluster 2.1
> > disk mirroring using Veritas Volume Manager (Raid 0/1)
> > Gigabit Network Adapters
> > Informix Dynamic Server 7.30UC9
> >
> > Workload: for example:
> >
> > 400.000 inserts and updates per day in one non-fragmented table, row
> > size 265 bytes
> > via continuous batch processes on server.
> > 400.000 inserts of blob data (average blob size 10K)
> > (total sum per month are >2 Mio data rows and blobs)
> >
> > 100 OLTP users selecting small portions of data and blobs, updating
> > data table
> >
> > several batch programs for statistic analysis and invoice generation,
> > mainly
> > selecting from data table
> >
> > All dbspaces are located on raw devices.
> > Root-, physical log and logical log dbspaces are on separated disk.
> > Table dbspace is separated from index dbspaces on a stripe with 3 disks
> > each (
> > stripe interleave is 16K which seems to be the same as Informix'
> > BIGREADS).
> > The blobs are stored in their own blobspaces with blob page size = 10K.
> >
> > Extract from onconfig:
> > =============================================================
> > NETTYPE tlitcp,2,100,NET # Configure poll thread(s) for nettype
> > DEADLOCK_TIMEOUT 60 # Max time to wait of lock in> > distributed env.
> > RESIDENT 0 # Forced residency flag (Yes = 1, No =
> > 0)
> >
> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for> > multi-processor
> > # on E4500 5 CPUs
> > # on E3500 3 CPUs
> > NUMCPUVPS 5 # Number of user (cpu) vps
> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps> > to one
> >
> > NOAGE 1 # Process aging
> > AFF_SPROC 1 # Affinity start processor
> > AFF_NPROCS 5 # Affinity number of processors> >
> > # Shared Memory Parameters
> >
> > LOCKS 1000000 # Maximum number of locks
> > BUFFERS 100000 # Maximum number of shared buffers
> > NUMAIOVPS 2 # Number of IO vps
> > PHYSBUFF 1024 # Physical log buffer size (Kbytes)
> > LOGBUFF 1024 # Logical log buffer size (Kbytes)> > LOGSMAX 30 # Maximum number of logical log files
> > CLEANERS 10 # Number of buffer cleaner processes
> > SHMBASE 0xa000000 # Shared memory base address
> > SHMVIRTSIZE 32768 # initial virtual shared memory segment> > size
> > SHMADD 32768 # Size of new shared memory segments
> > (Kbytes)
> > SHMTOTAL 0 # Total shared memory (Kbytes).
> > 0=>unlimited
> > CKPTINTVL 900 # Check point interval (in sec)
> > LRUS 20 # Number of LRU queues
> > #> > # lowered for better checkpoint time
> > LRU_MAX_DIRTY 5 # LRU percent dirty begin cleaning limit
> > LRU_MIN_DIRTY 3 # LRU percent dirty end cleaning limit
> > LTXHWM 50 # Long transaction high water mark> > percentage
> > LTXEHWM 60 # Long transaction high water mark
> > (exclusive)
> > TXTIMEOUT 0x12c # Transaction timeout (in sec)
> > STACKSIZE 32> >
> > =======================
> >
> > Problems: it seems not possible to get more than 300K/s throughput from
> > individual disks
> > (more than 1.2MB/s from a 3 disk stripe) if we do a sequential scan
> > from the data
> > table with a result set of very few rows (<10).
> > But... monitoring the i/o with iostat shows that all disks are
> > only 10-20% busy
> > and doing just 70 reads per second. The whole system shows 70 to 80%
> > idle time, each
> > CPU is only 8% busy !!
> >
> > But.. a level 0 or level 1 backup of data dbspaces and blobspaces
> > with onbar
> > and Legato Networker results in a throughput between 10MB/s and
> > 36MB/s for each
> > stripe !!!!
> >
> > In the last months we had several dates with Sun and Informix
> > consultants.
> >
> > Sun says: there is no problem w
In article <88cafc$rmc$1@news.xmission.com>, Thomas Vogt
<t.vogt@drgueldener.de> writes
>
>Thanks for your reply !
>
>The size of physical dbspace is 50000 KB.
>
Dave comes in fast from the wasteland that is his latest weekend
upgrade...
First I need:-
Hardware
========
Number/speed of CPU's?
Number/size of physical disks (ignore strip/raid what are the
underlying disks)?
Memory?
Software
========
OS and version
Informix products installed + version.
Software used for striping/mirroring.
Since Borland was mentioned I assume a client/server over the network
system.
Network
=======
If os is UNIX then
netstat -i
netstat -e
else NT equivalent (network i/o information)
Informix info
=============
onstat -p
onstat -u
onstat -d
onstat -D
onstat -c
onstat -g ath
onstat -g ioq
>We have made a call to Informix Techsupport a few months ago because we
>had long checkpoints (>20sec)
>but i/o on the disks for physical and logical logs was very low. So we
>both decided
>that we may have some problems with small i/o request sizes (as
>mentioned in my first
>posting) and raised the values for PHYSBUFF and LOGBUFF. Now our
Are your databases using buffered or unbuffered logging?
>checkpoint
>duration is about 1-7 seconds. But we also changed the number of
>CLEANERS and LRUs>and raised the checkpoint interval up to 900 seconds, because we had
>very little time
>to find a solution for those long checkpoints (users were not able to do
>their work).
>So I do not know what action caused better checkpoint times.
>
>Regarding the tunables like RESIDENT or SHMVIRTSIZE or larger stripes I
>must say
>it doesn't make ANY difference. After two months of operation IDS
Correct. Resident stops Informix being swapped and also long as
onstat -g seg
only show 1 segment of type V (virtual) then increases SHMVIRTSIZE has
no effect.
>allocated
>two additional shared mem segments and that sounds not very high for me.
Depends. HP processors only have 3 registers on the CPU to handle
shmem segment permissions and hence HP-UX slows down a lot when
>3 segments for Online. May also apply to other CPU's. Best to
have 1 segment of type V.
>Residency whether enabled or disabled didn't lead to more disk
>throughput anyway.
>And we tried several types of disk layout: stripes with more disks,
>stripes with greater
>and smaller unit sizes, modified Veritas config (read from mirror plexes
>either by
>'preferred plex' method or round robin) and so on.... Nothing had lead
>to any
>performance enhancement.
>
???
>In my opinion it must be possible to do a full table scan with more
>than 10MB/s stripe
>throughput. Or am I wrong ?? How does anybody handle large DSS with a
>lot of sequential scans ?
>
>Please help !!!!
>
>Thomas Vogt
>Firmengruppe Dr. Gueldener
>
>
>
>Steve Romankiw wrote:
>>
>> Thomas:
>>
>> I have a very similiar environment (almost identical!). The only difference
>is I use
>> QualixHA+ for failover.
>>
>> A few that I would recommend:
>>
>> ** Turn on Residency
>> RESIDENT 1 # Forced residency flag (Yes = 1, No =0)>>
>> ** Do you have a large number of virtual segments created (onstat -g seg)? If
>so
>> increase the initial size
>> SHMVIRTSIZE 32768 # initial virtual shared memory segment size>>
>> ** As for disk layout, I created stripe sets (RAID0+1) with 5 members, The
>stripe
>> unit is set to 16K for OLTP and 64K for DSS. This seemed to get the best all-
>around
>> performance.
>>
>> ** How big is your Physical dbspace?
>>
>> ** Why is your PHYSBUFF and LOGBUFF so big?
>>
>> Steve Romankiw
>>
>> Thomas Vogt wrote:
>>
>> > Hello everybody,
>> >
>> > we have huge performance problems on our 2-node Sun-Cluster with
>> > Informix:
>> >
>> > Environment:
>> > Sun E4500, 6 CPU's, 1GB Memory (primary host for database)
>> > Sun E3500, 4 CPU's, 1GB Memory (secondary host for takeover)
>> > 4 A5000-Diskarrays with 14 9,2GB disks each connected via FCAL
>> > Solaris 2.6, Sun Cluster 2.1
>> > disk mirroring using Veritas Volume Manager (Raid 0/1)
>> > Gigabit Network Adapters
>> > Informix Dynamic Server 7.30UC9
>> >
>> > Workload: for example:
>> >
>> > 400.000 inserts and updates per day in one non-fragmented table,
>row
>> > size 265 bytes
>> > via continuous batch processes on server.
>> > 400.000 inserts of blob data (average blob size 10K)
>> > (total sum per month are >2 Mio data rows and blobs)
>> >
>> > 100 OLTP users selecting small portions of data and blobs,
>updating
>> > data table
>> >
>> > several batch programs for statistic analysis and invoice
>generation,
>> > mainly
>> > selecting from data table
>> >
>> > All dbspaces are located on raw devices.
>> > Root-, physical log and logical log dbspaces are on separated disk.
>> > Table dbspace is separated from index dbspaces on a stripe with 3 disks
>> > each (
>> > stripe interleave is 16K which seems to be the same as Informix'
>> > BIGREADS).
>> > The blobs are stored in their own blobspaces with blob page size = 10K.
>> >
>> > Extract from onconfig:
>> > =============================================================
>> > NETTYPE tlitcp,2,100,NET # Configure poll thread(s) for nettype
>> > DEADLOCK_TIMEOUT 60 # Max time to wait of lock in>> > distributed env.
>> > RESIDENT 0 # Forced residency flag (Yes = 1, No =
>> > 0)
>> >
>> > MULTIPROCESSOR 1 # 0 for single-processor, 1 for>> > multi-processor
>> > # on E4500 5 CPUs
>> > # on E3500 3 CPUs
>> > NUMCPUVPS 5 # Number of user (cpu) vps
>> > SINGLE_CPU_VP 0 # If non-zero, limit number of cpu vps>> > to one
>> >
>> > NOAGE 1 # Process aging
>> > AFF_SPROC 1 # Affinity start processor
>> > AFF_NPROCS 5 # Affinity number of processors>> >
>> > # Shared Memory Parameters
>> >
>> > LOCKS 1000000 # Maximum number of locks
>> > BUFFERS 100000 # Maximum number of shared buffers
>> > NUMAIOVPS 2 # Number of IO vps
>> > PHYSBUFF 1024 # Physical log buffer size (Kbytes)
>> > LOGBUFF 1024 # Logical log buffer size (Kbytes)>> > LOGSMAX 30 # Maximum number of logical log files
>> > CLEANERS 10 # Number of buffer cleaner processes
>> > SHMBASE 0xa000000 # Shared memory base address
>> > SHMVIRTSIZE 32768 # initial virtual shared memory segment>> > size
>> > SHMADD 32768 # Size of new
Related threads
- Posting from the Informix-list
- Migrating from IDS 9.40.UC6 to 11.50.UC3
- Ip for a network session
- questions onstat -g