IDS 2000 log errors, PANIC and freezing
Posted in 2000
Hi all,
I have a big big probleme with an IDS 2000 db.
The version is
Informix Dynamic Server 2000 Version 9.20.UC2 -- On-Line -- Up 2 days
14:04:14 -- 325632 Kbytes
on Solaris 2.6 on an ultra sparc 4500/4CPU/2GbRAM
The base, which is less than 2 Gb of data, is used for our portal. There
are about 10 static tables which never changes, and 20 that are changed
multiples times during a day (delete everything and the insert new
data). (don't ask me why they choose to use a database... :)
This database was used for our old portal without problem. The new
version do more connections to the database. I have problem since this.
Most of the tables are made like this :
Column name Type Nulls
xmlcode_id nchar(30) no
xmlcode_stamp datetime year to fraction(3) no
xmlcode_doc text yes
Text is in a blobspace, and is XML pages.
The biggest problem in the database is an error, showing up everytime
the database is in a 'load' state. (in fact the server is never in a
load state) :
18:49:47 Action: Run 'oncheck -cD 1048716'
18:49:47 stack trace for pid 15106 written to /tmp/af.493aa89
18:49:47 See Also: /tmp/af.493aa89
18:50:16 Assert Failed: Incorrect BLOB stamps.
18:50:16 Informix Dynamic Server 2000 Version 9.20.UC2
18:50:16 Who: Session(151, portalp@host.net, -1, 240144352)
Thread(171, sqlexec, e4da798, 1)
File: rsdebug.c Line: 1799
18:50:16 Results: BLOBSpace blobdbs, BLOB addr: 0x60073c, BLOB stamp
-13149
Informix said it was a corruption in a blobspace. They said the only
solution was to unload data, initialise the instance and load data to a
fresh database. I the build a new instance, the same as the old one, and
loaded the data into it. I put the news instance live, and waited. At
18H, as usual, errors appears.
They say it could be a problem of the "unload/reload" system used to
transfert data. As data is unloaded in an ASCII form, and every data
seems to be there (the portal is working well), I'm sure they are right.
I'm now asking you ???
I don't know if it is linked, but last week the database went to PANIC,
after having one error every second in the log. 3 times during the last
week the database freezed. I had to kill every oninit process by hand,
run onmode -ky and the oninit. It was also at the same time, when the
portal is used most.
Here is one of the assert failure produced :
>cat /tmp/af.211fa6ba
18:05:14
18:05:14 Informix Dynamic Server 2000 Version 9.20.UC2 SoftwareSerial Number AAC#xxhiddenxx
18:05:14 Assert Failed: Incorrect BLOB stamps.
18:05:14 Who: Session(7467, portalp@ac13021e, -1, 240150680)
Thread(7479, sqlexec, e4df818, 1)
File: rsdebug.c Line: 1799
18:05:14 Results: BLOBSpace blobdbs, BLOB addr: 0x6003f4, BLOB stamp
-30452
18:05:14 Action: Run 'oncheck -cD 5242933'
18:05:14 Stack for thread: 7479 sqlexec
base: 0x12810018
len: 66048
pc: 0x0068f150
tos: 0x1281ebf8
state: running
vp: 1
0x0068e3d0 (oninit)afhandler(0x1, 0x14c30fc8, 0x99e1dc, 0x99e5dc, 0x401,
0x1)
0x0068db38 (oninit)afwarn_interface(0x14c30fc8, 0x99e1dc, 0x99e5dc,
0x94ba68, 0x707, 0xe511a80)
0x004caea4 (oninit)bfbcheck(0x99e1dc, 0x4400, 0x99db90, 0x1, 0x6003f4,
0x18)
0x004b3ce0 (oninit)bfbget (0x10263ecc, 0x1, 0x0, 0xfa8cfc0, 0x48,
0x1033fca0)
0x004b071c (oninit)rsbread (0xf985c00, 0x10263ecc, 0xf986494, 0x0,
0xf986018, 0x10263ec0)
0x001bad80 (oninit)cpblob2pipe(0x1000, 0x19d, 0x9a6bb0, 0x892ad8,
0xef766ad0, 0xef506888)
0x001ba9e8 (oninit)sq_fetchblob(0x26, 0x0, 0x1346, 0x12, 0xfb3d060,
0xf81f018)
0x00397a38 (oninit)sqmain (0x9b22b0, 0x992d98, 0x9a6bb0, 0x4, 0xa, 0x0)
0x00670054 (oninit)startup (0x9b228c, 0x12, 0x0, 0xe71ded0, 0x0,
0x1472195c)
0x0068b034 (oninit)kaiothread(0x0, 0x0, 0x0, 0x0, 0x0, 0x0)
0x00000000 (*nosymtab*)0x0
18:05:14 See Also: /tmp/af.211fa6ba
---------------------------------
Begin System Alarm Program Output
---------------------------------
Assertion Failure Type: Warning
Host Name: host
Database Server Name: portalp_p_shm
Time of failure: Thu Oct 5 18:05:15 MET DST 2000
AF file: /tmp/af.211fa6ba
Shared memory file: None
System Blocking: OFF
-------------------------------
End System Alarm Program Output
-------------------------------
18:05:15 sh /opt/informix/informix-9.20.UC2/etc/evidence.sh 1 0/tmp/af.211fa6ba 7467 0xe4df818 7479 0x10d54d80 1025 0 0 0 0
18:05:15
------------------ End of assertion failure 0 -----------------
Here is an explanation of how the database is used :
on one side we have a crontab deleting old data and inserting new one.
This is done by a java applet. We use JDBC.
On the other side, we have many (3 :) servers openning 10 connections
each (sometime more) to the database. These connections never end. Each
time the application server need to talk to the database, it uses one of
his connections. This is full java too (JDBC).
I don't know exactly how I could know what the JDBC version is ? Maybe
this is the problem ? Maybe there is a patch or a new JDBC ?
I actualy have no clue.
Just contact me if you have any idea, or need more details. Informix
France is getting ssh, so they will be able to log into the server
(what ? ssh ? is it like a telnet ? :)))))
Thanks for you help
(just try our portal when the database is running : www.freesbee.fr ;
come to see me on the chat : chat.freesbee.fr, web or IRC)
--
_______________________________________________
Sebastien THOMAS none networks
System Engineer freesbee
geo:153, rue Saint-Denis, 75002 Paris, France
vox:+33 1 45 08 23 10 - fax:+33 1 45 08 25 29
mailto:sebastien.thomas@none.net