DR: Needed to send a ping message but failed. 1
Posted in 2013
Topics: High Availability & Replication, Performance & Tuning, Logging & Checkpoints, Platform-Specific Issues
Hi All,
I am running IDS 11.70 FC7 on AIX server.
I am running primary, HDR (async) and RSS server, I have noticed an increase
in the frequency of following messages:
DR: Needed to send a ping message but failed. 1
I have checked the network connectivity between primary and secondary and it
is normal and its not showing any glitches.
Can someone help me that what particular i can monitor when it start logging
these messages in the online.log, so that i could tune that particular item to
improve the performance of both primary & secondary.
One thing i noticed that when this message appears in the log, the next
checkpoint is taking more time to complete than normal.
Also i noticed that during this time, there were some user sessions (in the
output of onstat -u command) which were showing the following flags:
G-BPX--
(which indicates the waiting for write of the logical log buffer)
I have checked the IO activity for the disk and CPU utilizations and they are
very normal.
Thanks in advance
Regards,
Khan
The failing to send a ping rather than the more common ping timeout, seems to mean that the server was for whatever reason unable to start the OS ping command. At least I've always assumed the DR pings were passed out to the OS. Anyway, you may want to check on some of your OS resources like max files, max procs and max user procs (I don't know what AIX calls them). --EEM > -----Original Message----- > From: ids-bounces@iiug.org [mailto:ids-bounces@iiug.org] On Behalf Of O > KHAN > Sent: Tuesday, November 12, 2013 4:26 PM > To: ids@iiug.org > Subject: Re: DR: Needed to send a ping message but fail.... [31938] > > can someone help me? > > > *********************************************************************** > ******** > Forum Note: Use "Reply" to post a response in the discussion forum.
Original post:
Hi All,
I am running IDS 11.70 FC7 on AIX server.
I am running primary, HDR (async) and RSS server, I have noticed an increase
in the frequency of following messages:
DR: Needed to send a ping message but failed. 1
I have checked the network connectivity between primary and secondary and it
is normal and its not showing any glitches.
Can someone help me that what particular i can monitor when it start logging
these messages in the online.log, so that i could tune that particular item to
improve the performance of both primary & secondary.
One thing i noticed that when this message appears in the log, the next
checkpoint is taking more time to complete than normal.
Also i noticed that during this time, there were some user sessions (in the
output of onstat -u command) which were showing the following flags:
G-BPX--
(which indicates the waiting for write of the logical log buffer)
I have checked the IO activity for the disk and CPU utilizations and they are
very normal.
Thanks in advance
Regards,
Khan
Response:
Ok, so in newer releases the HDR ping mechanism was changed. Any type of HDR
message sent counts as the connection still being ok (since that's all the
ping is trying to verify is that data is still getting to the secondary). This
new message you are seeing can show up if the dr_prping thread tries to send a
ping message, but there are currently no free hdr buffers to put a ping
message into (no free hdr buffers when doing async hdr would mean that at the
moment your primary is generating hdr buffers faster then the secondary can
receive and apply them so it's likely your secondary is falling behind the
primary a bit). When the ping thread can't get a hdr buffer (as there is only
a small fixed number of them) it sends a message to the secondary to see if
it's in a checkpoint (as checkpoints on the secondary can temporarily delay
applying incoming log records) and if it is not, it then checks to see if
anything has been sent/received since it's last check, and if nothing as been
sent/received, it reports that message in the MSGPATH file. I think the code
thinks that a longer time would have passed between when we 1st check if
anything has been sent/received and the 2nd time we check...in my opinion it's
very likely that the time between the 1st and 2nd check is not very large at
all and it could easily not have had enough time passed for us to noticed new
data sent/received (as we are only looking in second increments).
So as long as you are not ping timing out, the message itself is pretty
harmless. But we you seeing users waiting on the log buffer it would indicate
that either the secondary is at times falling behind the primary, or you are
generating logical log data faster then your network will allow it to be sent
and ack'd. How fast is the network between your primary and secondary? What is
the round trip network travel time (the round trip time will have a
significant impact on the number of hdr buffers that can be sent per second
between the hdr pair and if you sometimes generate more log data then you will
see users start to block waiting to get to the logical log buffers like you
show in your post).
Jacques Renaut
IBM Informix Advanced Support
APD Team