An observation/question.
With 750,000 buffers and your current LRU cleaner settings, you would
have 22,500 dirty buffers before cleaning would start. Is that many
buffers actually doing you a measurable amount of good?
Along that same line, Chunk writes still outnumber LRU write (4032951
to 2454820). What are the length of your checkpoints? (onstat -m).
In the onstat -p sequence you posted, you're really spinning through
the lock requests. From this I assume you might be using row level
locks. There doesn't seem to be too many conflicts, rollbacks, or
lock waits. Assuming lock requests take some amount of time, would
you achieve anything by changing to page level locks?