Oracle RAC node evicted or rebooted due to guest-level DM-Multipath queueing I/O on VMware Virtual Disks.
Issue
- Virtual machines hosting an Oracle RAC database cluster crash and reboot within a very tight window (coinciding within seconds of each other) during transient storage-level latency or path failover on the ESXi hypervisor.
- In
/var/log/messages, the Oracle watchdog (such ascssdmonitororcssdagent) detects that the heartbeat daemon is unresponsive and triggers a hard reboot or kernel panic (fencing/eviction) to protect database integrity:
[cssdagent(4612)]crsd has been unresponsive for 30 seconds. Fencing node.
kernel: SysRq : Trigger a crash dump
- Active DM-Multipath configurations map single-path virtual disks (e.g., VMware VMDKs or virtual RDMs presented as SCSI devices like
/dev/sda,/dev/sdb, or/dev/sdf) inside the guest operating system.
Environment
- Red Hat Enterprise Linux 8
- Red Hat Enterprise Linux 9
- VMware ESXi hypervisor
- Oracle Real Application Clusters (RAC)
device-mapper-multipath(DM-Multipath) active in the guest OS
Subscriber exclusive content
A Red Hat subscription provides unlimited access to our knowledgebase, tools, and much more.