The system crashed with a watchdog device due to grep to /dev/watchdog

Solution Verified - Updated -

Issue

  • On HPE ProLiant systems, an abrupt kernel panic occurs with the following messages in the logs:
[93166.104720] watchdog: watchdog0: watchdog did not stop!
[93187.224628] Kernel panic - not syncing: 02: An NMI occurred. Depending on your system the reason for the NMI is logged in any one of the following resources:
[93187.224629] 1. Integrated Management Log (IML)
[93187.224629] 2. OA Syslog
[93187.224629] 3. OA Forward Progress Log
[93187.224629] 4. iLO Event Log
[93187.224630] CPU: 0 PID: 0 Comm: swapper/0 Kdump: loaded Not tainted 4.18.0-348.23.1.el8_5.x86_64 #1
[93187.224630] Hardware name: XXXXX, BIOS U30 11/24/2021
[93187.224630] Call Trace:
[93187.224630]  <NMI>
[93187.224631]  dump_stack+0x5c/0x80
[93187.224631]  panic+0xe7/0x2a9
[93187.224631]  ? native_apic_msr_write+0x27/0x30
[93187.224631]  nmi_panic.cold.9+0xc/0xc
[93187.224632]  hpwdt_pretimeout+0x7f/0xc2 [hpwdt]
[93187.224632]  nmi_handle+0x63/0x110
[93187.224632]  unknown_nmi_error+0x16/0x30
[93187.224632]  do_nmi+0x183/0x1e0
[93187.224632]  end_repeat_nmi+0x16/0x6f
[93187.224633] RIP: 0010:intel_idle+0x6b/0xb0
[93187.224633] Code: 40 5c 01 00 48 89 d1 0f 01 c8 48 8b 00 a8 08 75 19 e9 07 00 00 00 0f 00 2d 7e f0 51 00 c1 ee 18 b9 01 00 00 00 89 f0 0f 01 c9 <65> 48 8b 04 25 40 5c 01 00 f0 80 60 02 df f0 83 44 24 fc 00 48 8b
[93187.224633] RSP: 0018:ffffffff86403e38 EFLAGS: 00000002
[93187.224635] RAX: 0000000000000020 RBX: ffffffff86535cc8 RCX: 0000000000000001
[93187.224635] RDX: 0000000000000000 RSI: 0000000000000020 RDI: 0000000000000003
[93187.224635] RBP: ffffce3fbfa00338 R08: 0000000000000002 R09: 0000000000029a00
[93187.224635] R10: 0000dc1c61feff24 R11: ffff9551ff828ec4 R12: 0000000000000003
[93187.224636] R13: ffffffff86535b60 R14: 0000000000000003 R15: 0000000000000003
[93187.224636]  ? intel_idle+0x6b/0xb0
[93187.224636]  ? intel_idle+0x6b/0xb0
[93187.224636]  </NMI>
[93187.224636]  cpuidle_enter_state+0x87/0x3d0
[93187.224637]  cpuidle_enter+0x2c/0x40
[93187.224637]  do_idle+0x239/0x270
[93187.224637]  cpu_startup_entry+0x6f/0x80
[93187.224637]  start_kernel+0x51d/0x53d
[93187.224637]  secondary_startup_64_no_verify+0xc2/0xcb
  • On DELL systems, the system simply resets with the following message in the logs.
[83166.305720] watchdog: watchdog0: watchdog did not stop!

Environment

  • Red Hat Enterprise Linux
  • HP ProLiant system with the [hpwdt] module loaded
  • Dell PowerEdge systems with the [iTCO_wdt or wdat_wdt] module loaded
  • Other systems with the [iTCO_wdt] module loaded

Subscriber exclusive content

A Red Hat subscription provides unlimited access to our knowledgebase, tools, and much more.

Current Customers and Partners

Log in for full access

Log In

New to Red Hat?

Learn more about Red Hat subscriptions

Using a Red Hat product through a public cloud?

How to access this content