Kernel panic after "Poison overwritten" in kmalloc-64 allocated in request_key_auth_new() and freed in keyctl_instantiate_key_common() on RHEL 8.10.z kernel 4.18.0-553.166.1.el8_10 and later

Solution Verified - Updated -

Issue

Crash pattern without slub_debug

  • Without slub_debug enabled, the bug can show as kernel crashes in different places, depending on what has reused the freed memory.

  • The crash can be in the keys code itself, for example in free_request_key_auth() or search_process_keyrings() called from the request-key helper such as nfsidmap, or in an unrelated code path. One example is a fault in kfree() called from free_request_key_auth() and key_revoke() while the request-key helper (here nfsidmap) instantiates a key:

stack segment: 0000 [#1] SMP PTI
CPU: 2 PID: 163271 Comm: nfsidmap Kdump: loaded Not tainted 4.18.0-553.166.1.el8_10.x86_64 #1
RIP: 0010:kfree+0x66/0x250
Call Trace:
 ? __die_body+0x1a/0x60
 ? die+0x2a/0x50
 ? do_trap+0xe7/0x110
 ? do_stack_segment+0x21/0x30
 ? stack_segment+0x1e/0x30
 ? free_request_key_auth.part.4+0x32/0x50
 ? kfree+0x66/0x250
 ? generic_file_buffered_read+0x84e/0xbb0
 free_request_key_auth.part.4+0x32/0x50
 key_revoke+0x3a/0x80
 __key_instantiate_and_link+0xb5/0x140
 key_instantiate_and_link+0x16a/0x180
  • The bug can also cause a hang instead of a panic. In one such case, one CPU reports Soft lockups or Hard lockups spinning on a socket's sk_lock.slock, and the holder is the other CPU, stuck with interrupts off in tcp_v4_rcv() -> sock_def_readable() -> __wake_up_common_lock() on the same socket's sk->sk_wq->wait.lock, a kmalloc-64 object.
[ 4215.465256] watchdog: BUG: soft lockup - CPU#0 stuck for 22s! [swapper/0:0]
[ 4220.598257] Kernel panic - not syncing: NMI: Not continuing

PID: 0        TASK: ffff8e3780c74000  CPU: 1    COMMAND: "swapper/1"
    [exception RIP: native_queued_spin_lock_slowpath+0x24]
    RIP: ffffffffbc15ee64  RSP: ffffb6f90001cae8  RFLAGS: 00000086
    RAX: 00000000c0000000  RBX: 0000000000000246  RCX: 0000000000000002
    RDX: 0000000000000001  RSI: 00000000c0000000  RDI: ffff8e36c56c5ec0 <<-----
 #5 [ffffb6f90001cae8] native_queued_spin_lock_slowpath at ffffffffbc15ee64
 #6 [ffffb6f90001cae8] _raw_spin_lock_irqsave at ffffffffbca28764
 #7 [ffffb6f90001caf8] __wake_up_common_lock at ffffffffbc14def6
 #8 [ffffb6f90001cb68] sock_def_readable at ffffffffbc82b2e7
 #9 [ffffb6f90001cb78] tcp_rcv_established at ffffffffbc90e990
#10 [ffffb6f90001cbb8] tcp_v4_do_rcv at ffffffffbc91b4b7
#11 [ffffb6f90001cbd8] tcp_v4_rcv at ffffffffbc91dab6

crash> px &((struct sock *)0xffff8e37a1810000)->sk_wq->wait->lock
$5 = (spinlock_t *) 0xffff8e36c56c5ec0

crash> kmem 0xffff8e36c56c5ec0
CACHE             OBJSIZE  ALLOCATED     TOTAL  SLABS  SSIZE  NAME
ffff8e37800028c0       64     120483    121728   1902     4k  kmalloc-64
  SLAB              MEMORY            NODE  TOTAL  ALLOCATED  FREE
  ffffeb6c4115b140  ffff8e36c56c5000     0     64         59     5
  FREE / [ALLOCATED]
  [ffff8e36c56c5ec0]
  • The characteristic is the spinlock word value 0xc0000000 (INT_MIN / 2) written by the x86 refcount exception handler ex_handler_refcount() when the second put of the double put in request_key_auth_destroy() decrements the unlocked spinlock of a socket_wq that reused the freed request_key_auth address, from 0 to -1.
crash> px ((struct sock *)0xffff8e37a1810000)->sk_wq->wait->lock->rlock->raw_lock
$9 = {
  {
    val = {
      counter = 0xc0000000 <<-----
    },
    {
      locked = 0x0,
      pending = 0x0
    },
    {
      locked_pending = 0x0,
      tail = 0xc000
    }
  }
}

Crash pattern with slub_debug

  • The kernel panics or reports memory corruption in the kmalloc-64 slab cache.
  • With slub_debug enabled, a "Poison overwritten" report with the allocation point request_key_auth_new() and the free point keyctl_instantiate_key_common() is printed, followed by a list_del corruption panic.
BUG kmalloc-64 (Not tainted): Poison overwritten
0x00000000be29508a-0x00000000be29508a @offset=7232. First byte 0x69 instead of 0x6b
Allocated in request_key_auth_new+0x5d/0x1f0 age=39 cpu=1 pid=29322
    request_key_auth_new+0x5d/0x1f0
    request_key_and_link+0x2dc/0x6f0
    request_key+0x3c/0x80
    nfs_idmap_get_key+0x12f/0x1e0 [nfsv4]
    nfs_idmap_lookup_id+0x30/0x80 [nfsv4]
    nfs_map_group_to_gid+0x11e/0x140 [nfsv4]
Freed in keyctl_instantiate_key_common+0x140/0x1a0 age=36 cpu=0 pid=29338
    keyctl_instantiate_key_common+0x140/0x1a0
    keyctl_instantiate_key+0x4d/0x80
    do_syscall_64+0x5b/0x1d0
    entry_SYSCALL_64_after_hwframe+0x66/0xcb
FIX kmalloc-64: Restoring 0x00000000be29508a-0x00000000be29508a=0x6b
FIX kmalloc-64: Marking all objects used
list_del corruption, fffff924c6452a08->next is LIST_POISON1 (dead000000000100)
kernel BUG at lib/list_debug.c:47!
RIP: 0010:__list_del_entry_valid.cold.1+0x12/0x48
Call Trace:
 free_debug_processing+0x390/0x4c0
 kfree+0x22e/0x250
Kernel panic - not syncing: Fatal exception in interrupt

Environment

  • Red Hat Enterprise Linux 8.10.z with affected version of kernel:

  • Any system that uses request-key upcalls, for example NFSv4 with id mapping (nfsidmap), CIFS with Kerberos, the kernel DNS resolver, or applications that call request_key()

Subscriber exclusive content

A Red Hat subscription provides unlimited access to our knowledgebase, tools, and much more.

Current Customers and Partners

Log in for full access

Log In

New to Red Hat?

Learn more about Red Hat subscriptions

Using a Red Hat product through a public cloud?

How to access this content