why vllm inference fails with Nvidia B300 gpu with cuda error?

Solution Verified - Updated -

Issue

  • why vllm inference fails with Nvidia B300 gpu with cuda error?
>>> print(torch.cuda.is_available())
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:180: UserWarning: CUDA initialization: Unexpected error from cudaGetDeviceCount(). Did you run some cuda functions before calling NumCudaDevices() that might have already set an error? Error 802: system not yet initialized (Triggered internally at /opt/pytorch/pytorch/c10/cuda/CUDAFunctions.cpp:119.)

Environment

  • RHELAI 3.3.
  • Nvidia B300 GPU.
  • RHAIIS.

Subscriber exclusive content

A Red Hat subscription provides unlimited access to our knowledgebase, tools, and much more.

Current Customers and Partners

Log in for full access

Log In

New to Red Hat?

Learn more about Red Hat subscriptions

Using a Red Hat product through a public cloud?

How to access this content