Skip to content

[Bug]: Zombie vNPU leak on container creation failure leads to "check aicore num failed" and pod scheduling deadlock #138

Description

@xrwang8

What happened

When a pod requesting Ascend vNPU resources fails during container creation (for example, due to an OCI runtime initialization error in ascend-docker-runtime or containerd shim failure), the dynamically created vNPU on the Ascend chip is not released.

Because containerd never successfully completes container creation (failed to create containerd task), it does not trigger the standard post-stop/delete hook for ascend-docker-runtime. Meanwhile, ascend-docker-runtime does not roll back the created vNPU upon creation failure. Consequently, the virtual device remains orphaned on the physical NPU with Container ID: 000000000000.

Over time or with repeated retries, these zombie vNPUs exhaust the physical AI Cores (e.g. 6 of 8 cores occupied on Ascend 310P), causing all subsequent pods requesting vNPU to fail with:

ascend-docker-runtime: [ERROR] check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22)
ascend-docker-runtime: [ERROR] [drvCreateVdevice]ioctl failed, devid(0), ret(8).
kubelet: Error: failed to create containerd task: failed to create shim task: OCI runtime create failed: ... /usr/local/Ascend/Ascend-Docker-Runtime/ascend-docker-runtime did not terminate successfully: exit status 1

Currently, ascend-device-plugin sets --enable_periodic_idle_vnpu_cleanup=false by default (deliberately done via commit 084b97b / Issue #86 to prevent prematurely destroying vNPUs of newly started running pods that haven't yet invoked NPU APIs). As a result, orphaned vNPUs are never cleaned up while the plugin is running, resulting in a persistent resource deadlock.


What did you expect to happen

  1. OCI Runtime Rollback: If ascend-docker-runtime fails during container creation after provisioning a vNPU, it should roll back and destroy the newly created vNPU instead of leaving it on the chip.
  2. Pod-Aware Reconciliation in Device Plugin: Instead of relying solely on driver-level IsContainerUsed == 0 (which causes premature destruction for valid workloads as noted in Discussion: vNPU cleanup policy #86), the device plugin should cross-reference with active Pod allocations on the node (e.g., via Kubelet pod-resources API or Pod status). A vNPU should only be considered an orphan and safely destroyed if both:
    • IsContainerUsed == 0 (no active container process using it), AND
    • No active Pod on the node (Pending, ContainerCreating, or Running) is allocated or requesting that physical card and template.

How to reproduce it

  1. Deploy ascend-device-plugin on a node with Ascend 310P (8 AI Cores total).
  2. Deploy a pod requesting vNPU (e.g., huawei.com/Ascend310P: 1, huawei.com/Ascend310P-memory: 6k or 8k).
  3. Trigger a container startup failure during OCI runtime create (e.g., misconfigured /usr/bin/runc or containerd task failure).
  4. Inspect the NPU status using npu-smi info -t info-vnpu -i <card_id> -c <chip_id>:
    • Notice that orphaned vNPUs (e.g., vir02, vir04) are listed with Container ID: 000000000000.
    • npu-smi info shows No running processes found in NPU.
  5. Deploy a new pod requesting 4 AI Cores (vir04).
  6. The new pod fails to start due to check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22).

Environment

  • OS / Distro: Linux (CentOS / openEuler / RHEL / Ubuntu)
  • Kubernetes / K3s Version: K3s (containerd CRI)
  • Hardware: Huawei Ascend 310P3 (8 AI Cores, 24GB memory)
  • Ascend Driver / Tool Version: npu-smi 26.1.1
  • Component: Project-HAMi/ascend-device-plugin
  • HAMi Slicing Mode: Template vNPU (vir02 / vir04)

Logs & Diagnostics

1. Container Runtime Log

Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [EVENT][dsmi_create_vdevice]The create-virtual-device details. (uid=0; devid=0; vdev_id=4294967295)
Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [ERROR][share_log_read_in_single_module]check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22)
Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [ERROR][drvCreateVdevice]ioctl failed, devid(0), ret(8).
Sep 15 14:09:05 node190 k3s[2924]: E0915 14:09:05.828642 2924 log.go:32] "StartContainer from runtime service failed" err="rpc error: code = Unknown desc = failed to create containerd task: failed to create shim task: OCI runtime create failed: unable to retrieve OCI runtime error (.../log.json: no such file or directory): /usr/local/Ascend/Ascend-Docker-Runtime/ascend-docker-runtime did not terminate successfully: exit status 1"

2. npu-smi showing zombie vNPUs with Container ID: 000000000000

+-------------------------------------------------------------------------------+
| NPU resource static info as follow:                                           |
| Format:Free/Total                   NA: Currently, query is not supported.    |
| AICORE    Memory    AICPU    VPC    VENC    VDEC    JPEGD    JPEGE    PNGD    |
|            GB                                                                 |
|===============================================================================|
| 2/8       5/21      1/7      3/12   0/3     3/12    4/16     2/8      NA/NA   |
+-------------------------------------------------------------------------------+
| Total number of vnpu: 2                                                       |
+-------------------------------------------------------------------------------+
|  Vnpu ID  |  Vgroup ID     |  Container ID  |  Status  |  Template Name       |
+-------------------------------------------------------------------------------+
|  100      |  0             |  000000000000  |  0       |  vir02               |
+-------------------------------------------------------------------------------+
|  101      |  1             |  000000000000  |  0       |  vir04               |
+-------------------------------------------------------------------------------+

Analysis on Why Simple Approaches Are Insufficient

  1. Why we cannot simply turn --enable_periodic_idle_vnpu_cleanup=true back on:
    As established in Discussion: vNPU cleanup policy #86, IsContainerUsed == 0 is true during the normal startup window of valid pods (image pull, model loading, etc.). Re-enabling periodic cleanup with only IsContainerUsed == 0 as the criterion re-introduces the bug where healthy workloads have their vNPUs deleted from underneath them.
  2. Why "cleanup on Allocate failure" does not work:
    In the Kubelet Device Plugin lifecycle, PluginServer.Allocate() succeeds before container creation and returns the device environment variables. The failure occurs later during ascend-docker-runtime's execution within containerd. Therefore, Allocate() never observes this failure.

Suggested Solution Directions

  1. Upstream OCI Runtime Fix:
    ascend-docker-runtime should implement transactional rollback on container creation errors, ensuring that any vNPU created during the failed create invocation is cleaned up before returning an exit error.
  2. Pod-Aware / Kubelet-Resources Reconciliation in ascend-device-plugin:
    The plugin already mounts /var/lib/kubelet/pod-resources. The periodic cleaner should cross-reference vNPUs against active Pods/containers on the node. Only vNPUs that have no corresponding Pod allocating that device/card should be eligible for destruction.
  3. Grace Period / Multi-Tick Candidate Confirmation:
    Track idle vNPUs across multiple reconciliation intervals (or require IsContainerUsed == 0 continuously for e.g. 10+ minutes) to ensure transient startup states are never mistaken for orphaned devices.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions