You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When a pod requesting Ascend vNPU resources fails during container creation (for example, due to an OCI runtime initialization error in ascend-docker-runtime or containerd shim failure), the dynamically created vNPU on the Ascend chip is not released.
Because containerd never successfully completes container creation (failed to create containerd task), it does not trigger the standard post-stop/delete hook for ascend-docker-runtime. Meanwhile, ascend-docker-runtime does not roll back the created vNPU upon creation failure. Consequently, the virtual device remains orphaned on the physical NPU with Container ID: 000000000000.
Over time or with repeated retries, these zombie vNPUs exhaust the physical AI Cores (e.g. 6 of 8 cores occupied on Ascend 310P), causing all subsequent pods requesting vNPU to fail with:
ascend-docker-runtime: [ERROR] check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22)
ascend-docker-runtime: [ERROR] [drvCreateVdevice]ioctl failed, devid(0), ret(8).
kubelet: Error: failed to create containerd task: failed to create shim task: OCI runtime create failed: ... /usr/local/Ascend/Ascend-Docker-Runtime/ascend-docker-runtime did not terminate successfully: exit status 1
Currently, ascend-device-plugin sets --enable_periodic_idle_vnpu_cleanup=false by default (deliberately done via commit 084b97b / Issue #86 to prevent prematurely destroying vNPUs of newly started running pods that haven't yet invoked NPU APIs). As a result, orphaned vNPUs are never cleaned up while the plugin is running, resulting in a persistent resource deadlock.
What did you expect to happen
OCI Runtime Rollback: If ascend-docker-runtime fails during container creation after provisioning a vNPU, it should roll back and destroy the newly created vNPU instead of leaving it on the chip.
Pod-Aware Reconciliation in Device Plugin: Instead of relying solely on driver-level IsContainerUsed == 0 (which causes premature destruction for valid workloads as noted in Discussion: vNPU cleanup policy #86), the device plugin should cross-reference with active Pod allocations on the node (e.g., via Kubelet pod-resources API or Pod status). A vNPU should only be considered an orphan and safely destroyed if both:
IsContainerUsed == 0 (no active container process using it), AND
No active Pod on the node (Pending, ContainerCreating, or Running) is allocated or requesting that physical card and template.
How to reproduce it
Deploy ascend-device-plugin on a node with Ascend 310P (8 AI Cores total).
Deploy a pod requesting vNPU (e.g., huawei.com/Ascend310P: 1, huawei.com/Ascend310P-memory: 6k or 8k).
Trigger a container startup failure during OCI runtime create (e.g., misconfigured /usr/bin/runc or containerd task failure).
Inspect the NPU status using npu-smi info -t info-vnpu -i <card_id> -c <chip_id>:
Notice that orphaned vNPUs (e.g., vir02, vir04) are listed with Container ID: 000000000000.
npu-smi info shows No running processes found in NPU.
Deploy a new pod requesting 4 AI Cores (vir04).
The new pod fails to start due to check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22).
Environment
OS / Distro: Linux (CentOS / openEuler / RHEL / Ubuntu)
Kubernetes / K3s Version: K3s (containerd CRI)
Hardware: Huawei Ascend 310P3 (8 AI Cores, 24GB memory)
Ascend Driver / Tool Version: npu-smi 26.1.1
Component: Project-HAMi/ascend-device-plugin
HAMi Slicing Mode: Template vNPU (vir02 / vir04)
Logs & Diagnostics
1. Container Runtime Log
Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [EVENT][dsmi_create_vdevice]The create-virtual-device details. (uid=0; devid=0; vdev_id=4294967295)
Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [ERROR][share_log_read_in_single_module]check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22)
Sep 15 14:09:05 node190 ascend-docker-runtime[987045]: [ERROR][drvCreateVdevice]ioctl failed, devid(0), ret(8).
Sep 15 14:09:05 node190 k3s[2924]: E0915 14:09:05.828642 2924 log.go:32] "StartContainer from runtime service failed" err="rpc error: code = Unknown desc = failed to create containerd task: failed to create shim task: OCI runtime create failed: unable to retrieve OCI runtime error (.../log.json: no such file or directory): /usr/local/Ascend/Ascend-Docker-Runtime/ascend-docker-runtime did not terminate successfully: exit status 1"
2. npu-smi showing zombie vNPUs with Container ID: 000000000000
+-------------------------------------------------------------------------------+
| NPU resource static info as follow: |
| Format:Free/Total NA: Currently, query is not supported. |
| AICORE Memory AICPU VPC VENC VDEC JPEGD JPEGE PNGD |
| GB |
|===============================================================================|
| 2/8 5/21 1/7 3/12 0/3 3/12 4/16 2/8 NA/NA |
+-------------------------------------------------------------------------------+
| Total number of vnpu: 2 |
+-------------------------------------------------------------------------------+
| Vnpu ID | Vgroup ID | Container ID | Status | Template Name |
+-------------------------------------------------------------------------------+
| 100 | 0 | 000000000000 | 0 | vir02 |
+-------------------------------------------------------------------------------+
| 101 | 1 | 000000000000 | 0 | vir04 |
+-------------------------------------------------------------------------------+
Analysis on Why Simple Approaches Are Insufficient
Why we cannot simply turn --enable_periodic_idle_vnpu_cleanup=true back on:
As established in Discussion: vNPU cleanup policy #86, IsContainerUsed == 0 is true during the normal startup window of valid pods (image pull, model loading, etc.). Re-enabling periodic cleanup with only IsContainerUsed == 0 as the criterion re-introduces the bug where healthy workloads have their vNPUs deleted from underneath them.
Why "cleanup on Allocate failure" does not work:
In the Kubelet Device Plugin lifecycle, PluginServer.Allocate() succeeds before container creation and returns the device environment variables. The failure occurs later during ascend-docker-runtime's execution within containerd. Therefore, Allocate() never observes this failure.
Suggested Solution Directions
Upstream OCI Runtime Fix: ascend-docker-runtime should implement transactional rollback on container creation errors, ensuring that any vNPU created during the failed create invocation is cleaned up before returning an exit error.
Pod-Aware / Kubelet-Resources Reconciliation in ascend-device-plugin:
The plugin already mounts /var/lib/kubelet/pod-resources. The periodic cleaner should cross-reference vNPUs against active Pods/containers on the node. Only vNPUs that have no corresponding Pod allocating that device/card should be eligible for destruction.
Grace Period / Multi-Tick Candidate Confirmation:
Track idle vNPUs across multiple reconciliation intervals (or require IsContainerUsed == 0 continuously for e.g. 10+ minutes) to ensure transient startup states are never mistaken for orphaned devices.
What happened
When a pod requesting Ascend vNPU resources fails during container creation (for example, due to an OCI runtime initialization error in
ascend-docker-runtimeor containerd shim failure), the dynamically created vNPU on the Ascend chip is not released.Because containerd never successfully completes container creation (
failed to create containerd task), it does not trigger the standard post-stop/delete hook forascend-docker-runtime. Meanwhile,ascend-docker-runtimedoes not roll back the created vNPU upon creation failure. Consequently, the virtual device remains orphaned on the physical NPU withContainer ID: 000000000000.Over time or with repeated retries, these zombie vNPUs exhaust the physical AI Cores (e.g. 6 of 8 cores occupied on Ascend 310P), causing all subsequent pods requesting vNPU to fail with:
Currently,
ascend-device-pluginsets--enable_periodic_idle_vnpu_cleanup=falseby default (deliberately done via commit 084b97b / Issue #86 to prevent prematurely destroying vNPUs of newly started running pods that haven't yet invoked NPU APIs). As a result, orphaned vNPUs are never cleaned up while the plugin is running, resulting in a persistent resource deadlock.What did you expect to happen
ascend-docker-runtimefails during container creation after provisioning a vNPU, it should roll back and destroy the newly created vNPU instead of leaving it on the chip.IsContainerUsed == 0(which causes premature destruction for valid workloads as noted in Discussion: vNPU cleanup policy #86), the device plugin should cross-reference with active Pod allocations on the node (e.g., via Kubelet pod-resources API or Pod status). A vNPU should only be considered an orphan and safely destroyed if both:IsContainerUsed == 0(no active container process using it), ANDPending,ContainerCreating, orRunning) is allocated or requesting that physical card and template.How to reproduce it
ascend-device-pluginon a node with Ascend 310P (8 AI Cores total).huawei.com/Ascend310P: 1,huawei.com/Ascend310P-memory: 6kor8k)./usr/bin/runcor containerd task failure).npu-smi info -t info-vnpu -i <card_id> -c <chip_id>:vir02,vir04) are listed withContainer ID: 000000000000.npu-smi infoshowsNo running processes found in NPU.vir04).check aicore num failed, phy_id:0, core_need:4, used:6, total:8 Failed to create vdev info. (dev_id=0; ret=-22).Environment
npu-smi 26.1.1Project-HAMi/ascend-device-pluginvir02/vir04)Logs & Diagnostics
1. Container Runtime Log
2.
npu-smishowing zombie vNPUs withContainer ID: 000000000000Analysis on Why Simple Approaches Are Insufficient
--enable_periodic_idle_vnpu_cleanup=trueback on:As established in Discussion: vNPU cleanup policy #86,
IsContainerUsed == 0is true during the normal startup window of valid pods (image pull, model loading, etc.). Re-enabling periodic cleanup with onlyIsContainerUsed == 0as the criterion re-introduces the bug where healthy workloads have their vNPUs deleted from underneath them.In the Kubelet Device Plugin lifecycle,
PluginServer.Allocate()succeeds before container creation and returns the device environment variables. The failure occurs later duringascend-docker-runtime's execution within containerd. Therefore,Allocate()never observes this failure.Suggested Solution Directions
ascend-docker-runtimeshould implement transactional rollback on container creation errors, ensuring that any vNPU created during the failedcreateinvocation is cleaned up before returning an exit error.ascend-device-plugin:The plugin already mounts
/var/lib/kubelet/pod-resources. The periodic cleaner should cross-reference vNPUs against active Pods/containers on the node. Only vNPUs that have no corresponding Pod allocating that device/card should be eligible for destruction.Track idle vNPUs across multiple reconciliation intervals (or require
IsContainerUsed == 0continuously for e.g. 10+ minutes) to ensure transient startup states are never mistaken for orphaned devices.