https://project.feishu.cn/dynamia-rd/story/detail/7103930628
What would you like to be added:
Add ubs-virt-enpu (vCANN-RT) as an opt-in Ascend sharing backend alongside hami-vnpu-core and template-based hardware vNPU, with optional integration for the upstream mem-swap feature and metrics in HAMi's existing Ascend monitoring endpoint.
The proposed scope is:
- Select the backend with
huawei.com/vnpu-mode: enpu, reuse existing Ascend memory/core resources, and support fixed-share, elastic, and best-effort policies.
- Generate per-container runtime configuration for the physical DIE selected by HAMi and inject
libvruntime.so, enpu-monitor, and ld.so.preload following the upstream container configuration. Keep runtimeClassName: ascend and expose only the selected device by default.
- Build the ordinary-slicing runtime from official release
1.0.0 using the official build image, record provenance/checksums, and package the assets in the device-plugin image alongside the existing hami-vnpu-core assets.
- Add optional node-local
enpu-manager integration for mem-swap. Use the HAMi allocation as the memory reservation, allow a separate huawei.com/enpu-memory-limit, and validate the returned physical device, Pod/container identity, quotas, and policy before mounting the manager-generated configuration. Workloads should not have to construct these mappings manually.
- Export ENPU metrics through the existing device-plugin
:9395/metrics endpoint. Attribute DCMI process HBM usage to the current Pod/container through host PID and cgroup mappings, cross-check the selected physical DIE and UUID, and expose configured memory request/limit, compute share, scheduling policy, allocation identity, and collection status.
- Provide Helm/node configuration, ordinary-slicing and mem-swap Pod examples, bilingual monitoring instructions, an optional PodMonitor example, and links to upstream hardware/runtime prerequisites.
Why is this needed:
Users should be able to deploy ENPU workloads through normal HAMi scheduling and allocation. Centralizing runtime injection and manager configuration avoids mismatches between the device HAMi selects and the device/configuration used inside the workload. Optional memory oversubscription should preserve the same scheduler reservation contract. ENPU resource usage should also be observable through the same Ascend Prometheus endpoint, without requiring operators to run a command inside every workload.
Anything else we need to know?:
- Coordinated implementation PRs: ascend-device-plugin #141 and HAMi #3110. Deploy both changes together for admission, node capability discovery, scheduling, runtime allocation, and chart configuration.
- ENPU must be disabled by default. Existing hami-vnpu-core and hardware-template behavior must remain unchanged when it is disabled. ENPU and hami-core must not share a physical DIE, and ENPU workloads sharing a DIE must use compatible scheduling policies.
- Initial ENPU scope is one physical DIE per container. Target configurations are A2/910B and A3/910C; A3 requires independent-DIE mode. This does not imply validation on every device model.
- Official release
1.0.0 provides ordinary slicing, not mem-swap. Mem-swap integration must remain optional and require compatible upstream runtime/manager components; it must not silently replace the default release build with preview sources. Without oversubscription settings, memory request and limit remain equal. fixed-share must reject unequal request/limit values.
- No changes to ubs-virt's internal swap implementation are proposed here. Automatic cleanup of manager allocation records after Pod exit is not implemented in the current plugin PR; manager allocation lifecycle and failure recovery remain explicit follow-up work.
Monitoring scope and status:
The monitor integration has been pushed to device-plugin PR #141 in commit 850eda1 for review. Deployment of this update remains pending.
- Reuse
hami_vgpu_memory_used_bytes and hami_vgpu_memory_limit_bytes for container HBM usage/limit, and add ENPU memory reservation, configured compute-share, allocation-info, and collection-success metrics. Preserve the existing Pod/container/device labels.
- Follow the official monitor's DCMI memory-query semantics. The collector does not invoke or parse the
enpu-monitor CLI on each scrape. It uses the plugin's existing privileged context and a read-only host /proc mount at /host/proc; no new workload environment variables, hostPID, or pods/exec permission are required.
- Export AI Core utilization at physical-DIE scope only. Official release
1.0.0 has no HTTP endpoint or per-Pod compute-utilization output. Configured compute share is not measured utilization and is not a strict ceiling under every scheduling policy; a best-effort CLI quota of 0 must not be interpreted as a zero HAMi allocation.
- Report resident HBM in bytes. Mem-swap configurations may have different memory request and limit values, but CPU swap bytes and sleep state are not exported. Collection failures omit affected measurements and set the corresponding success metric to
0, instead of fabricating zero usage.
- Preserve hami-core-only and hardware-template collection behavior. When hami-core and ENPU are both enabled, shared metric descriptors are registered together and whole-device series are exported once.
- The PodMonitor is an optional manual example, with no default Prometheus Operator CRD dependency. Its namespaces, Helm release selectors, and Prometheus selectors must match the deployment.
Usage at this commit: English guide, 中文说明, PodMonitor example.
Upstream references: ubs-virt-enpu source, official vCANN-RT configuration guide.
What would you like to be added:
Add
ubs-virt-enpu(vCANN-RT) as an opt-in Ascend sharing backend alongsidehami-vnpu-coreand template-based hardware vNPU, with optional integration for the upstream mem-swap feature and metrics in HAMi's existing Ascend monitoring endpoint.The proposed scope is:
huawei.com/vnpu-mode: enpu, reuse existing Ascend memory/core resources, and supportfixed-share,elastic, andbest-effortpolicies.libvruntime.so,enpu-monitor, andld.so.preloadfollowing the upstream container configuration. KeepruntimeClassName: ascendand expose only the selected device by default.1.0.0using the official build image, record provenance/checksums, and package the assets in the device-plugin image alongside the existing hami-vnpu-core assets.enpu-managerintegration for mem-swap. Use the HAMi allocation as the memory reservation, allow a separatehuawei.com/enpu-memory-limit, and validate the returned physical device, Pod/container identity, quotas, and policy before mounting the manager-generated configuration. Workloads should not have to construct these mappings manually.:9395/metricsendpoint. Attribute DCMI process HBM usage to the current Pod/container through host PID and cgroup mappings, cross-check the selected physical DIE and UUID, and expose configured memory request/limit, compute share, scheduling policy, allocation identity, and collection status.Why is this needed:
Users should be able to deploy ENPU workloads through normal HAMi scheduling and allocation. Centralizing runtime injection and manager configuration avoids mismatches between the device HAMi selects and the device/configuration used inside the workload. Optional memory oversubscription should preserve the same scheduler reservation contract. ENPU resource usage should also be observable through the same Ascend Prometheus endpoint, without requiring operators to run a command inside every workload.
Anything else we need to know?:
1.0.0provides ordinary slicing, not mem-swap. Mem-swap integration must remain optional and require compatible upstream runtime/manager components; it must not silently replace the default release build with preview sources. Without oversubscription settings, memory request and limit remain equal.fixed-sharemust reject unequal request/limit values.Monitoring scope and status:
The monitor integration has been pushed to device-plugin PR #141 in commit
850eda1for review. Deployment of this update remains pending.hami_vgpu_memory_used_bytesandhami_vgpu_memory_limit_bytesfor container HBM usage/limit, and add ENPU memory reservation, configured compute-share, allocation-info, and collection-success metrics. Preserve the existing Pod/container/device labels.enpu-monitorCLI on each scrape. It uses the plugin's existing privileged context and a read-only host/procmount at/host/proc; no new workload environment variables,hostPID, orpods/execpermission are required.1.0.0has no HTTP endpoint or per-Pod compute-utilization output. Configured compute share is not measured utilization and is not a strict ceiling under every scheduling policy; a best-effort CLI quota of0must not be interpreted as a zero HAMi allocation.0, instead of fabricating zero usage.Usage at this commit: English guide, 中文说明, PodMonitor example.
Upstream references: ubs-virt-enpu source, official vCANN-RT configuration guide.