Skip to content

Polygraphy: find the CUDA runtime from CUDA 13+ pip wheels - #4871

Open
LinCanNerd wants to merge 1 commit into
NVIDIA:mainfrom
LinCanNerd:polygraphy-find-cuda13-runtime-wheel
Open

LinCanNerd wants to merge 1 commit into
NVIDIA:mainfrom
LinCanNerd:polygraphy-find-cuda13-runtime-wheel

Conversation

@LinCanNerd

Copy link
Copy Markdown

What does this PR do?

On a system whose CUDA runtime comes only from pip (no CUDA Toolkit, no LD_LIBRARY_PATH entry), every TensorRT runner in Polygraphy fails:

$ polygraphy run model.onnx --trt
...
OSError: libcudart.so: cannot open shared object file: No such file or directory

_cuda_runtime_wheel_dirs only searches the CUDA 12 wheel location, <site-packages>/nvidia/cuda_runtime/{lib,bin}. The CUDA 13 runtime wheel (nvidia-cuda-runtime 13.x; the -cu13 name is deprecated, see #4614) uses a different layout. From the published nvidia-cuda-runtime==13.0.96 wheels:

Wheel CUDA runtime path
manylinux2014_x86_64, manylinux2014_aarch64 nvidia/cu13/lib/libcudart.so.13
win_amd64 nvidia/cu13/bin/x86_64/cudart64_13.dll
(CUDA 12, nvidia-cuda-runtime-cu12) nvidia/cuda_runtime/{lib,bin}

Changes in polygraphy/cuda/cuda.py:

  • Also search existing nvidia/cu<major> package directories (globbed as cu[0-9]*, so other nvidia/cu* packages such as cuBLAS are not matched). They are appended after the current locations, so environments that work today resolve the same library.
  • On Windows, also search the wheel's bin/x86_64 subdirectory.
  • The missing-runtime hint now names nvidia-cuda-runtime (CUDA 13+) and nvidia-cuda-runtime-cu12 instead of the deprecated nvidia-cuda-runtime-cuXX.

Added TestFindCudaLibDirs::test_cuda13_runtime_wheel_dirs_included (Linux and Windows) and a CHANGELOG entry.

Testing

On NVIDIA Jetson AGX Thor (SM110, JetPack 7.2 / L4T R39.2, no CUDA Toolkit installed), TensorRT 11.3.0.99 (tensorrt-cu13), nvidia-cuda-runtime==13.0.96, Polygraphy from main, Python 3.12, LD_LIBRARY_PATH unset:

Before After
polygraphy run model.onnx --trt --onnxrt OSError: libcudart.so: cannot open shared object file Finds nvidia/cu13/lib/libcudart.so.13; outputs match (PASSED)
pytest tests/cuda/test_cuda.py 17 failed, 14 passed (CUDA tests fail with the OSError) 30 passed, 1 failed

The one remaining failure is TestDeviceBuffer::test_copy_from_overhead, which is marked flaky and asserts copy_from is within 12% of a raw memcpy; on this device it measures 1.2–2.2x across runs. It could not run at all before this change and is unrelated to it.

black --check passes on the changed files.

The CUDA 13 runtime wheel (`nvidia-cuda-runtime`) installs libcudart under
`<site-packages>/nvidia/cu13/lib` on Linux and `nvidia/cu13/bin/x86_64` on
Windows, while Polygraphy only searched the CUDA 12 wheel location
`nvidia/cuda_runtime/{lib,bin}`. On systems whose CUDA runtime comes only
from pip, such as a Jetson AGX Thor without a CUDA Toolkit, every TensorRT
runner failed with `OSError: libcudart.so: cannot open shared object file`.

Also search `nvidia/cu<major>` wheel directories, after the existing
locations, and update the missing-runtime hint to name the current wheels.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: LinCanNerd <lincanecdl@gmail.com>
@LinCanNerd
LinCanNerd requested a review from a team as a code owner October 6, 2026 08:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants