Skip to content

[WIP] fix(quickstart): specify model precision before engine building - #11

Open
yenhao-huang wants to merge 2 commits into
mainfrom
fix/4581-intro-strong-typing
Open

yenhao-huang wants to merge 2 commits into
mainfrom
fix/4581-intro-strong-typing

Conversation

@yenhao-huang

Copy link
Copy Markdown
Owner

IntroNotebooks still uses BuilderFlag.FP16, which raises AttributeError on TensorRT 11, and its PyTorch notebook exports FP32 before requesting FP16 through removed builder options. Migrate the helper to a strongly typed network and make precision explicit in ONNX. Retain fp16_mode through graph conversion with preserved I/O types; use False for already typed models.

The notebook now selects dtype before export, uses consistent normalization, allocates from engine I/O metadata, rejects mismatched inputs, and compares logits with framework references. Disable TF32 for the FP32 reference comparison. Document explicit conversion's accuracy and performance implications, dependencies, and selective mixed precision.

Validation: 5 GPU/error-path tests passed on each of TensorRT 10.16 and 11.2 (11.2 baseline: 3 failed, 2 passed). Pretrained ResNet50 on TRT 11.2 passed FP32/FP16 comparisons at batch 1 and the notebook's batch 32, including the default conversion path and FP16 versus FP32 reference. Batch-32 top-1 predictions all matched. Executed export, preprocessing, allocation, inference, and cleanup notebook cells; Torch references ran on CPU, and Jupyter timing/Torch CUDA cells were not run as a complete notebook. Black, notebook syntax, and git diff --check passed. Report: docs/howard/4581.md.

Addresses NVIDIA#4581. The separate ONNXClassifierWrapper runtime issue covered by NVIDIA#4612/NVIDIA#4613 is outside this precision migration.

Signed-off-by: yenhao <46972327+yenhao-huang@users.noreply.github.com>
Signed-off-by: yenhao <46972327+yenhao-huang@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant