[WIP] fix(quickstart): specify model precision before engine building - #11
Open
yenhao-huang wants to merge 2 commits into
Open
yenhao-huang wants to merge 2 commits into
yenhao-huang wants to merge 2 commits into
Conversation
Signed-off-by: yenhao <46972327+yenhao-huang@users.noreply.github.com>
Signed-off-by: yenhao <46972327+yenhao-huang@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
IntroNotebooks still uses BuilderFlag.FP16, which raises AttributeError on TensorRT 11, and its PyTorch notebook exports FP32 before requesting FP16 through removed builder options. Migrate the helper to a strongly typed network and make precision explicit in ONNX. Retain fp16_mode through graph conversion with preserved I/O types; use False for already typed models.
The notebook now selects dtype before export, uses consistent normalization, allocates from engine I/O metadata, rejects mismatched inputs, and compares logits with framework references. Disable TF32 for the FP32 reference comparison. Document explicit conversion's accuracy and performance implications, dependencies, and selective mixed precision.
Validation: 5 GPU/error-path tests passed on each of TensorRT 10.16 and 11.2 (11.2 baseline: 3 failed, 2 passed). Pretrained ResNet50 on TRT 11.2 passed FP32/FP16 comparisons at batch 1 and the notebook's batch 32, including the default conversion path and FP16 versus FP32 reference. Batch-32 top-1 predictions all matched. Executed export, preprocessing, allocation, inference, and cleanup notebook cells; Torch references ran on CPU, and Jupyter timing/Torch CUDA cells were not run as a complete notebook. Black, notebook syntax, and git diff --check passed. Report: docs/howard/4581.md.
Addresses NVIDIA#4581. The separate ONNXClassifierWrapper runtime issue covered by NVIDIA#4612/NVIDIA#4613 is outside this precision migration.