Skip to content

#4866 - Fail trtexec engine build when buildSerializedNetworkToStream fails - #4867

Open
SammyTourani wants to merge 1 commit into
NVIDIA:mainfrom
SammyTourani:fix/issue-4866
Open

SammyTourani wants to merge 1 commit into
NVIDIA:mainfrom
SammyTourani:fix/issue-4866

Conversation

@SammyTourani

Copy link
Copy Markdown

Fixes #4866

Related to #4866 (question 3, the "0 MiB engine" output). The engine build failure itself happens inside TensorRT and is
not addressed here.

With --saveEngine, trtexec builds through IBuilder::buildSerializedNetworkToStream() in networkToSerializedEngine()
(samples/common/sampleEngines.cpp), but ignored the bool it returns. When the build failed, trtexec logged
Created engine with size: 0 MiB and Engine built in ... sec. after the builder's error. It then exited on
Assertion failure: false && "Attempting to access an empty engine!" while saving the engine.

The return value is now checked with SMP_RETVAL_IF_FALSE. The error message is the one the buildSerializedNetwork()
path in the same function already uses (Engine could not be created from network). A failed build now goes through
trtexec's normal failure path:

[E] Error[10]: IBuilder::buildSerializedNetworkToStream: Error Code 10: Internal Error (...)
[E] Engine could not be created from network
[E] Building engine failed
[E] Failed to create engine from model or file.
[E] Engine set up failed

The exit status was already non-zero before this change, because the assertion calls exit(EXIT_FAILURE). What changes
is that trtexec no longer reports a 0 MiB engine as built and no longer ends on an internal assertion. Successful builds
are unaffected.

Verification

No GPU was available, so this was not tested against a real TensorRT build failure. Instead I built samples/common with
clang (C++20, TRT_BUILD_ONNX_PARSER=1) against fake builder, network, config and ONNX parser objects. A small driver,
not part of this PR, calls sample::getEngineBuildEnv() with the options trtexec --onnx=model.onnx --saveEngine=<file>
sets, and handles a failure the same way trtexec's main() does. The fake buildSerializedNetworkToStream() works in
one of two modes:

The driver was built once from 98adec8 (before) and once from this commit (after). Log timestamps are removed below;
the output is otherwise verbatim.

Command: build-before/harness fail before-fail.engine
Result:

[I] Start parsing network model.
[I] Finished parsing network model. Parse time: 5e-07
[E] Error[10]: IBuilder::buildSerializedNetworkToStream: Error Code 10: Internal Error (Could not find any implementation for node {ForeignNode[ones...node_sum_11]}.)
[I] Created engine with size: 0 MiB
[I] Engine built in 0.000109459 sec.
[E] Assertion failure: false && "Attempting to access an empty engine!"
exit status: 1

Command: build-after/harness fail after-fail.engine
Result:

[I] Start parsing network model.
[I] Finished parsing network model. Parse time: 7.92e-07
[E] Error[10]: IBuilder::buildSerializedNetworkToStream: Error Code 10: Internal Error (Could not find any implementation for node {ForeignNode[ones...node_sum_11]}.)
[E] Engine could not be created from network
[E] Building engine failed
[E] Failed to create engine from model or file.
[E] Engine set up failed
exit status: 1

Command: build-before/harness succeed before-succeed.engine and build-after/harness succeed after-succeed.engine
Result: identical before and after. Both print Created engine with size: 3.00002 MiB and Engine built in ... sec.,
save a 3145745-byte engine file, and exit with status 0.

Command: build-tests-before/trt_samples_common_test and build-tests-after/trt_samples_common_test. This is the
samples/common gtest suite from samples/common/CMakeLists.txt, built the same way.
Result: [==========] 81 tests from 12 test suites ran. / [ PASSED ] 81 tests. both before and after. None of these
tests exercise the build path in sampleEngines.cpp, and the gtest binary does not link getEngineBuildEnv at all. This
run only shows that nothing else in samples/common broke.

clang-format -style=file (23.1.2) reports no changes on the edited lines.

…Stream fails

With --saveEngine, trtexec builds through
IBuilder::buildSerializedNetworkToStream() but ignored its return value.
When the build failed, trtexec still logged "Created engine with size:
0 MiB" and "Engine built in ... sec.", then exited on the assertion
"Attempting to access an empty engine!" while saving the engine.

Check the return value and fail with the same error that the
buildSerializedNetwork() path reports.

Signed-off-by: Sammy Tourani <sammytourani@gmail.com>
@SammyTourani
SammyTourani requested a review from a team as a code owner October 4, 2026 06:19
@george-auradine

Copy link
Copy Markdown

The assertion calls exit(EXIT_FAILURE), so the exit status was
already non-zero — our own runs recorded rc=1. @SammyTourani is right about
this in the linked PR.

The actual defect is that two [E] errors are followed by [I] Created engine with size: 0 MiB and [I] Engine built in 336.072 sec., which read as success,
and the process then ends on an internal assertion rather than a clean error
path. Still worth fixing. Thanks

@george-auradine

Copy link
Copy Markdown

The failing op is F.avg_pool2d in S2M2's CostVolume
(src/s2m2/core/model/submodules.py):

b, h, w, w2 = cv.shape
self.cv    = cv.reshape(b * h * w, 1, 1, w2)
self.cv_2x = F.avg_pool2d(self.cv, kernel_size=[1, 2])   # <-- fails here

The pooled tensor is [b*h*w, 1, 1, w2], so its batch dimension is the full
quarter-resolution pixel count
. That is what crosses a threshold — not the
image size directly:

input resolution ¼-res pooling batch dim vs 65,536 build
960 × 608 240 × 152 36,480 under builds
1280 × 800 320 × 200 64,000 1,536 under builds
1536 × 960 384 × 240 92,160 over fails
1920 × 1216 480 × 304 145,920 over fails

1280 × 800 sits 1,536 below 2¹⁶ and builds; the next size up exceeds it and
fails. This looks like a 16-bit limit on the pooling layer's batch
dimension
, not the "~1 MPix" threshold in my original title and description —
the megapixel correlation was incidental, since quarter-res pixel count scales
with image size.

The Myelin error in the 11.3.0 log is consistent with that:

MyelinCheckException: autotuner.cpp:3361:
CHECK(sorted_ids.size() > 0) failed. Must have costs
In compileGraph at /_src/optimizer/myelin/codeGenerator.cpp:1962

No viable tactic for the fused node containing that pooling op.

Workaround

Rewriting the pool as reshape + mean avoids the pooling layer and emits
ReduceMean:

# before
self.cv_2x = F.avg_pool2d(self.cv, kernel_size=[1, 2])

# after - numerically identical
_n, _c, _h, _w2 = self.cv.shape
self.cv_2x = self.cv.reshape(_n, _c, _h, _w2 // 2, 2).mean(dim=-1)

Verified bit-exact against avg_pool2d beforehand (torch.equal True, max abs
diff 0.0, including at the 145,920-batch shape).

All previously failing builds now succeed:

configuration before after engine
1536 × 960, TRT 10.14.1 fp16 SIGSEGV 187.29 ms 122 MiB
1920 × 1216, TRT 10.14.1 fp16 SIGSEGV 318.15 ms 172 MiB
1536 × 960, TRT 11.3.0 fp32 0 MiB engine 415.70 ms 155 MiB
1920 × 1216, TRT 11.3.0 fp32 0 MiB engine 723.93 ms 215 MiB

(11.3.0 builds are fp32 — no --fp16 flag, strongly typed from an fp32 ONNX.
The ~2.2× ratio to the fp16 builds is consistent with precision alone.)

The patch changes only the CostVolume pool. The exported graph still
contains 18 other AveragePool nodes — UNet and stacked-MRT downsampling on
ordinary [B,C,H,W] maps with batch 1–2 — and those build fine. So the problem
is pooling with a very large batch dimension, not AveragePool generally.

Remaining question

Is the ~65,536 batch limit on pooling layers expected and documented? If it is a
known constraint, a build-time error naming the layer and the offending
dimension would be far more useful than either failure mode. If it is not
expected, the reproducer above is deterministic and uses a public model.

@george-auradine

Copy link
Copy Markdown

@SammyTourani
Thanks for picking this up, and for being explicit that it addresses the
reporting rather than the build failure.

You noted no GPU was available. We hit this on Jetson AGX Thor (sm_110),
TensorRT 11.3.0, and the real log matches your fail mode exactly:

[E] Error[9]: Error Code: 9: Skipping tactic 0x0000000000000000 due to exception
    [myelin_graph.h:1292: attachExceptionMsgToGraph] MyelinCheckException:
    autotuner.cpp:3361: CHECK(sorted_ids.size() > 0) failed. Must have costs
    In compileGraph at /_src/optimizer/myelin/codeGenerator.cpp:1962
[E] Error[10]: IBuilder::buildSerializedNetworkToStream: Error Code 10:
    Internal Error (Could not find any implementation for node
    {ForeignNode[node_convert_element_type_default_1...node_sum_11]}.
    In computeCosts at /_src/optimizer/common/tactic/optimizer.cpp:4254)
[I] Created engine with size: 0 MiB
[I] Engine built in 336.072 sec.
[E] Assertion failure: false && "Attempting to access an empty engine!"

Same Error Code 10 from buildSerializedNetworkToStream, same 0 MiB line, same
assertion — so the path you patched is the one real failures take. The 336 s is
a genuine build attempt, which is part of why the following "Engine built in
336.072 sec." reads as success.

Happy to test a patched trtexec against a real failing build on Thor if that
is useful before merge — the reproducer is a public model and deterministic.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Engine build failure of TensorRT 10.14.1 and 11.3.0 when running S2M2 stereo matching above ~1 MPix input on Jetson AGX Thor (sm_110)

2 participants