Skip to content

Engine build failure of TensorRT 10.14.1 and 11.3.0 when running S2M2 stereo matching above ~1 MPix input on Jetson AGX Thor (sm_110) #4866

Description

@george-auradine

Description

On Jetson AGX Thor (sm_110), TensorRT cannot build an engine for a stereo
matching network above roughly 1 megapixel of input. The same ONNX graph builds
and runs correctly at smaller resolutions.

The failure reproduces across a major version boundary with two different
symptoms:

TensorRT behaviour above ~1 MPix
10.14.1.48 SIGSEGV inside the Myelin compiler, ~2 min into the build
11.3.0.99 build runs to completion (285–420 s), reports success, produces a 0 MiB engine, then Assertion failure: false && "Attempting to access an empty engine!"

In both versions TensorRT fuses the entire network into a single Myelin
ForeignNode. 10.14.1 crashes while compiling it; 11.3.0 completes without
error but emits nothing.

The 11.3.0 behaviour is the more dangerous of the two: a build pipeline that
checks only the builder's exit status sees success and writes an empty engine.
The failure surfaces later, at load or inference time, far from its cause.

This is not resource exhaustion. ~116 GB of 122 GB were free throughout,
and workspace caps of 2/8/24 GB all fail identically.

Failing output, 11.3.0:

[I] Created engine with size: 0 MiB
[I] Engine built in 285.128 sec.
[E] Assertion failure: false && "Attempting to access an empty engine!"

Failing output, 10.14.1 (verbose, final lines before SIGSEGV):

[V] [TRT] After concat removal: 1 layers
[V] [TRT] Graph optimization time: 0.15853 seconds.
[V] [TRT] Building graph using backend strategy 2
[V] [TRT] =============== Computing costs for
          {ForeignNode[node_convert_element_type_default_1...output_occ_castOut]}
[V] [TRT] --------------- Timing Runner:
          {ForeignNode[...]} (Myelin[0x80000...])
[I] [TRT] Compiler backend is used during engine build.
<SIGSEGV, exit 139>

Environment

TensorRT Version: 10.14.1.48 and 11.3.0.99 (both affected)

NVIDIA GPU: Jetson AGX Thor Developer Kit (T5000), sm_110, MAXN power mode

NVIDIA Driver Version: 580.00

CUDA Version: 13.0

CUDNN Version: as shipped in the containers below

Operating System: JetPack R38 (release) REVISION 4.0, GCID 43443517, aarch64

Python Version (if applicable): 3.12

PyTorch Version (if applicable): 2.10.0a0+b4e4ee81d3.nv25.12

Baremetal or Container (if so, version):

  • nvcr.io/nvidia/pytorch:25.12-py3 → TensorRT 10.14.1.48
  • nvcr.io/nvidia/pytorch:26.09-py3 → TensorRT 11.3.0.99

Note: the stock JetPack R38.4 apt repo pins TensorRT to 10.13.3.9, which is
older than both versions tested, so default Jetson installs are likely affected
as well. There is no newer TensorRT reachable from Jetson via apt (pinned to the
JetPack release) or PyPI (tensorrt-cu13-libs publishes no aarch64 wheels —
only the Python bindings).

Relevant Files

Model link: S2M2 stereo matching, "S" variant (26.5 M parameters)

Graph: 3,123 ONNX nodes, 780 initializers. Two image inputs [1,3,H,W], three
outputs (disparity, occlusion, confidence).

I can attach the full --verbose build log (101,824 lines, 285 KB gzipped) and
~28 log files covering every control described below — happy to upload on
request or attach here.

Steps To Reproduce

Commands or scripts:

# 1. clone and fetch the S variant weights
git clone https://github.com/junhong-3dv/s2m2 && cd s2m2
mkdir -p weights/pretrain_weights
curl -L -o weights/pretrain_weights/CH128NTR1.pth \
  https://huggingface.co/minimok/s2m2/resolve/main/CH128NTR1.pth

# 2. export ONNX at a failing resolution
python demo/export_onnx.py --model_type S --img_width 1536 --img_height 960

# 3a. TensorRT 10.14.1 -> SIGSEGV (exit 139)
trtexec --onnx=weights/onnx_save/S2M2_S_1536_960_v2_torch21.onnx \
        --saveEngine=out.engine --fp16

# 3b. TensorRT 11.3.0 -> 0 MiB engine + assertion
#     (--fp16 was removed in 11.x; networks are strongly typed from the ONNX)
trtexec --onnx=weights/onnx_save/S2M2_S_1536_960_v2_torch21.onnx \
        --saveEngine=out.engine

Works at --img_width 1280 --img_height 800; fails at 1536 × 960 and above.

Resolution boundary (same pipeline, only dimensions varied):

Resolution Pixels TRT 10.14.1 TRT 11.3.0
960 × 608 0.584 MPix builds, 61.85 ms fp16 builds, 55.48 ms fp16
1280 × 800 1.024 MPix builds, 124.36 ms fp16 not tested
1536 × 960 1.475 MPix SIGSEGV 0 MiB engine
1920 × 1216 2.335 MPix SIGSEGV 0 MiB engine

The boundary lies between 1.024 and 1.475 megapixels.

Ruled out by controlled test:

Hypothesis Test Result
Workspace / memory --memPoolSize=workspace: 2048, 8192, 24576 MB all SIGSEGV identically
Precision fp32 and fp16, both versions identical failure — 4 combinations
The workspace flag itself 960 × 608 with workspace:8192 builds fine, 61.54 ms
System memory free during build ~116 GB available throughout
Export pipeline same script, smaller resolutions builds and runs correctly
Builder optimisation levels 0, 1, 2, 3 1–3 SIGSEGV; level 0 builds then fails at enqueueV3

Control — the working resolution fuses identically. At 960 × 608 the
verbose log shows the same whole-graph fusion (1 layers, 1 ForeignNode,
1 Timing Runner) and builds successfully. Input dimensions are the only
variable between success and failure.

Attempted workaround that does NOT work. Marking intermediate tensors as
additional graph outputs (fully typed via ONNX shape inference;
onnx.checker.check_model passes) to force graph splits still fails — tried
with 2 and 5 cut points, both SIGSEGV under 10.14.1.

Second failure mode at --builderOptimizationLevel=0 (10.14.1): the engine
builds, then fails at execution:

[E] Error[1]: IExecutionContext::enqueueV3: Error Code 1:
    Myelin ([cask.cpp:2974: exec] Platform (Cuda) error
    In executeMyelinGraph at /_src/runtime/myelin/runner.cpp:820)

Have you tried the latest release?: Yes — TensorRT 11.3.0.99 via
nvcr.io/nvidia/pytorch:26.09-py3. The crash becomes a silent 0 MiB engine but
the build still fails at the same resolution boundary.

Can this model run on other frameworks?: Yes. torch.compile runs the same
model at all resolutions on the same hardware — 73.87 ms at 960 × 608 and
399.34 ms at 1920 × 1216. Where TensorRT does build it is faster (55.48 ms at
960 × 608, 11.3.0 fp16), which is why the resolution ceiling matters.

Questions

  1. Is there a supported way to prevent whole-graph Myelin fusion so the compiler
    handles smaller subgraphs? Marking intermediate tensors as outputs does not
    work.
  2. Is there a documented input-size limit for Myelin-compiled ForeignNodes on
    sm_110?
  3. In 11.3.0, why does the builder report success and emit a 0 MiB engine rather
    than raising a build error?
  4. Is a fix expected in a release that reaches Jetson through JetPack? Neither
    apt nor PyPI currently offers anything newer than 10.13.3.9 for aarch64.

Activity

  1. george-auradine commented on Oct 4, 2026

    @george-auradine
    Author
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions