Description
On Jetson AGX Thor (sm_110), TensorRT cannot build an engine for a stereo
matching network above roughly 1 megapixel of input. The same ONNX graph builds
and runs correctly at smaller resolutions.
The failure reproduces across a major version boundary with two different
symptoms:
| TensorRT |
behaviour above ~1 MPix |
| 10.14.1.48 |
SIGSEGV inside the Myelin compiler, ~2 min into the build |
| 11.3.0.99 |
build runs to completion (285–420 s), reports success, produces a 0 MiB engine, then Assertion failure: false && "Attempting to access an empty engine!" |
In both versions TensorRT fuses the entire network into a single Myelin
ForeignNode. 10.14.1 crashes while compiling it; 11.3.0 completes without
error but emits nothing.
The 11.3.0 behaviour is the more dangerous of the two: a build pipeline that
checks only the builder's exit status sees success and writes an empty engine.
The failure surfaces later, at load or inference time, far from its cause.
This is not resource exhaustion. ~116 GB of 122 GB were free throughout,
and workspace caps of 2/8/24 GB all fail identically.
Failing output, 11.3.0:
[I] Created engine with size: 0 MiB
[I] Engine built in 285.128 sec.
[E] Assertion failure: false && "Attempting to access an empty engine!"
Failing output, 10.14.1 (verbose, final lines before SIGSEGV):
[V] [TRT] After concat removal: 1 layers
[V] [TRT] Graph optimization time: 0.15853 seconds.
[V] [TRT] Building graph using backend strategy 2
[V] [TRT] =============== Computing costs for
{ForeignNode[node_convert_element_type_default_1...output_occ_castOut]}
[V] [TRT] --------------- Timing Runner:
{ForeignNode[...]} (Myelin[0x80000...])
[I] [TRT] Compiler backend is used during engine build.
<SIGSEGV, exit 139>
Environment
TensorRT Version: 10.14.1.48 and 11.3.0.99 (both affected)
NVIDIA GPU: Jetson AGX Thor Developer Kit (T5000), sm_110, MAXN power mode
NVIDIA Driver Version: 580.00
CUDA Version: 13.0
CUDNN Version: as shipped in the containers below
Operating System: JetPack R38 (release) REVISION 4.0, GCID 43443517, aarch64
Python Version (if applicable): 3.12
PyTorch Version (if applicable): 2.10.0a0+b4e4ee81d3.nv25.12
Baremetal or Container (if so, version):
nvcr.io/nvidia/pytorch:25.12-py3 → TensorRT 10.14.1.48
nvcr.io/nvidia/pytorch:26.09-py3 → TensorRT 11.3.0.99
Note: the stock JetPack R38.4 apt repo pins TensorRT to 10.13.3.9, which is
older than both versions tested, so default Jetson installs are likely affected
as well. There is no newer TensorRT reachable from Jetson via apt (pinned to the
JetPack release) or PyPI (tensorrt-cu13-libs publishes no aarch64 wheels —
only the Python bindings).
Relevant Files
Model link: S2M2 stereo matching, "S" variant (26.5 M parameters)
Graph: 3,123 ONNX nodes, 780 initializers. Two image inputs [1,3,H,W], three
outputs (disparity, occlusion, confidence).
I can attach the full --verbose build log (101,824 lines, 285 KB gzipped) and
~28 log files covering every control described below — happy to upload on
request or attach here.
Steps To Reproduce
Commands or scripts:
# 1. clone and fetch the S variant weights
git clone https://github.com/junhong-3dv/s2m2 && cd s2m2
mkdir -p weights/pretrain_weights
curl -L -o weights/pretrain_weights/CH128NTR1.pth \
https://huggingface.co/minimok/s2m2/resolve/main/CH128NTR1.pth
# 2. export ONNX at a failing resolution
python demo/export_onnx.py --model_type S --img_width 1536 --img_height 960
# 3a. TensorRT 10.14.1 -> SIGSEGV (exit 139)
trtexec --onnx=weights/onnx_save/S2M2_S_1536_960_v2_torch21.onnx \
--saveEngine=out.engine --fp16
# 3b. TensorRT 11.3.0 -> 0 MiB engine + assertion
# (--fp16 was removed in 11.x; networks are strongly typed from the ONNX)
trtexec --onnx=weights/onnx_save/S2M2_S_1536_960_v2_torch21.onnx \
--saveEngine=out.engine
Works at --img_width 1280 --img_height 800; fails at 1536 × 960 and above.
Resolution boundary (same pipeline, only dimensions varied):
| Resolution |
Pixels |
TRT 10.14.1 |
TRT 11.3.0 |
| 960 × 608 |
0.584 MPix |
builds, 61.85 ms fp16 |
builds, 55.48 ms fp16 |
| 1280 × 800 |
1.024 MPix |
builds, 124.36 ms fp16 |
not tested |
| 1536 × 960 |
1.475 MPix |
SIGSEGV |
0 MiB engine |
| 1920 × 1216 |
2.335 MPix |
SIGSEGV |
0 MiB engine |
The boundary lies between 1.024 and 1.475 megapixels.
Ruled out by controlled test:
| Hypothesis |
Test |
Result |
| Workspace / memory |
--memPoolSize=workspace: 2048, 8192, 24576 MB |
all SIGSEGV identically |
| Precision |
fp32 and fp16, both versions |
identical failure — 4 combinations |
| The workspace flag itself |
960 × 608 with workspace:8192 |
builds fine, 61.54 ms |
| System memory |
free during build |
~116 GB available throughout |
| Export pipeline |
same script, smaller resolutions |
builds and runs correctly |
| Builder optimisation |
levels 0, 1, 2, 3 |
1–3 SIGSEGV; level 0 builds then fails at enqueueV3 |
Control — the working resolution fuses identically. At 960 × 608 the
verbose log shows the same whole-graph fusion (1 layers, 1 ForeignNode,
1 Timing Runner) and builds successfully. Input dimensions are the only
variable between success and failure.
Attempted workaround that does NOT work. Marking intermediate tensors as
additional graph outputs (fully typed via ONNX shape inference;
onnx.checker.check_model passes) to force graph splits still fails — tried
with 2 and 5 cut points, both SIGSEGV under 10.14.1.
Second failure mode at --builderOptimizationLevel=0 (10.14.1): the engine
builds, then fails at execution:
[E] Error[1]: IExecutionContext::enqueueV3: Error Code 1:
Myelin ([cask.cpp:2974: exec] Platform (Cuda) error
In executeMyelinGraph at /_src/runtime/myelin/runner.cpp:820)
Have you tried the latest release?: Yes — TensorRT 11.3.0.99 via
nvcr.io/nvidia/pytorch:26.09-py3. The crash becomes a silent 0 MiB engine but
the build still fails at the same resolution boundary.
Can this model run on other frameworks?: Yes. torch.compile runs the same
model at all resolutions on the same hardware — 73.87 ms at 960 × 608 and
399.34 ms at 1920 × 1216. Where TensorRT does build it is faster (55.48 ms at
960 × 608, 11.3.0 fp16), which is why the resolution ceiling matters.
Questions
- Is there a supported way to prevent whole-graph Myelin fusion so the compiler
handles smaller subgraphs? Marking intermediate tensors as outputs does not
work.
- Is there a documented input-size limit for Myelin-compiled
ForeignNodes on
sm_110?
- In 11.3.0, why does the builder report success and emit a 0 MiB engine rather
than raising a build error?
- Is a fix expected in a release that reaches Jetson through JetPack? Neither
apt nor PyPI currently offers anything newer than 10.13.3.9 for aarch64.
Description
On Jetson AGX Thor (sm_110), TensorRT cannot build an engine for a stereo
matching network above roughly 1 megapixel of input. The same ONNX graph builds
and runs correctly at smaller resolutions.
The failure reproduces across a major version boundary with two different
symptoms:
SIGSEGVinside the Myelin compiler, ~2 min into the buildAssertion failure: false && "Attempting to access an empty engine!"In both versions TensorRT fuses the entire network into a single Myelin
ForeignNode. 10.14.1 crashes while compiling it; 11.3.0 completes withouterror but emits nothing.
The 11.3.0 behaviour is the more dangerous of the two: a build pipeline that
checks only the builder's exit status sees success and writes an empty engine.
The failure surfaces later, at load or inference time, far from its cause.
This is not resource exhaustion. ~116 GB of 122 GB were free throughout,
and workspace caps of 2/8/24 GB all fail identically.
Failing output, 11.3.0:
Failing output, 10.14.1 (verbose, final lines before SIGSEGV):
Environment
TensorRT Version: 10.14.1.48 and 11.3.0.99 (both affected)
NVIDIA GPU: Jetson AGX Thor Developer Kit (T5000), sm_110, MAXN power mode
NVIDIA Driver Version: 580.00
CUDA Version: 13.0
CUDNN Version: as shipped in the containers below
Operating System: JetPack R38 (release) REVISION 4.0, GCID 43443517, aarch64
Python Version (if applicable): 3.12
PyTorch Version (if applicable): 2.10.0a0+b4e4ee81d3.nv25.12
Baremetal or Container (if so, version):
nvcr.io/nvidia/pytorch:25.12-py3→ TensorRT 10.14.1.48nvcr.io/nvidia/pytorch:26.09-py3→ TensorRT 11.3.0.99Note: the stock JetPack R38.4 apt repo pins TensorRT to 10.13.3.9, which is
older than both versions tested, so default Jetson installs are likely affected
as well. There is no newer TensorRT reachable from Jetson via apt (pinned to the
JetPack release) or PyPI (
tensorrt-cu13-libspublishes no aarch64 wheels —only the Python bindings).
Relevant Files
Model link: S2M2 stereo matching, "S" variant (26.5 M parameters)
CH128NTR1.pth)demo/export_onnx.py(opset 18)Graph: 3,123 ONNX nodes, 780 initializers. Two image inputs
[1,3,H,W], threeoutputs (disparity, occlusion, confidence).
I can attach the full
--verbosebuild log (101,824 lines, 285 KB gzipped) and~28 log files covering every control described below — happy to upload on
request or attach here.
Steps To Reproduce
Commands or scripts:
Works at
--img_width 1280 --img_height 800; fails at1536 × 960and above.Resolution boundary (same pipeline, only dimensions varied):
The boundary lies between 1.024 and 1.475 megapixels.
Ruled out by controlled test:
--memPoolSize=workspace:2048, 8192, 24576 MBworkspace:8192freeduring buildenqueueV3Control — the working resolution fuses identically. At 960 × 608 the
verbose log shows the same whole-graph fusion (
1 layers, 1ForeignNode,1
Timing Runner) and builds successfully. Input dimensions are the onlyvariable between success and failure.
Attempted workaround that does NOT work. Marking intermediate tensors as
additional graph outputs (fully typed via ONNX shape inference;
onnx.checker.check_modelpasses) to force graph splits still fails — triedwith 2 and 5 cut points, both SIGSEGV under 10.14.1.
Second failure mode at
--builderOptimizationLevel=0(10.14.1): the enginebuilds, then fails at execution:
Have you tried the latest release?: Yes — TensorRT 11.3.0.99 via
nvcr.io/nvidia/pytorch:26.09-py3. The crash becomes a silent 0 MiB engine butthe build still fails at the same resolution boundary.
Can this model run on other frameworks?: Yes.
torch.compileruns the samemodel at all resolutions on the same hardware — 73.87 ms at 960 × 608 and
399.34 ms at 1920 × 1216. Where TensorRT does build it is faster (55.48 ms at
960 × 608, 11.3.0 fp16), which is why the resolution ceiling matters.
Questions
handles smaller subgraphs? Marking intermediate tensors as outputs does not
work.
ForeignNodes onsm_110?
than raising a build error?
apt nor PyPI currently offers anything newer than 10.13.3.9 for aarch64.