Skip to content

TensorRT silently produces wrong results for batch > 1 on SegFormer (b0/b1/b2 verified, fp32, static and dynamic engines, TRT 10.11 & 11.1, sm89 + sm120) #4853

Description

@Daniel595

Summary

TensorRT silently produces wrong results for batch > 1 when compiling SegFormer models. This was verified with all three public checkpoints nvidia/segformer-b0/b1/b2-finetuned-ade-512-512 and a reduced single encoder block of the b2 variant. Batch=1 is bit-accurate against ONNX Runtime; batch=2 is wrong by orders of magnitude while trtexec reports PASSED. The bug is:

  • precision-independent (fp32 AND fp16 engines affected),
  • shape-mode-independent (fully static batch=2 engines min=opt=max=2 AND dynamic 1–2 engines affected),
  • architecture-independent (reproduced on sm_120 Blackwell laptop and sm_89 L4),
  • present in TensorRT 10.11 (10.11.0.x) and 11.1 with identical error values,
  • not affected by fusion level (--builderOptimizationLevel=0 still fails, though less catastrophically).

ONNX Runtime (CPU and CUDA EP) executes the same ONNX with batch=2 correctly on the same GPUs.

Environment

  • TensorRT 10.11 — trtexec [TensorRT v101100] (CUDA 12) and TensorRT 11.1 — trtexec [TensorRT v110100] (from nvcr.io/nvidia/tritonserver:26.07-py3)
  • GPUs: NVIDIA RTX PRO 2000 Blackwell Generation Laptop (sm_120), driver 595.84; NVIDIA L4 (sm_89)
  • ONNX: opset 19, exported with torch.onnx.export (legacy TorchScript path) + onnxsim, dynamic batch axis
  • Model: nvidia/segformer-b2-finetuned-ade-512-512 (public), 512×512 NHWC input, fp32

Symptom (full public models)

Build a plain fp32 engine with dynamic batch 1–2 and run the same random input as batch=1 and as b2[0] of a batch=2 request. The batch=2 input's first sample is identical to the batch=1 input, so a correct engine must produce identical outputs:

full model (512×512, fp32) b2[0] vs b1, TRT 10.11 b2[0] vs b1, TRT 11.1 argmax match b2[0] vs b1
segformer-b0-finetuned-ade-512-512 9.94 (34% of out magnitude) not tested —
segformer-b1-finetuned-ade-512-512 11.35 (37%) not tested —
segformer-b2-finetuned-ade-512-512 19.2 19.2 49.2%
ORT CUDA EP: b2[0] vs b1 max_abs_diff = 0.018, argmax match = 100 %
ORT CPU    : batch=2 bit-identical to batch=1

The segmentation output is essentially garbage for batch≥2 (with a fine-tuned SegFormer variant at 1024×1024, batch=2 collapsed to all-background). The b0/b1/b2 checkpoints above were tested via generate/export_public.py from the attached zip, which regenerates each ONNX and the input files from the HuggingFace hub.

Bisection

Prefix-bisection (cut the graph after every encoder block, compare TRT vs ORT at batch=2, fp32):

prefix cut b2 max_abs_diff vs ORT
patch embeds + stage-0 blocks (0.0–0.2) ~0.002–0.007 (noise)
block.1.0 end 0.62 (first corruption)
block.1.1 end 1.57
… monotonically growing downstream … …
full model ~19

Node-level cuts inside block.1.0 show: everything up to the mlp/dense1 output is clean at batch=2; the first wrong values appear at the mlp/dwconv (3×3 depthwise conv, groups=512) output (diff 3.1 vs fp32 noise 0.004).

However, the dwconv alone does not reproduce the bug:

  • single Conv (same weights, shapes, real activations): correct at batch=2
  • dense1 → transpose → reshape → dwconv: correct
  • LayerNorm → dense1 → (dynamic Shape/Gather/Concat reshape chain) → dwconv: correct

Only the complete block (LayerNorm → efficient attention → residual/LN → MLP with dwconv, 64 nodes, 1.9 MB) reproduces. Ending the same graph at dense1 (removing the dwconv) heals it. So the broken kernel/tactic is only selected in the full-block compilation context.

Minimal repro (attached, trt_batch_bug_repro.zip)

repro.onnx — one SegFormer-b2 encoder block, fully synthetic weights (seeded RNG), no third-party data. Input x: [B, 4096, 128], output y: [B, 512, 64, 64]. The batch=2 input file's first sample is identical to the batch=1 input, so a correct engine must produce identical outputs — no reference model needed:

trtexec --onnx=repro.onnx --saveEngine=repro.engine \
        --minShapes=x:1x4096x128 --optShapes=x:2x4096x128 --maxShapes=x:2x4096x128

trtexec --loadEngine=repro.engine --shapes=x:1x4096x128 \
        --loadInputs=x:repro_input_b1.bin --exportOutput=out_b1.json \
        --iterations=1 --warmUp=0 --duration=0

trtexec --loadEngine=repro.engine --shapes=x:2x4096x128 \
        --loadInputs=x:repro_input_b2.bin --exportOutput=out_b2.json \
        --iterations=1 --warmUp=0 --duration=0

python3 verify.py out_b1.json out_b2.json

Measured max_abs_diff(out_b2[0], out_b1[0]) — identical input, same engine:

TRT 10.11 (v101100) TRT 11.1 (v110100)
repro.onnx (synthetic weights) 0.0133 0.0133
repro_public_weights.onnx (public SegFormer weights) 3.105 3.105

For comparison: ONNX Runtime CPU executes batch=2 bit-identical to batch=1,
and a correct fp32 engine gives ~1e-6 for the same comparison. Every trtexec
run prints &&&& PASSED.

Additional observations

  • Same failure through Triton Inference Server (26.07 / TRT backend), so it is not trtexec-specific.
  • Two identical images in the batch produce two identical wrong outputs (deterministic miscompile, not input mixing).
  • With the 1024×1024 fine-tuned SegFormer variant we additionally verified: fp16 AND fp32 fail; static batch=2 (min=opt=max=2) fails; --builderOptimizationLevel=0 still fails (max_abs_diff ~22 instead of ~38; batch=2 no longer fully degenerate).
  • Current workaround: build engines for batch=1 only.

Related reports

Happy to provide engine files, layer info dumps or run additional experiments on request.

trt_batch_bug_repro.zip

Activity

  1. kvnloo commented on Sep 28, 2026

    @kvnloo

    The first-sample equivalence test is a particularly strong oracle here:

    if batch2[0] is byte-identical to the batch=1 input, then the corresponding output should differ only by normal numerical noise.

    Since corruption first appears around the stage-1 MLP depthwise convolution but the same Conv alone is correct, I would continue reducing the context rather than the Conv itself.

    Useful toggles:

    • remove the preceding attention;
    • remove the residual;
    • replace dynamic reshape/shape chain with constants;
    • force the suspect depthwise Conv to a different tactic if possible.

    For every arm keep:

    • FP32;
    • static batch=2;
    • TensorRT vs ORT;
    • batch2[0] vs batch1.

    That should identify the smallest preceding layout/state that makes the otherwise-correct depthwise Conv consume batch data incorrectly.

  2. OhtaTetsuya commented on Oct 7, 2026

    @OhtaTetsuya

    Confirming the same issue on a different model, with a bit more localization and a graph-level workaround.

    Setup / symptom: a detector with a PVT-style transformer backbone (pre-norm residual blocks),
    ONNX opset 18 with LayerNormalization function ops. TensorRT 10.14.1.48, RTX A4500 (sm86).
    The same ONNX is correct with TensorRT 10.3.0.26. Symptom is identical to this issue: batch=1
    matches ONNX Runtime, batch>=2 is wrong for every image in the batch, in fp32 and fp16, with
    dynamic (1-16) and static (8) engines alike, while trtexec reports PASSED. A batch of 8 copies of
    the same image gives 8 identical but wrong outputs, so it is a batch-size-dependent layout error
    rather than cross-sample contamination.

    Localization: the first divergence is in backbone stage 2, block 0 (stage 1 is fine), consistent
    with the block.1.0 finding above. Marking intermediate tensors as outputs changes fusion and hides
    the bug, so I extracted a single pre-norm block as a standalone ONNX; it reproduces (max abs err 5.2
    vs 3e-6 at batch 1). --profilingVerbosity=detailed shows the residual Add and the following
    LayerNorm inside one Myelin kernel,
    __myl_ReshReshTranReshAddReshTranMeanSubMulMeanAddSqrtDivMulMulAdd_*, followed by a
    __myl_MoveReshTranReshMove_* kernel that copies the residual stream to the next Add. Marking the
    LN output (or the input of the attention out-projection GEMM) as a network output fixes it; marking
    the GEMM output or the residual Add output does not. --fp16, --builderOptimizationLevel=0,
    --stronglyTyped, and decomposing LN into primitive ops (Myelin re-fuses it) do not help.

    Workaround: in the ONNX, for every Add -> LayerNormalization pair insert
    Reshape(-1, C) -> LayerNormalization -> Reshape(original shape). Outputs then match ONNX Runtime
    (max abs diff 8e-6 vs the unmodified ONNX on CPU) for batch 1/8/16, static and dynamic engines.
    The fused LN kernel is still generated, the MoveReshTranReshMove copy kernel disappears, and the
    engine is slightly faster (b1 5.48 -> 5.24 ms, b8 27.2 -> 25.9 ms).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Module:AccuracyOutput mismatch between TensorRT and other frameworks

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions