Skip to content

TransferBench v1.71.00 - #359

Open
AtlantaPepsi wants to merge 21 commits into
developfrom
candidate-1.71
Open

AtlantaPepsi wants to merge 21 commits into
developfrom
candidate-1.71

Conversation

@AtlantaPepsi

@AtlantaPepsi AtlantaPepsi commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

  • NIC executor check against max message size
  • fixing logical CU ID report
  • inclusion of gfx1250-strict target
  • Pingpong latency testing integration and presets

Technical Details

Test Plan

  • Validate HSA_DISABLE_GFX12_STRICT=0 ./TransferBench output of gfx1250-strict target
  • On gfx1250, gfx942 and gfx950 compare SHOW_ITERATIONS output CU ID against CU_MASK bits
  • Latency Testing
    • p2p_latency preset: in/cross-domain, and AMD and Nvidia platform
    • Concurrent transfers + pingpongs run, single/multistream
    • validation of all memory types support

Test Result

p2p_latency example output

[Latency Related]
GPU_MEM_TYPE         =            0 : Using default GPU memory for flags (0=default, 1=fine-grained, 2=uncached, 3=managed)
NUM_GPU_DEVICES      =            8 : Using 8 GPUs
NUM_LAPS             =         1000 : Timing 1000 round trips per iteration
USE_REMOTE_READ      =            0 : Executors write to their partner's memory and poll their own

Pingpong round-trip latency per lap (us), each pair run by itself
[1000 laps] [default GPU memory flags] [remote write / local poll]
 PING\PONG     GPU 00     GPU 01     GPU 02     GPU 03     GPU 04     GPU 05     GPU 06     GPU 07
    GPU 00      0.550      1.973      4.093      4.504      4.155      4.501      4.092      4.504
    GPU 01      1.956      0.585      4.502      4.930      4.507      4.860      4.505      4.925
    GPU 02      4.079      4.501      0.548      1.993      4.095      4.503      4.067      4.503
    GPU 03      4.501      4.919      1.995      0.592      4.503      4.873      4.502      5.010
    GPU 04      4.111      4.501      4.050      4.501      0.578      1.973      4.111      4.501
    GPU 05      4.501      4.901      4.501      4.837      1.993      0.581      4.503      4.818
    GPU 06      4.060      4.501      4.034      4.500      4.088      4.501      0.562      1.954
    GPU 07      4.502      4.957      4.502      4.873      4.511      4.841      1.956      0.575

Concurrent transfer + pingpong hybrid run example output

./TransferBench cmdline 16M "6 16 (G0->G1->G1) (G1->G2->G2) (G2->G3->G3) (G3->G0->G0
) (N G0 G1 +1000 N G1 G0) (C0 G0 G1 +50 C0 G1 G0)"

Test 1:
-------------------┬--------------┬------------┬-------------------┬---------------------------
  Executor: GPU 00 │ 302.684 GB/s │   0.055 ms │    16777216 bytes │ 332.617 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 3    │ 332.617 GB/s │   0.050 ms │    16777216 bytes │ G3 -> G0:16 -> G0
     PingPong 4    │     1.046 us │   1.046 ms │         1000 laps │ N->G0->G1 <+> N->G1->G0
     PingPong 5    │     1.544 us │   0.077 ms │           50 laps │ C0->G0->G1 <+> C0->G1->G0
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 01 │ 331.512 GB/s │   0.051 ms │    16777216 bytes │ 355.359 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 0    │ 355.359 GB/s │   0.047 ms │    16777216 bytes │ G0 -> G1:16 -> G1
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 02 │ 345.778 GB/s │   0.049 ms │    16777216 bytes │ 350.870 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 1    │ 350.870 GB/s │   0.048 ms │    16777216 bytes │ G1 -> G2:16 -> G2
-------------------┼--------------┼------------┼-------------------┼---------------------------
  Executor: GPU 03 │ 358.855 GB/s │   0.047 ms │    16777216 bytes │ 364.532 GB/s (sum)
-------------------┼--------------┼------------┼-------------------┼---------------------------
     Transfer 2    │ 364.532 GB/s │   0.046 ms │    16777216 bytes │ G2 -> G3:16 -> G3
-------------------┼--------------┼------------┼-------------------┼---------------------------
   Aggregate (CPU) │  58.392 GB/s │   1.149 ms │    67108864 bytes │ Overhead 1.094 ms
-------------------┴--------------┴------------┴-------------------┴---------------------------

Submission Checklist

Copilot AI lite review requested due to automatic review settings September 23, 2026 15:27
@AtlantaPepsi
AtlantaPepsi requested review from a team as code owners September 23, 2026 15:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical and moderate correctness, portability, parsing, and build-target issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 4 High severity · 4 Medium severity

Open (8)
What changed in this PR

Adds pingpong latency testing, NIC message-size validation, logical CU reporting fixes, and gfx1250-strict support.

Changes:

  • Adds pingpong parsing, execution, timing, and latency presets.
  • Adds NIC max_msg_sz validation and reporting.
  • Updates GPU architecture and TDM handling.
File Review summary
src/​header/​TransferBench.hpp Critical and moderate issues in XCC handling, pingpong validation, NIC limits, subindex selection, dump parsing, and timing scaling.
src/​header/​tdmCopy.h Critical architecture guard mismatch enables unsupported TDM targets.
src/​client/​Utilities.hpp Reviewed result and topology utilities.
src/​client/​Topology.hpp Reviewed NIC message-size reporting.
src/​client/​Presets/​Presets.hpp Reviewed latency preset registration.
src/​client/​Presets/​Latency.hpp Moderate issues with failed-run result handling and diagonal pair measurement.
src/​client/​EnvVars.hpp Reviewed pingpong configuration variables.
src/​client/​Client.cpp Reviewed pingpong transfer display updates.
docs/​install/​build_from_source.rst Reviewed strict GPU target documentation.
CMakeLists.txt Moderate issue: package build targets omit gfx1250-strict.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/header/TransferBench.hpp
exeInfo.totalSubExecs += t.numSubExecs;
} else {
exeInfo.totalPingpong ++;
}
exeInfo.useSubIndices |= (t.exeSubIndex != -1 || (t.exeDevice.exeType == EXE_GPU_GFX && !cfg.gfx.prefXccTable.empty()));
Comment thread src/header/TransferBench.hpp Outdated
Comment thread src/header/tdmCopy.h
Comment thread src/client/Presets/Latency.hpp Outdated
Comment thread src/header/TransferBench.hpp
Comment thread src/header/TransferBench.hpp
Comment thread src/header/TransferBench.hpp
Copilot AI review requested due to automatic review settings September 23, 2026 15:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread src/header/TransferBench.hpp
@AtlantaPepsi AtlantaPepsi changed the title Candidate 1.71 TransferBench v1.71.00 Sep 23, 2026
Copilot AI review requested due to automatic review settings September 23, 2026 21:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

#else
useSubIndexCount[exe]++;
int numSubIndices = GetNumExecutorSubIndices(exe);
if (subIndex >= numSubIndices) {
Comment on lines +6660 to +6661
dim3 const gridSize(xccDim, numPingpong, 1);
dim3 const blockSize(1);
Comment on lines 3111 to +3115
if (t.numSubExecs <= 0)
errors.push_back({ERR_FATAL, "Transfer %d: # of subexecutors must be positive", i});
else
else if (isPingpong) {
if (t.numSubExecs != 1)
errors.push_back({ERR_WARN,
Copilot AI review requested due to automatic review settings September 28, 2026 15:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +5562 to +5567
} else if (partnerMem.memIndex != exeDevice.exeIndex) {
if (System::Get().IsVerbose()) {
System::Get().Log("[INFO] Enabling pingpong peer access: GPU %d -> GPU %d\n",
exeDevice.exeIndex, partnerMem.memIndex);
}
ERR_CHECK(EnablePeerAccess(exeDevice.exeIndex, partnerMem.memIndex));
Comment thread src/client/Presets/Latency.hpp
Copilot AI lite review requested due to automatic review settings October 8, 2026 13:47
* fix (client): pre-resolve master address in LaunchTransferBench

The host list was forwarded verbatim as TB_MASTER_ADDR, so an ssh_config
alias that the local SSH client understands would fail getaddrinfo() on
the workers, leaving rank 0 waiting on connections that never arrive.
Expand the entry via ssh -G and prefer a literal IPv4, since workers
resolve the master address remotely and only over AF_INET.

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Apply suggestion from @nileshnegi

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

p.srcMem[0] = nullptr;
p.srcMem[1] = nullptr;
p.localFlagMem = nullptr;
p.flagMem = static_cast<volatile uint8_t*>(static_cast<void*>(rss.dstMem[0]));
// Expand pong half
std::vector<Transfer> pongTransfers;
for (int r = 0; r < numRanks; r++) {
if (!RecursiveWildcardTransferExpansion(pongWct, r, numBytes, numSubExecs, pongTransfers))
Copilot AI lite review requested due to automatic review settings October 8, 2026 13:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +6566 to +6570
// Advance one stride, plus an extra stride every hp laps so that a slot is never
// revisited an even number of laps later (which would leave a stale matching value)
off += sx; if (off >= y) off -= y;
if (--hopCnt == 0) {
hopCnt = hp;
Copilot AI lite review requested due to automatic review settings October 8, 2026 14:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines 6810 to +6812
for (int i = 0; i < exeInfo.resources.size(); i++) {
TransferResources& rss = exeInfo.resources[i];
if (rss.numLaps != 0) continue;
Comment on lines +11 to +30
## Single-node bandwidth presets

| Preset | Purpose |
|---|---|
| `a2a` | All-to-all parallel transfers between every pair of GPUs. |
| `a2asweep` | GFX-based a2a swept across CU counts and unroll factors (`MEM_TYPE`, `NUM_SUB_EXECS`). |
| `bmasweep` | Compares DMA vs. Batched-DMA for one-to-many copies (HIP 7.1 / CUDA 12.8+). |
| `gfxsweep` | Sweeps GFX kernel options for one Transfer. |
| `hbm` | Local HBM read bandwidth on each GPU. |
| `healthcheck` | Quick correctness/perf health check (AMD MI300 series only). |
| `one2all` | All subsets of parallel transfers from one GPU to all others. |
| `p2p` | Peer-to-peer device-memory matrix between every GPU pair. |
| `pcopy` | Parallel copies from a single GPU to other GPUs. |
| `rsweep` | Random sweep through Transfer combinations. |
| `rwrite` | Parallel remote writes from a single GPU to others. |
| `scaling` | Scaling test: one GPU → all others, varying SEs, mem types (`CPU_MEM_TYPE`, `GPU_MEM_TYPE`). |
| `schmoo` | Local/remote read/write/copy scaling between two GPUs. |
| `smoketest` | Quick DMA/GFX correctness sweep. |
| `sweep` | Ordered sweep through Transfer combinations. |
| `wallclock` | Compares wallclock counters across XCCs within one GPU. |
Copies standardised security scanning config from
ROCm/rocm-repo-template.

Co-authored-by: haribabug <haribabug@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 8, 2026 14:46
- docs/sphinx/requirements.txt: bump gitpython, tornado, pyjwt, urllib3,
  cryptography (plus cffi and typing-extensions, which it requires),
  soupsieve and jupyter-core past their HIGH/CRITICAL CVEs; all other
  pins unchanged
- workflows: pin actions to commit SHAs and container images to digests
- build-relocatable-packages: grant id-token: write only to the two jobs
  that upload to S3 instead of the whole workflow

Co-authored-by: Cursor <cursoragent@cursor.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Unresolved security-scan allowlist gaps and functional correctness issues remain.

12 open findings
Previously missed (3)

In code that hasn't changed since last review

Medium severity tee pipeline aborts on SIGPIPE when truncated with head

.claude/​skills/​transferbench-debug/​examples/​topology-probe.sh:24

With set -o pipefail, truncating the tee pipeline with head can close the pipe while the binary is still writing; tee/the binary then exits on SIGPIPE and the script aborts before completing the probe. Use a reader that consumes the full stream (for example sed -n '1,60p') or capture first and truncate afterward.

Medium severity Documented presets are not registered in presetFuncMap

.claude/​skills/​transferbench-run/​references/​presets.md:25

This reference presents these rows as runnable presets, but neither pcopy nor rwrite is registered in the current presetFuncMap (src/client/Presets/Presets.hpp:68-99), so following either command fails instead of launching a benchmark. Remove them or register the corresponding presets before documenting them here.

Medium severity Packaging defaults omit gfx1250 GPU targets

CMakeLists.txt:224

Adding the target only to CMake's default list does not add it to the package build path: build_packages_local.sh passes its separate DEFAULT_GPU_TARGETS list to -DGPU_TARGETS, and that list still omits both gfx1250 and gfx1250-strict. Packages built without an explicit GPU_TARGETS override therefore will not contain the target advertised by this change; update the packaging defaults or explicitly document the scope.

🧠 Review effort: Lite


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment on lines +35 to +44
[[allowlists]]
# Allowlist by location only where a real first-party secret structurally
# can't live. Dependency lock files qualify: their contents are generated
# from a manifest, and the high-entropy strings they carry are artifact
# digests rather than credentials. Add vendored third-party trees here as
# callers bring them in, one explicit path per entry.
description = "Generated dependency lock files"
paths = [
'''.*\.lock$''',
]
Comment on lines +53 to +55
'''(?i)\bINVALID_TOKEN\s*=''',
'''(?i)\bwrong_secret\s*=''',
'''(?i)\bsecret\s*='''
Copilot AI lite review requested due to automatic review settings October 8, 2026 14:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Unresolved moderate correctness, packaging, launcher, documentation, and security-configuration issues remain.

12 open findings
Previously missed (2)

In code that hasn't changed since last review

Medium severity Packaging defaults omit the new GPU targets

CMakeLists.txt:224

Adding the target only to CMakeLists.txt does not add it to packaged builds: build_packages_local.sh overrides GPU_TARGETS with its own hard-coded default list, which currently omits both gfx1250 and gfx1250-strict. Update that packaging default as well, otherwise the release packages will not contain the target this change advertises.

Low severity Referenced presets are not registered

.claude/​skills/​transferbench-run/​references/​presets.md:25

These entries are not available presets in the current source: Presets.hpp registers neither pcopy nor rwrite, and no source-side dispatch for either name exists. The run-side reference therefore directs users to commands that will be rejected; remove them or document the actual registered preset names.

🧠 Review effort: Lite


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Copilot AI lite review requested due to automatic review settings October 9, 2026 00:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines 7533 to +7535
Transfer const& t = transfers[resource->transferIdx];
// Ping and Pong will start with src value of 0
if (t.numLaps != 0) continue;

echo
echo "=== Quick parser sanity check ==="
"$BINARY" dryrun "1 4 (G0->G0->G1)" 2>&1 | head -10
Comment thread CMakeLists.txt
Comment on lines +223 to +224
gfx1250
gfx1250-strict)
Comment on lines +48 to +49
./TransferBench dryrun "<expression>" # validate parsing, expand wildcards
TB_DUMP_CFG_FILE=dump.cfg ./TransferBench p2p # dump what a preset actually emits
- Quoting issues on the shell side when using `cmdline` (e.g. `G*` getting glob-expanded).

### Fix
1. **Always quote** `cmdline` arguments: `./TransferBench cmdline "1 4 (G0->G0->G1)"`.
TB_DUMP_CFG_FILE=p2p_dump.cfg ./TransferBench p2p

# "Is the slowness in iter 0 only, or every iter?"
NUM_WARMUPS=0 NUM_ITERATIONS=20 SHOW_ITERATIONS=1 ./TransferBench cmdline "1 4 (G0->G0->G1)" 256M
Comment on lines +39 to +40
- `cmdline "<transfer expression>"` — run one ad-hoc transfer
- `dryrun "<transfer expression>"` — parse and print without executing
-2 (G0->G0->G1 4 1M) (G1->G1->G0 8 2M)
# Copies 1MiB GPU0->GPU1 with 4 CUs, in parallel with 2MiB GPU1->GPU0 with 8 CUs
```

Comment on lines +100 to +101
./TransferBench dryrun "1 4 (G0->G0->G1)"
./TransferBench dryrun my.cfg
Copilot AI lite review requested due to automatic review settings October 10, 2026 03:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +3047 to +3049
"Transfer %d: Cross-rank GPU executor (R%d%c%d) cannot access remote host memory "
"(%s on rank %d is %s). Fabric-handle sharing only supports GPU memory for 1.67; use a NIC "
"executor (e.g. R%dN..) for cross-rank transfers involving host memory.",
Comment on lines +23 to +25
| `pcopy` | Parallel copies from a single GPU to other GPUs. |
| `rsweep` | Random sweep through Transfer combinations. |
| `rwrite` | Parallel remote writes from a single GPU to others. |
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants