You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This RFC proposes a single TENT operator CLI, tentatively named tent, and a versioned sanitized diagnostic bundle format.
The goal is to let an operator answer the following questions without reading TENT source code or manually correlating many logs:
What CPU/GPU/NPU/NIC topology did TENT discover?
Which transports, devices, policies, QP pools, SL/TC values, and runtime features are enabled?
Are the local configuration and remote peer compatible?
Why did TENT select a Direct or Staged path, a particular backend/rail, or a fallback?
Is a failure caused by configuration, metadata, connectivity, runtime queue pressure, receiver credit, staging, or transport execution?
What evidence can be attached to a GitHub issue without exposing credentials, rkeys, raw addresses, or other sensitive deployment data?
The first implementation milestone is read-only. It must not hot-reload configuration, quarantine a rail, trigger failover, change QoS policy, or upload diagnostics automatically.
The CLI should reuse existing functionality rather than create parallel implementations:
The TENT roadmap in #1058 explicitly lists an isolated transfer/diagnostics CLI as a community-contributable production-readiness item. TENT has accumulated many useful mechanisms, but their operational interfaces remain fragmented.
These are useful data sources, but users still need different binaries, configuration knowledge, log patterns, and manual interpretation. A failed cross-node run often requires separately checking:
local NIC enumeration and GID selection;
remote metadata and endpoint names;
transport enablement and build flags;
policy matching and device masks;
QP pool, SL, and TC settings;
Direct versus Staged routing;
peer/control connectivity;
queue, staging, or receiver-side pressure;
the exact Mooncake build and sanitized configuration.
This increases support cost and makes bug reports difficult to reproduce. The CLI should provide a common diagnostic model and explain what it observed, what it inferred, and what it could not verify.
Goals
Provide one discoverable tent command with stable subcommands.
Reuse existing collectors, probes, and benchmarks.
Separate offline checks from live/remote checks.
Produce both human-readable text and versioned machine-readable JSON.
Define consistent check severity, evidence, remediation, and exit-code semantics.
Redact sensitive values by default.
Generate a bounded diagnostic bundle suitable for offline support and GitHub issues.
Keep all v1 operations read-only and explicitly bounded.
Work with partial builds where some transports/platforms are unavailable.
Provide a stable integration point for future TENT health, Execution Plan, and QoS diagnostics.
Non-goals
A configuration hot-reload or remote administration service.
Automatic rail quarantine, failover, or remediation.
A second topology discovery implementation.
A new benchmark engine that competes with tebench.
Dumping all environment variables, process memory, rkeys, or raw transport handles.
Automatically uploading a diagnostic bundle.
A cluster-wide scheduler or monitoring platform.
Replacing Prometheus, tracing, NCCL RAS, or vendor fabric tools.
Guaranteeing that a successful control-plane probe proves the data path is healthy.
Requiring every optional backend to be installed for the base CLI to build.
Proposed command surface
The exact spelling can be adjusted during review, but the command responsibilities should remain distinct.
tent version
Print:
Mooncake/TENT version and Git commit;
build type and enabled platform/transport features;
CLI diagnostic schema versions;
compiler, CUDA/ROCm/CANN, and relevant runtime versions where available.
tent show-topology
Display the local topology known to TENT:
CPU NUMA nodes;
GPU/NPU/device identifiers;
NIC name, type, NUMA node, link state, speed, port, and selected GID;
memory locations;
GPU/NPU-to-NIC affinity;
NVLink/MNNVL or other supported local links;
unavailable or partially discovered devices with reasons.
This command should call the same topology/prober code used by TENT, not parse a second set of sysfs files independently.
tent check-config
Validate an effective configuration without starting transfer traffic.
Checks should include:
parse/type/range validation;
unknown keys and deprecated aliases;
build-time availability of enabled transports;
policy references to valid transports and devices;
QP pool, SL, TC, and rail-topology consistency;
conflicting allow/deny lists;
impossible Direct/Staged combinations;
QoS Contract validation when configured;
fields that require restart rather than future hot reload;
source/provenance of effective values when available.
The output must distinguish:
PASS: verified and valid;
WARN: valid but suspicious or unverifiable;
FAIL: invalid or unsafe;
SKIP: collector/feature unavailable in this build.
tent show-link
Reuse #2820's show_link implementation and output model. The existing binary may remain as a compatibility wrapper or alias.
In later milestones, an optional remote target may add:
local-to-remote NIC mapping;
selected GID/port and the reason;
control-plane reachability;
expected Direct/Staged candidate paths;
unavailable mappings and specific rejection reasons.
tent test-run
Perform a small, bounded connectivity test using registered scratch memory.
It should report separately:
metadata resolution;
control-plane probe;
remote descriptor compatibility;
selected policy and Execution Plan;
transport/rail/QP pool/SL/TC;
submit, completion, and data-integrity result;
stage timing when Staged;
structured failure origin and recommended next check.
Safety requirements:
no arbitrary raw remote address supplied by the user;
bounded default buffer size, request count, duration, and concurrency;
explicit opt-in for any larger test;
no long-running server exposed on a public interface by default;
clear distinction between a control probe and a real data-path test.
tent show-plan
Consume #2863's planner or legacy decision adapter and explain:
effective intent and matched policy;
Direct/Staged physical path;
Primary/Fallback/Degraded role;
backend, device/rail, QP pool, SL/TC;
resource charge vector;
cost/health inputs and freshness;
accepted and rejected candidates with reasons.
Before #2863 is implemented, this subcommand may be absent or limited to the current selector decision. The CLI RFC must not define a second Execution Plan model.
tent status
Query a local running TENT instance through a bounded local interface in a later milestone. Candidate fields:
instance/config generation and uptime;
registered segment/buffer counts without raw addresses;
peer/control state;
queue depth and inflight work;
staging occupancy;
receiver-credit state;
rail/QP/backend status;
last-progress timestamp and recent categorized errors.
The local query interface should default to a Unix socket or localhost and must not block transfer progress. This command is observation-only; health-policy automation belongs to a separate RFC.
tent collect-diagnostics
Collect a bounded snapshot from available commands and data sources.
Suggested contents:
manifest.json with bundle/schema version and checksums;
version/build information;
sanitized effective configuration and provenance;
topology and link information;
policy/QoS validation;
plan explain for a supplied target/request shape, if requested;
bounded status/metrics snapshot;
bounded recent categorized errors;
results of explicitly requested probes/tests;
collector errors and skipped sections.
The bundle is created locally and is never uploaded automatically.
tent bench
This is an optional convenience frontend to tebench, not a new benchmark implementation. It should preserve access to the underlying tebench command line and output schema. More complex benchmark changes should continue to land in tebench first.
Diagnostic data model
All subcommands should build a normalized diagnostic result before rendering text or JSON.
Diagnostics frequently contain deployment-sensitive information. Redaction must happen at collection/model boundaries, not as a final regex over serialized JSON.
Always excluded by default
rkeys and lkeys;
raw CUDA IPC/Fabric handles;
credentials, tokens, passwords, certificates, and private keys;
arbitrary process environment dumps;
raw memory addresses;
payload data;
unbounded logs;
cloud instance credentials or metadata-service responses.
Redaction modes
Suggested modes:
default -> safe for ordinary support; preserve enough topology to diagnose
strict -> pseudonymize hosts, IPs, device serials, paths, tenant/policy names
none -> explicit local-only opt-in with a warning; secrets remain excluded
Strict mode can use a random per-bundle salt so the same host/device remains correlatable inside one bundle without being correlatable across bundles.
Bundle handling
output files should be created with restrictive permissions;
archive size and per-section size must be bounded;
collection timeout must be bounded;
symlinks and arbitrary file inclusion are forbidden;
the manifest records redaction mode and omitted sections;
the tool must not upload or transmit the bundle without a separate explicit user action outside this RFC.
Remote testing
test-run and remote show-link must use existing authenticated/authorized control paths where available. The CLI must not introduce an unauthenticated arbitrary-memory test server. A standalone validation server, if needed later, requires an explicit security design and is out of v1 scope.
Offline and live modes
Commands fall into two groups.
Offline/local discovery
version
show-topology
check-config
local show-link
These should work without a running TENT instance or metadata service when possible.
Live/remote diagnosis
remote show-link
test-run
show-plan with remote metadata
status
live sections of collect-diagnostics
These must report partial/unavailable state clearly when the peer or metadata service cannot be reached.
Exit codes
Suggested stable exit codes:
0 command completed and no FAIL checks were produced
1 command completed and at least one diagnostic check failed
2 invalid command line or invalid input
3 required collector/service unavailable or command timed out
4 internal tool error or incompatible diagnostic schema
Warnings alone return zero unless --strict is specified. JSON output must still be emitted for diagnostic failures whenever possible.
Collectors should be reusable libraries rather than logic embedded directly in main(). Existing commands such as show_link can call the same collector/rendering library.
Optional collectors must be isolated by build feature. A CPU-only build should still provide config/topology checks and report GPU/RDMA collectors as unavailable rather than failing to build the CLI.
Compatibility and rollout
Existing show_link behavior remains available during migration.
Existing tebench remains the benchmark implementation and may remain a separate binary.
No current environment variable or config key changes meaning.
The CLI is additive and does not initialize transports unless a command explicitly requires it.
Read-only commands must avoid creating GPU contexts where discovery can be completed without one.
Commands that initialize a real engine must say so in help and output.
JSON schemas and exit codes require compatibility tests.
Proposed PR sequence
PR1: Diagnostic core and schema
Add DiagnosticSnapshot, checks, evidence, renderers, redaction primitives, and schema tests.
Add tent version.
No transport initialization or remote calls.
PR2: Topology and existing show-link integration
Add tent show-topology and tent show-link using existing collectors.
Preserve show_link as a wrapper/alias if maintainers prefer.
Add fake topology and real Linux discovery tests.
PR3: Config validation
Add tent check-config.
Reuse TENT config parsing, policy validation, QoS Contract validation, and effective-config provenance where available.
Add invalid/mismatched configuration fixtures.
PR4: Bounded connectivity test
Add tent test-run using scratch registered buffers.
Separate control probe, metadata resolution, data-path completion, and integrity results.
tent bench forwards to tebench without changing benchmark semantics.
Real-cluster validation
On the existing H20/RoCE nodes, cover:
healthy two-node RDMA and TCP paths;
multiple NUMA domains and NICs;
an unreachable or wrong GID/port;
a disabled/down NIC;
policy referring to a missing device;
local/remote transport capability mismatch;
Direct GPU↔GPU transfer;
forced Staged transfer;
peer process stopped or killed;
runtime queue and receiver-credit enabled/disabled where available;
default and strict diagnostic bundles inspected for secrets.
For every failure, the CLI should identify the furthest verified stage and avoid claiming later stages passed. For example, a successful control probe must not be reported as successful RDMA data connectivity.
Overhead
offline commands add no data-path overhead;
status collection must not block transfer progress;
any live monitoring interval is bounded and configurable;
diagnostic collection must complete or time out without leaking threads, QPs, registered memory, or temporary files.
Expected benefits and claims boundary
This RFC should provide:
reproducible and more complete issue reports;
faster diagnosis of GID, port, remote QP, NUMA, rail topology, policy, and capability problems;
a common explain surface for Execution Plan and QoS decisions;
stable JSON for automation and CI qualification;
a secure default path for sharing diagnostics;
less duplicate diagnostic logic across examples and ad-hoc scripts.
It does not claim to fix a failing path automatically or improve transfer performance. Any future remediation or automatic health policy requires separate design and validation.
Summary
This RFC proposes a single TENT operator CLI, tentatively named
tent, and a versioned sanitized diagnostic bundle format.The goal is to let an operator answer the following questions without reading TENT source code or manually correlating many logs:
The first implementation milestone is read-only. It must not hot-reload configuration, quarantine a rail, trigger failover, change QoS policy, or upload diagnostics automatically.
The CLI should reuse existing functionality rather than create parallel implementations:
show-linkand NIC discovery from [TransferEngine] Add show-link diagnostic tool for NIC topology #2820;tebenchfor performance testing;Motivation
The TENT roadmap in #1058 explicitly lists an isolated transfer/diagnostics CLI as a community-contributable production-readiness item. TENT has accumulated many useful mechanisms, but their operational interfaces remain fragmented.
Examples:
show_link, including local NIC discovery, NUMA affinity, link speed, and a readable/JSON topology matrix.tebenchexercises real transfers and already contains multiple benchmark modes.probePeerAliveByID()can test control-plane liveness.These are useful data sources, but users still need different binaries, configuration knowledge, log patterns, and manual interpretation. A failed cross-node run often requires separately checking:
This increases support cost and makes bug reports difficult to reproduce. The CLI should provide a common diagnostic model and explain what it observed, what it inferred, and what it could not verify.
Goals
tentcommand with stable subcommands.Non-goals
tebench.Proposed command surface
The exact spelling can be adjusted during review, but the command responsibilities should remain distinct.
tent versionPrint:
tent show-topologyDisplay the local topology known to TENT:
This command should call the same topology/prober code used by TENT, not parse a second set of sysfs files independently.
tent check-configValidate an effective configuration without starting transfer traffic.
Checks should include:
The output must distinguish:
PASS: verified and valid;WARN: valid but suspicious or unverifiable;FAIL: invalid or unsafe;SKIP: collector/feature unavailable in this build.tent show-linkReuse #2820's
show_linkimplementation and output model. The existing binary may remain as a compatibility wrapper or alias.In later milestones, an optional remote target may add:
tent test-runPerform a small, bounded connectivity test using registered scratch memory.
It should report separately:
Safety requirements:
tent show-planConsume #2863's planner or legacy decision adapter and explain:
Before #2863 is implemented, this subcommand may be absent or limited to the current selector decision. The CLI RFC must not define a second Execution Plan model.
tent statusQuery a local running TENT instance through a bounded local interface in a later milestone. Candidate fields:
The local query interface should default to a Unix socket or localhost and must not block transfer progress. This command is observation-only; health-policy automation belongs to a separate RFC.
tent collect-diagnosticsCollect a bounded snapshot from available commands and data sources.
Suggested contents:
manifest.jsonwith bundle/schema version and checksums;The bundle is created locally and is never uploaded automatically.
tent benchThis is an optional convenience frontend to
tebench, not a new benchmark implementation. It should preserve access to the underlying tebench command line and output schema. More complex benchmark changes should continue to land in tebench first.Diagnostic data model
All subcommands should build a normalized diagnostic result before rendering text or JSON.
Collector failures should be represented in the result instead of aborting the entire bundle, unless a required command input is invalid.
JSON compatibility
Machine-readable output is a public operational contract.
Rules:
schema_versionis required;latency_us,bandwidth_bps);Example:
{ "schema_version": 1, "command": "check-config", "build": { "commit": "<git-sha>", "tent": true, "cuda": true, "rdma": true }, "checks": [ { "id": "policy.device.exists", "outcome": "pass", "summary": "all configured policy devices were discovered" }, { "id": "rdma.gid.reachability", "outcome": "warn", "summary": "remote reachability was not tested", "remediation": ["run tent test-run --target <peer>"] } ], "redaction": { "mode": "default", "fields_redacted": 7 } }Redaction and security
Diagnostics frequently contain deployment-sensitive information. Redaction must happen at collection/model boundaries, not as a final regex over serialized JSON.
Always excluded by default
Redaction modes
Suggested modes:
Strict mode can use a random per-bundle salt so the same host/device remains correlatable inside one bundle without being correlatable across bundles.
Bundle handling
Remote testing
test-runand remoteshow-linkmust use existing authenticated/authorized control paths where available. The CLI must not introduce an unauthenticated arbitrary-memory test server. A standalone validation server, if needed later, requires an explicit security design and is out of v1 scope.Offline and live modes
Commands fall into two groups.
Offline/local discovery
versionshow-topologycheck-configshow-linkThese should work without a running TENT instance or metadata service when possible.
Live/remote diagnosis
show-linktest-runshow-planwith remote metadatastatuscollect-diagnosticsThese must report partial/unavailable state clearly when the peer or metadata service cannot be reached.
Exit codes
Suggested stable exit codes:
Warnings alone return zero unless
--strictis specified. JSON output must still be emitted for diagnostic failures whenever possible.Architecture
Collectors should be reusable libraries rather than logic embedded directly in
main(). Existing commands such asshow_linkcan call the same collector/rendering library.Optional collectors must be isolated by build feature. A CPU-only build should still provide config/topology checks and report GPU/RDMA collectors as unavailable rather than failing to build the CLI.
Compatibility and rollout
show_linkbehavior remains available during migration.Proposed PR sequence
PR1: Diagnostic core and schema
DiagnosticSnapshot, checks, evidence, renderers, redaction primitives, and schema tests.tent version.PR2: Topology and existing show-link integration
tent show-topologyandtent show-linkusing existing collectors.show_linkas a wrapper/alias if maintainers prefer.PR3: Config validation
tent check-config.PR4: Bounded connectivity test
tent test-runusing scratch registered buffers.PR5: Plan explain
tent show-planby consuming [RFC]: TENT Backend Capability Model and First-Class Execution Plan #2863.PR6: Live status and diagnostic bundle
tent statusandtent collect-diagnostics.Later work
tent benchconvenience integration with tebench;Validation plan
Unit and schema tests
Differential tests
tent show-linkmatches the existing [TransferEngine] Add show-link diagnostic tool for NIC topology #2820 collector output;tent check-configaccepts configurations accepted by the runtime and rejects known-invalid configurations before runtime initialization;tent show-planmatches [RFC]: TENT Backend Capability Model and First-Class Execution Plan #2863 explain output;tent benchforwards to tebench without changing benchmark semantics.Real-cluster validation
On the existing H20/RoCE nodes, cover:
For every failure, the CLI should identify the furthest verified stage and avoid claiming later stages passed. For example, a successful control probe must not be reported as successful RDMA data connectivity.
Overhead
Expected benefits and claims boundary
This RFC should provide:
It does not claim to fix a failing path automatically or improve transfer performance. Any future remediation or automatic health policy requires separate design and validation.
Relationship to existing work
show-link; do not duplicate its discovery code.Open questions
tentthe preferred binary name, or should the initial commands live under an existing Mooncake CLI namespace?show_linkbecome a compatibility wrapper aroundtent show-link, or remain independently installed?show-topology,check-config, andshow-link, or shouldtest-runalso be included?collect-diagnosticscreate a directory, JSON file, or compressed archive by default?benchremain visibly separate from tebench until the rest of the CLI stabilizes?Before submitting a new issue...