Skip to content

[Lumen-RL] Receive trained weights over RCCL, transactionally - #2298

Open
i-chaochen wants to merge 25 commits into
mainfrom
chao/recv_weights
Open

i-chaochen wants to merge 25 commits into
mainfrom
chao/recv_weights

Conversation

@i-chaochen

Copy link
Copy Markdown

Motivation

Stacked on #2297

ATOM rollout has no direct weight path from a trainer: weights arrive via
safetensors on shared storage or a ZMQ CUDA-IPC hop. This lets each TP worker
join an independent process group alongside the trainer, so weights land
straight in resident GPU memory.

Technical Details

The wire format is frozen by the sending side and matched byte for byte: a
4x int64 header [command, metadata_bytes, payload_bytes, version], a uint8
JSON metadata list, then a uint8 payload each entry views into. Decoded tensors
are views over the payload rather than copies, since a transfer already sized in
tens of GB cannot afford to double its peak footprint. A test reads the sender's
own constants from source, so drift on either side fails here rather than
decoding tens of GB into garbage.

Two things make this more than a memcpy.

Fused parameters: qkv_proj and gate_up_proj combine several checkpoint
tensors, and one fused parameter's shards can arrive in different buckets, so
state must survive across them. update_weights' application loop is extracted
into _apply_named_tensors and shared, leaving finalisation -- expert relayout,
KV clear, accumulator teardown -- with the caller. A bucketed stream does that
once at commit; per bucket it would relayout a half-built parameter.

In-place application: a stream that fails halfway leaves the model a mix of two
versions, and inference would keep serving, quietly wrong. So a stream is a
transaction -- begin/apply/commit/abort -- with monotonic versions so a replayed
or out-of-order stream is refused, and commit-time coverage verification against
named_parameters(). forward now calls assert_weight_update_ready(), which
refuses to serve mid-reload or after a partial one. The cost of a false stop is
a raised error; the cost of not stopping is rollouts from a half-updated model,
surfacing much later as unexplained divergence.

Coverage accounting credits weight_scale values that requantisation writes as
a side effect. They are derived, never sent, so without that they look
permanently missing and verification would reject every stream.

init_independent_process_group is vendored into atom/utils rather than
imported from the RL framework, so a rollout container carrying only ATOM can
still build the group.

Note the coverage check currently gates the RDMA path only; extending it to the
shm and IPC transports is follow-up work.

Test Plan

11 new CPU-only tests covering the wire-format round trip, view-not-copy, and
rejection of every corrupt-metadata shape, plus a test pinning the sender's
constants. Re-ran #2028's weight-sync tests, which are the real check on the
shared-loop extraction.

Test Result

Not yet validated on hardware. The 9-rank run needs two GPU nodes
simultaneously and has not been schedulable yet.

Submission Checklist

@i-chaochen
i-chaochen requested review from ZhangDanyang-AMD and valarLip and a lite review from Copilot September 19, 2026 17:54
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2298 --add-label <label>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical weight-transaction and collective-coordination issues remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 5 High severity · 2 Medium severity · 1 Low severity

Open (8)
What changed in this PR

Adds transactional RCCL weight streaming, collective RPC routing, and capability discovery for RLHF rollouts.

Changes:

  • Adds independent process groups and RDMA weight reception.
  • Adds transactional weight application and serving fences.
  • Adds collective RPC routing and capability negotiation.
  • Expands CPU-only transport and wire-format tests.
File Description
tests/​test_rdma_weight_receiver.py Tests RDMA wire decoding and contracts.
tests/​test_collective_rpc_transport.py Tests TP RPC transport.
tests/​test_collective_rpc_dp.py Tests DP routing and correlation.
tests/​test_collective_rpc_dispatch.py Tests RPC dispatch behavior.
tests/​test_capabilities.py Tests capability negotiation.
atom/​utils/​independent_process_group.py Creates isolated distributed groups.
atom/​rollout/​weight_updater.py Applies transactional weight updates.
atom/​rollout/​rdma_weight_receiver.py Receives and decodes RCCL weight streams.
atom/​rollout/​model_runner_ext.py Integrates RLHF runner extensions.
atom/​rollout/​capabilities.py Reports worker capabilities.
atom/​rollout/​async_engine.py Exposes RPC and capability APIs.
atom/​model_engine/​engine_utility.py Handles RPC dispatch and utility responses.
atom/​model_engine/​engine_core_mgr.py Routes DP-level RPC responses.
atom/​model_engine/​collective_rpc.py Defines RPC wire types and routing.
atom/​model_engine/​capabilities.py Aggregates capabilities.
atom/​model_engine/​async_proc.py Collects per-rank RPC replies.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread atom/rollout/rdma_weight_receiver.py
Comment thread atom/rollout/rdma_weight_receiver.py
Comment thread atom/rollout/weight_updater.py Outdated
Comment thread atom/rollout/weight_updater.py Outdated
Comment thread atom/rollout/weight_updater.py Outdated
Comment thread atom/model_engine/engine_core_mgr.py Outdated
Comment thread atom/model_engine/engine_utility.py Outdated
Comment thread tests/test_rdma_weight_receiver.py Outdated

@ZhangDanyang-AMD ZhangDanyang-AMD left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The end goal is to run a 200-step ATOM + Megatron Qwen3 disaggregated case and match vLLM Disaggregated 2-Node Megatron + vLLM RDMA Deployment then run Running the eight examples from the release image cases 1 and 3 for 1 step to confirm other functionality is unaffected.

Comment thread atom/model_engine/engine_core_mgr.py Outdated
"""
if not hasattr(self, "_expert_relayout_pending"):
self._expert_relayout_pending = {}
return self._expert_relayout_pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

abort_weight_update clears only the packed accumulator. A later reload can inherit expert-relayout state from the failed transaction. See 1177-1187

# Release DP load bookkeeping now: an aborted seq may never emit a
# finished STREAM output, so relying on the finish path alone would leak
# its in-flight count. _release_seq_load is idempotent.
self._release_seq_load(req_id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

engine_utility.py:312-329 writes an unconditional response while it is consumed synchronously so non-RPC responses enter the legacy queue.

Copilot AI review requested due to automatic review settings September 28, 2026 11:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Comment thread atom/model_engine/async_proc.py Outdated
Comment thread atom/rollout/rdma_weight_receiver.py Outdated
Comment thread atom/rollout/rdma_weight_receiver.py
Comment thread atom/rollout/rdma_weight_receiver.py Outdated
Comment thread atom/utils/independent_process_group.py
Comment thread atom/rollout/capabilities.py
Comment thread atom/rollout/capabilities.py Outdated
Copilot AI review requested due to automatic review settings September 28, 2026 13:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

# in place, so without this a failed stream would keep serving from a
# mix of two versions -- visible much later as unexplained divergence
# rather than as an error here.
self.assert_weight_update_ready()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, fixed in d0785b7. The direct, SHM and IPC paths now lift the fence when a reload finishes (update_weights on return, SHM/IPC on the is_last bucket), provided no transaction is open. They verify no coverage, so the lift is skipped when a fused parameter is still waiting on shards or the reload matched nothing, and it is logged as a warning saying so. Verified coverage on those paths is the planned integrity follow-up.

Comment thread atom/rollout/rdma_weight_receiver.py Outdated
Copilot AI review requested due to automatic review settings September 29, 2026 13:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread atom/rollout/weight_updater.py
Copilot AI review requested due to automatic review settings September 29, 2026 14:16

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread atom/rollout/model_runner_ext.py
Copilot AI review requested due to automatic review settings September 29, 2026 14:37

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Comment thread atom/utils/independent_process_group.py
Copilot AI lite review requested due to automatic review settings September 30, 2026 11:03
Chen and others added 12 commits October 2, 2026 15:27
The dead-rank check ran only on a poll that was about to keep waiting,
so the deadline and died branches returned first. A rank found dead on
the last poll -- or one whose turn came after the budget was spent --
left the others waiting at the barrier after the call had already given
up on them, and every later call to them hung.

The check now comes first in the empty-poll branch, ahead of both
returns, and runs on every call rather than only barrier ones: a barrier
call that timed out on a slow rank leaves the rest waiting, and if that
rank dies afterwards, only a later call is there to notice. It is still
a no-op until some rank has actually died.

Co-authored-by: Cursor <cursoragent@cursor.com>
The handler checked the request id only for truthiness, so an unhashable
one -- a list, say -- was broadcast, echoed back by every worker, and then
raised inside RpcResponseRouter.route() on the manager's output thread.
Nothing catches it there, so the thread ended and that engine's outputs
went unread from then on. The handler's own error reply echoed the same
id, so refusing it there alone was not enough; and an id-less
collective_rpc reply fell through to the shared queue, where the next
synchronous caller would have taken it as its own.

The handler now requires a non-empty string id before it broadcasts, and
the manager drops any collective_rpc reply without one instead of routing
or queueing it.

Co-authored-by: Cursor <cursoragent@cursor.com>
collective_rpc refused forward but not exit, so a caller could run
ModelRunner.exit() on every TP worker -- tearing down its NCCL groups, KV
connector and CUDA graphs -- after which busy_loop's own "exit" check
ended each worker's loop, behind a reply that read as success. The same
path also reached async_proc_aggregation, which drains the KV connector's
finished transfers into a reply the KV aggregator never sees, so the
requests waiting on them would stall.

These are now one reserved set, built from the worker loop's own KV names
so a new one cannot be missed, and refused before anything is enqueued.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ifecycle

The TP-level collective_rpc said request ids let several calls be
outstanding, but every caller reads the same per-rank reply queues, so two
concurrent calls took each other's replies, dropped them as stale and
timed out. Its one caller is the engine busy loop, one call at a time, so
calls are now serialized under a lock, and the docstring says what the id
is actually for: recognising a late reply from a call that already gave
up.

Capability discovery reported rdma_weight_receive with none of the methods
that drive it. The advertised list now names the RDMA lifecycle and the
weight-update status an orchestrator reads, each still listed only where
the runner has it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Replies are collected rank by rank against one deadline, so a silent rank
can spend the whole budget before the next rank's turn. That rank's reply
is not lost meanwhile: each rank has its own queue, its output thread
fills it as replies land, and a poll with no time left still takes what
is already there. Pinned, so a change that returns at the deadline before
looking cannot report a rank that answered as one that timed out.

Co-authored-by: Cursor <cursoragent@cursor.com>
Lets each TP worker join an independent process group alongside the trainer,
so weights land straight in resident GPU memory -- no safetensors round-trip,
no shared folder, no host bounce.

The wire format is frozen by the sending side and matched byte for byte: a
4x int64 header [command, metadata_bytes, payload_bytes, version], then a
uint8 JSON metadata list, then a uint8 payload each entry views into. Decoded
tensors are views over the payload rather than copies, because a transfer
already sized in tens of GB cannot afford to double its peak footprint.
A test reads the sender's own constants from source so a drift on either side
fails here instead of decoding 61 GB into garbage.

Two things make this more than a memcpy:

Fused parameters. `qkv_proj` and `gate_up_proj` combine several checkpoint
tensors, and one fused parameter's shards can arrive in different buckets, so
state has to survive across them. `update_weights`' application loop is
extracted into `_apply_named_tensors` and shared, leaving finalisation --
expert relayout, KV clear, accumulator teardown -- with the caller. A
bucketed stream does that once at commit; doing it per bucket would relayout
a half-built parameter.

In-place application. A stream that fails halfway leaves the model a mix of
two versions, and inference would keep serving, quietly wrong. So a stream is
a transaction: begin/apply/commit/abort, with monotonic versions so a
replayed or out-of-order stream is refused, and commit-time coverage
verification against named_parameters(). `forward` now calls
`assert_weight_update_ready()`, which refuses to serve mid-reload or after a
partial one -- the cost of a false stop is a raised error, against rollouts
from a half-updated model surfacing much later as unexplained divergence.

Coverage accounting credits `weight_scale` values that requantisation writes
as a side effect. They are derived, never sent, so without that they look
permanently missing and verification would reject every stream.

`init_independent_process_group` is vendored into atom/utils rather than
imported from the RL framework: a rollout container carrying only ATOM must
still be able to build the group.

11 new tests. Full suite 5712 -> 5723 passed, failures/skips/errors
unchanged; #2028's 34 weight-sync tests still pass, which is the check that
matters for the shared-loop extraction.

Co-authored-by: Cursor <cursoragent@cursor.com>
Review follow-ups on the RDMA receiver and its transaction.

Coverage was recorded under the incoming checkpoint name and checked
against named_parameters(). On every fused route the two differ -- q_proj
lands in qkv_proj, one expert's gate_proj in w13_weight -- so
verify_full_load rejected every reload of a fused model, Qwen3 included.
Coverage is now kept by parameter object. A packed parameter counts once
all of its shards have landed; an expert buffer once every expert's every
shard has, read off the relayout bookkeeping before the relayout consumes
it. A tied parameter counts under either name.

FP8 requantisation now reports whether it wrote. A shape it could not
shard, or a quant type it did not know, counted as written and credited
the scale as well.

abort clears the pending expert relayout, and begin starts clean. Left
behind, an aborted stream's slices were relaid out again by the next
reload, or failed it on shards the aborted stream never sent.

A rank whose bucket fails now keeps receiving until the end marker rather
than leaving the broadcasts, which hung the trainer and every other rank
in the next one. The decoder rejects repeated names and overlapping byte
ranges, and a reload refuses a weight sent twice.

The sender contract test no longer reads a hardcoded checkout path: it
runs against LUMENRL_ROOT when that is set, and also pins the header
field order.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ce on recovery

Review follow-ups on the receiver and its transaction.

A begin that refused the stream -- a replayed or older version, or a
reload already open -- ran before the drain logic, so the rank left
without receiving it while the trainer and every other rank waited in the
next broadcast. It now runs inside it: the stream is received, nothing is
applied, and the refusal is raised at the end marker like any failure.

The end marker was accepted with any sizes. The contract is
[END, 0, 0, version]; one carrying sizes now fails the stream instead of
committing it.

The DP stride was the logical TP width, but an engine runs tp_world_size
x prefill_context_parallel_size workers. Under PCP one DP engine's workers
took the next one's ranks, and under simulated TP ranks went unclaimed.
The stride is now the engine's own worker count, the one ModelRunner
places devices by.

Only a commit lifted the fence an aborted stream leaves, so recovering
through the direct, SHM or IPC path left every later forward refused. A
legacy reload that finishes -- its last bucket applied, no fused parameter
left waiting on shards, and not one that matched nothing -- now lifts it,
with a warning that those paths verify no coverage.

Co-authored-by: Cursor <cursoragent@cursor.com>
…l requantisations

receive_weight_stream synchronized the device after its try block, so an
asynchronous fault in the bucket writes, or in commit's own finalisation,
surfaced after commit had declared the version good, and nothing fenced
serving. The synchronize is now inside the try, and a fault there aborts.

The SHM and IPC loops counted every FP8 requantisation as updated, though
the requantiser returns without writing on a shape it cannot shard or a
quant type it does not know. A reload of nothing else then reported
success and lifted the fence over the old weights. They now branch on its
result, as the transactional path already did.

The helper's docstring said group_name namespaces the store. The isolation
between groups is _new_process_group_helper's, which keys everything under
the group name whatever store it is given; the extra prefix is only for a
store created here, as init_process_group does, and the trainer's copy
does the same. Said so, rather than changing the key layout on one end of
the group only.

A contract test composes the real mixins, so an advertised name the
receiver spells differently cannot drop out of discovery unnoticed.

Co-authored-by: Cursor <cursoragent@cursor.com>
…tream's state alone

A non-transactional reload lifted the fence once it finished, and finishing
says nothing about what it covered: an empty update_weights, or an SHM or
IPC reload whose last bucket held one weight, put serving back over
parameters the failed stream left as they were. While fenced, writes on
every path are now counted the way a transaction counts them, from the
moment the fence went up, and the fence lifts only once every parameter has
been rewritten -- commit's own test. The SHM and IPC loops now share
_apply_named_tensors with the direct and RDMA paths, which is where that
counting lives, instead of keeping copies of it; and a fused expert buffer
is credited expert by expert, after its relayout, so experts sent across
several reloads add up.

The receiver aborted on any failure, a begin that refused the stream
included. That abort ends whatever reload is open, so a stream refused
because another caller's reload was open tore that reload down, and a
replayed version fenced a runner whose committed weights were fine. It now
aborts only a reload it opened.

A frame of an unknown command failed at its header, leaving the trainer and
every other rank in the two broadcasts that follow every frame but the end
marker. Those are now received first, by the sizes the header carries, and
the frame fails the stream at the end marker like any bad bucket. Only a
negative size, which no buffer can be posted for, still fails at once.

Co-authored-by: Cursor <cursoragent@cursor.com>
main (#2419) narrowed the except around torch.cuda.ipc_collect() to
AttributeError and RuntimeError. On a CPU-only torch, which is what the
non-GPU CI job installs, ipc_collect raises AssertionError ("Torch not
compiled with CUDA enabled") instead, so every IPC test that reaches a last
bucket would fail there. Releasing the sender's mapping is not what those
tests check, so it is stubbed out.

Co-authored-by: Cursor <cursoragent@cursor.com>
…iver off a fence

The receive loop held a bucket until its names were reassigned, so the
previous one was still resident while the next was allocated -- two
buckets at the peak -- and the last one stayed resident through commit.
Draining after a failure held the failed bucket the same way, through the
traceback of the error kept for the end marker, while every later bucket
was allocated; running out of memory there would take this rank out of the
very broadcasts the drain exists to keep it in. Buckets are now released at
the top of each pass, and the kept error's frames are cleared.

verify_full_load=False let any caller commit a reload without the coverage
check. It stays -- LumenRL passes it to this backend and to vLLM alike --
but now waives only what it says: an update that resends part of the
model, over a committed version that keeps serving the rest. Over a fence
there is no such version, so there the check stands whatever was asked. An
unverified commit says so in a warning, and the manifest and the receive
stats carry how many parameters it did not rewrite.

The duplicate-name report counted each name against the whole list; it is
a Counter now.

The sender checks read LumenRL's source and skipped wherever LUMENRL_ROOT
was unset, CI included. They now read a verbatim snapshot of the sender's
sending half, kept in tests/fixtures, and also check that each tensor is
described by the keys the decoder reads; with LUMENRL_ROOT set, the
snapshot is held to the sender itself.

Co-authored-by: Cursor <cursoragent@cursor.com>
Copilot AI lite review requested due to automatic review settings October 2, 2026 16:08

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread atom/rollout/rdma_weight_receiver.py Outdated
Comment thread atom/rollout/rdma_weight_receiver.py
Comment thread tests/test_rdma_weight_receiver.py
…oad; test a real group

A header whose sizes this rank cannot allocate for -- [BUCKET, 1, 2**63-1,
v], say -- raised from torch.empty while the trainer and every peer were
already in that frame's broadcasts. A broadcast is received whole, so no
drain can follow it. Sizes are now held to explicit ceilings, and on a GPU
to the memory free for them, before anything is allocated; a frame that
fails either, or whose allocation fails anyway, raises RDMAStreamOutOfStep
naming the header, and receive_weights_rdma tears down this rank's end of
the group so nothing later rides it out of step. The sender and any peer
still in the broadcast are released by the group's timeout.

The full-load check can no longer be waived. A partial commit served the
rest of the model from another version, the state the transaction exists
to rule out: commit_weight_update takes no waiver, and receive_weights_rdma
still accepts verify_full_load -- LumenRL drives the vLLM worker with the
same arguments -- but ignores False, with a warning.

The receiver's tests stubbed the group and dist.broadcast throughout. A new
module runs the trainer and the receivers as separate processes over a
group init_independent_process_group builds: over gloo everywhere, CI
included -- a clean stream, then one in which a rank fails and drains -- and
over RCCL, with LumenRL's own sender, wherever there are two GPUs.

The sender contract now runs in CI against LumenRL itself: a workflow checks
LumenRL out and runs the contract tests with LUMENRL_ROOT set, on any change
to the receiver, its tests or the snapshot, and daily.

Co-authored-by: Cursor <cursoragent@cursor.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 13:33

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Comment thread atom/model_engine/engine_utility.py Outdated
Copilot AI lite review requested due to automatic review settings October 5, 2026 13:40

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical cross-rank transaction and teardown issues remain unresolved, along with a missing forward-readiness regression test.

Review effort: Lite
Findings: 4 High severity

Open (4)
Resolved since last review (2)

Comment thread atom/rollout/rdma_weight_receiver.py
Comment thread atom/rollout/rdma_weight_receiver.py
Chen and others added 3 commits October 5, 2026 14:31
update_weights, update_weights_shm and update_weights_ipc ran through
call_func(..., wait_out=True), which returns rank 0's result alone: a rank
that rejected a tensor while rank 0 succeeded reported success, and the
caller went on with a partly updated model -- when that rank's raise had
not ended its worker loop outright. All three now go through the generic
collective_rpc path, where every rank answers on its own channel and a
failure is caught where it happens; the reply is rank 0's count only if
every rank succeeded, and otherwise names each rank that failed.

Co-authored-by: Cursor <cursoragent@cursor.com>
…e group on any break-off

A begin refused on one TP rank left that rank serving its old weights while
its peers committed the same stream: one engine holding two versions. Commit
is now prepare and finish. Each rank prepares and synchronizes, the engine's
TP ranks vote over their CPU group, and a rank finishes only if every rank
can. If the vote fails and any rank began the stream, every rank ends
fenced -- one that refused through the new fence_weight_update, which leaves
a reload already open there to its caller. Every rank votes exactly once
per stream, on its way out too, so a peer in the vote is never left waiting.

Only the rank whose own frame could not be received tore its end of the
group down. Its peers failed in the broadcast it never joined, kept the
group, and a later stream rode it out of step. Any failure before the end
marker now raises RDMAStreamOutOfStep, so every rank that cannot finish the
stream leaves the group.

The multi-process test's receivers now stand for one engine's TP ranks and
vote over a gloo group of their own: a clean stream commits on both, and a
rank that fails a bucket, or refuses the stream, leaves neither serving it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Copilot AI lite review requested due to automatic review settings October 6, 2026 11:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

persist-credentials: false
- name: Run the sender contract tests
env:
LUMENRL_ROOT: ${{ github.workspace }}/lumenrl
Comment on lines +262 to +270
failed = [r for r in replies if not r.ok]
if failed:
error = "; ".join(f"TP rank {r.tp_rank}: {r.error}" for r in failed)
logger.error(
f"{self.label}: {cmd} failed on {len(failed)} of {len(replies)} "
f"TP rank(s): {error}"
)
return {"cmd": cmd, "error": error}
result = replies[0].value
application, and copying here would double the peak footprint of a transfer
already sized in tens of GB.
"""
metadata = json.loads(bytes(metadata_tensor.cpu().tolist()).decode("utf-8"))
Comment on lines +461 to +471
def _vote_with_tp_peers(self, ready: bool, began: bool) -> tuple[bool, bool]:
"""The stream's decision across this engine's TP ranks, every one of
which takes it: one rank's yes is not enough to serve it."""
if int(getattr(self, "world_size", 1) or 1) <= 1:
return ready, began
from aiter.dist.parallel_state import get_tp_group

tp = get_tp_group()
if tp.world_size <= 1:
return ready, began
return vote_on_stream(tp.cpu_group, ready, began)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants