Repository navigation
[Lumen-RL] Receive trained weights over RCCL, transactionally - #2298
i-chaochen wants to merge 25 commits into
Conversation
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical weight-transaction and collective-coordination issues remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 5
Open (8)
Bucket decoder accepts duplicate names and overlapping byte ranges · New Uncoordinated receive failures can deadlock collective broadcasts · New Commit coverage uses source names instead of canonical parameter names · New FP8 coverage is recorded when requantization writes nothing · New Abort fails to clear pending expert relayout state · New Failed DP engines omit required per-TP rank results · New Update failures omit error responses and cause synchronous timeouts · New Contract test silently skips when the external checkout is unavailable · New
What changed in this PR
Adds transactional RCCL weight streaming, collective RPC routing, and capability discovery for RLHF rollouts.
Changes:
- Adds independent process groups and RDMA weight reception.
- Adds transactional weight application and serving fences.
- Adds collective RPC routing and capability negotiation.
- Expands CPU-only transport and wire-format tests.
| File | Description |
|---|---|
tests/test_rdma_weight_receiver.py |
Tests RDMA wire decoding and contracts. |
tests/test_collective_rpc_transport.py |
Tests TP RPC transport. |
tests/test_collective_rpc_dp.py |
Tests DP routing and correlation. |
tests/test_collective_rpc_dispatch.py |
Tests RPC dispatch behavior. |
tests/test_capabilities.py |
Tests capability negotiation. |
atom/utils/independent_process_group.py |
Creates isolated distributed groups. |
atom/rollout/weight_updater.py |
Applies transactional weight updates. |
atom/rollout/rdma_weight_receiver.py |
Receives and decodes RCCL weight streams. |
atom/rollout/model_runner_ext.py |
Integrates RLHF runner extensions. |
atom/rollout/capabilities.py |
Reports worker capabilities. |
atom/rollout/async_engine.py |
Exposes RPC and capability APIs. |
atom/model_engine/engine_utility.py |
Handles RPC dispatch and utility responses. |
atom/model_engine/engine_core_mgr.py |
Routes DP-level RPC responses. |
atom/model_engine/collective_rpc.py |
Defines RPC wire types and routing. |
atom/model_engine/capabilities.py |
Aggregates capabilities. |
atom/model_engine/async_proc.py |
Collects per-rank RPC replies. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
ZhangDanyang-AMD
left a comment
There was a problem hiding this comment.
The end goal is to run a 200-step ATOM + Megatron Qwen3 disaggregated case and match vLLM Disaggregated 2-Node Megatron + vLLM RDMA Deployment then run Running the eight examples from the release image cases 1 and 3 for 1 step to confirm other functionality is unaffected.
| """ | ||
| if not hasattr(self, "_expert_relayout_pending"): | ||
| self._expert_relayout_pending = {} | ||
| return self._expert_relayout_pending |
There was a problem hiding this comment.
abort_weight_update clears only the packed accumulator. A later reload can inherit expert-relayout state from the failed transaction. See 1177-1187
| # Release DP load bookkeeping now: an aborted seq may never emit a | ||
| # finished STREAM output, so relying on the finish path alone would leak | ||
| # its in-flight count. _release_seq_load is idempotent. | ||
| self._release_seq_load(req_id) |
There was a problem hiding this comment.
engine_utility.py:312-329 writes an unconditional response while it is consumed synchronously so non-RPC responses enter the legacy queue.
7be3e9e to
a67bc0b
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical RPC routing, RDMA stream validation, transaction coordination, and process-group isolation issues remain.
Review effort: Lite
Findings: 5
Open (8)
Concurrent RPC replies are misrouted between request waiters · New Early transaction rejection can desynchronize collective broadcasts · New End marker accepts invalid metadata and payload sizes · New DP stride uses logical instead of physical worker count · New Supplied stores are not namespaced by group name · New RDMA lifecycle methods are missing from capability discovery · New FP8 support is advertised based on method presence · New Contract test silently skips when the external checkout is unavailable
Resolved since last review (7)
Abort fails to clear pending expert relayout state FP8 coverage is recorded when requantization writes nothing Commit coverage uses source names instead of canonical parameter names Uncoordinated receive failures can deadlock collective broadcasts Bucket decoder accepts duplicate names and overlapping byte ranges Update failures omit error responses and cause synchronous timeouts Failed DP engines omit required per-TP rank results
a67bc0b to
79d2ea4
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical and moderate correctness issues remain in reload fencing, RPC handling, and capability reporting.
Review effort: Lite
Findings: 7
Open (10)
Successful non-commit reloads leave the weight update fence active · New Commit can succeed before synchronization faults are handled · New Supplied stores are not namespaced by group name DP stride uses logical instead of physical worker count End marker accepts invalid metadata and payload sizes Early transaction rejection can desynchronize collective broadcasts Concurrent RPC replies are misrouted between request waiters FP8 support is advertised based on method presence RDMA lifecycle methods are missing from capability discovery Contract test silently skips when the external checkout is unavailable
| # in place, so without this a failed stream would keep serving from a | ||
| # mix of two versions -- visible much later as unexplained divergence | ||
| # rather than as an error here. | ||
| self.assert_weight_update_ready() |
There was a problem hiding this comment.
Valid, fixed in d0785b7. The direct, SHM and IPC paths now lift the fence when a reload finishes (update_weights on return, SHM/IPC on the is_last bucket), provided no transaction is open. They verify no coverage, so the lift is skipped when a fused parameter is still waiting on shards or the reload matched nothing, and it is logged as a warning saying so. Verified coverage on those paths is the planned integrity follow-up.
79d2ea4 to
10dd72a
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved moderate issues can cause distributed hangs, rank misconfiguration, stream desynchronization, or incorrect update reporting.
Review effort: Lite
Findings: 7
Open (11)
Commit can succeed before synchronization faults are handled Successful non-commit reloads leave the weight update fence active Supplied stores are not namespaced by group name DP stride uses logical instead of physical worker count End marker accepts invalid metadata and payload sizes Early transaction rejection can desynchronize collective broadcasts Concurrent RPC replies are misrouted between request waiters SHM and IPC callers ignore FP8 update results · New FP8 support is advertised based on method presence RDMA lifecycle methods are missing from capability discovery Contract test silently skips when the external checkout is unavailable
10dd72a to
283e6dc
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved reload-safety, stream-validation, rank-calculation, capability-discovery, and process-group issues remain.
Review effort: Lite
Findings: 8
Open (12)
Prefill path bypasses reload readiness fence · New Commit can succeed before synchronization faults are handled Successful non-commit reloads leave the weight update fence active Supplied stores are not namespaced by group name DP stride uses logical instead of physical worker count End marker accepts invalid metadata and payload sizes Early transaction rejection can desynchronize collective broadcasts Concurrent RPC replies are misrouted between request waiters SHM and IPC callers ignore FP8 update results FP8 support is advertised based on method presence RDMA lifecycle methods are missing from capability discovery Contract test silently skips when the external checkout is unavailable
283e6dc to
3fe9a00
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved process-group, transactional coordination, coverage, and capability issues block approval.
Review effort: Lite
Findings: 9
Open (13)
Pass correct arguments when creating RDMA process groups · New Prefill path bypasses reload readiness fence Commit can succeed before synchronization faults are handled Successful non-commit reloads leave the weight update fence active Supplied stores are not namespaced by group name DP stride uses logical instead of physical worker count End marker accepts invalid metadata and payload sizes Early transaction rejection can desynchronize collective broadcasts Concurrent RPC replies are misrouted between request waiters SHM and IPC callers ignore FP8 update results FP8 support is advertised based on method presence RDMA lifecycle methods are missing from capability discovery Contract test silently skips when the external checkout is unavailable
The dead-rank check ran only on a poll that was about to keep waiting, so the deadline and died branches returned first. A rank found dead on the last poll -- or one whose turn came after the budget was spent -- left the others waiting at the barrier after the call had already given up on them, and every later call to them hung. The check now comes first in the empty-poll branch, ahead of both returns, and runs on every call rather than only barrier ones: a barrier call that timed out on a slow rank leaves the rest waiting, and if that rank dies afterwards, only a later call is there to notice. It is still a no-op until some rank has actually died. Co-authored-by: Cursor <cursoragent@cursor.com>
The handler checked the request id only for truthiness, so an unhashable one -- a list, say -- was broadcast, echoed back by every worker, and then raised inside RpcResponseRouter.route() on the manager's output thread. Nothing catches it there, so the thread ended and that engine's outputs went unread from then on. The handler's own error reply echoed the same id, so refusing it there alone was not enough; and an id-less collective_rpc reply fell through to the shared queue, where the next synchronous caller would have taken it as its own. The handler now requires a non-empty string id before it broadcasts, and the manager drops any collective_rpc reply without one instead of routing or queueing it. Co-authored-by: Cursor <cursoragent@cursor.com>
collective_rpc refused forward but not exit, so a caller could run ModelRunner.exit() on every TP worker -- tearing down its NCCL groups, KV connector and CUDA graphs -- after which busy_loop's own "exit" check ended each worker's loop, behind a reply that read as success. The same path also reached async_proc_aggregation, which drains the KV connector's finished transfers into a reply the KV aggregator never sees, so the requests waiting on them would stall. These are now one reserved set, built from the worker loop's own KV names so a new one cannot be missed, and refused before anything is enqueued. Co-authored-by: Cursor <cursoragent@cursor.com>
…ifecycle The TP-level collective_rpc said request ids let several calls be outstanding, but every caller reads the same per-rank reply queues, so two concurrent calls took each other's replies, dropped them as stale and timed out. Its one caller is the engine busy loop, one call at a time, so calls are now serialized under a lock, and the docstring says what the id is actually for: recognising a late reply from a call that already gave up. Capability discovery reported rdma_weight_receive with none of the methods that drive it. The advertised list now names the RDMA lifecycle and the weight-update status an orchestrator reads, each still listed only where the runner has it. Co-authored-by: Cursor <cursoragent@cursor.com>
Replies are collected rank by rank against one deadline, so a silent rank can spend the whole budget before the next rank's turn. That rank's reply is not lost meanwhile: each rank has its own queue, its output thread fills it as replies land, and a poll with no time left still takes what is already there. Pinned, so a change that returns at the deadline before looking cannot report a rank that answered as one that timed out. Co-authored-by: Cursor <cursoragent@cursor.com>
Lets each TP worker join an independent process group alongside the trainer, so weights land straight in resident GPU memory -- no safetensors round-trip, no shared folder, no host bounce. The wire format is frozen by the sending side and matched byte for byte: a 4x int64 header [command, metadata_bytes, payload_bytes, version], then a uint8 JSON metadata list, then a uint8 payload each entry views into. Decoded tensors are views over the payload rather than copies, because a transfer already sized in tens of GB cannot afford to double its peak footprint. A test reads the sender's own constants from source so a drift on either side fails here instead of decoding 61 GB into garbage. Two things make this more than a memcpy: Fused parameters. `qkv_proj` and `gate_up_proj` combine several checkpoint tensors, and one fused parameter's shards can arrive in different buckets, so state has to survive across them. `update_weights`' application loop is extracted into `_apply_named_tensors` and shared, leaving finalisation -- expert relayout, KV clear, accumulator teardown -- with the caller. A bucketed stream does that once at commit; doing it per bucket would relayout a half-built parameter. In-place application. A stream that fails halfway leaves the model a mix of two versions, and inference would keep serving, quietly wrong. So a stream is a transaction: begin/apply/commit/abort, with monotonic versions so a replayed or out-of-order stream is refused, and commit-time coverage verification against named_parameters(). `forward` now calls `assert_weight_update_ready()`, which refuses to serve mid-reload or after a partial one -- the cost of a false stop is a raised error, against rollouts from a half-updated model surfacing much later as unexplained divergence. Coverage accounting credits `weight_scale` values that requantisation writes as a side effect. They are derived, never sent, so without that they look permanently missing and verification would reject every stream. `init_independent_process_group` is vendored into atom/utils rather than imported from the RL framework: a rollout container carrying only ATOM must still be able to build the group. 11 new tests. Full suite 5712 -> 5723 passed, failures/skips/errors unchanged; #2028's 34 weight-sync tests still pass, which is the check that matters for the shared-loop extraction. Co-authored-by: Cursor <cursoragent@cursor.com>
Review follow-ups on the RDMA receiver and its transaction. Coverage was recorded under the incoming checkpoint name and checked against named_parameters(). On every fused route the two differ -- q_proj lands in qkv_proj, one expert's gate_proj in w13_weight -- so verify_full_load rejected every reload of a fused model, Qwen3 included. Coverage is now kept by parameter object. A packed parameter counts once all of its shards have landed; an expert buffer once every expert's every shard has, read off the relayout bookkeeping before the relayout consumes it. A tied parameter counts under either name. FP8 requantisation now reports whether it wrote. A shape it could not shard, or a quant type it did not know, counted as written and credited the scale as well. abort clears the pending expert relayout, and begin starts clean. Left behind, an aborted stream's slices were relaid out again by the next reload, or failed it on shards the aborted stream never sent. A rank whose bucket fails now keeps receiving until the end marker rather than leaving the broadcasts, which hung the trainer and every other rank in the next one. The decoder rejects repeated names and overlapping byte ranges, and a reload refuses a weight sent twice. The sender contract test no longer reads a hardcoded checkout path: it runs against LUMENRL_ROOT when that is set, and also pins the header field order. Co-authored-by: Cursor <cursoragent@cursor.com>
…ce on recovery Review follow-ups on the receiver and its transaction. A begin that refused the stream -- a replayed or older version, or a reload already open -- ran before the drain logic, so the rank left without receiving it while the trainer and every other rank waited in the next broadcast. It now runs inside it: the stream is received, nothing is applied, and the refusal is raised at the end marker like any failure. The end marker was accepted with any sizes. The contract is [END, 0, 0, version]; one carrying sizes now fails the stream instead of committing it. The DP stride was the logical TP width, but an engine runs tp_world_size x prefill_context_parallel_size workers. Under PCP one DP engine's workers took the next one's ranks, and under simulated TP ranks went unclaimed. The stride is now the engine's own worker count, the one ModelRunner places devices by. Only a commit lifted the fence an aborted stream leaves, so recovering through the direct, SHM or IPC path left every later forward refused. A legacy reload that finishes -- its last bucket applied, no fused parameter left waiting on shards, and not one that matched nothing -- now lifts it, with a warning that those paths verify no coverage. Co-authored-by: Cursor <cursoragent@cursor.com>
…l requantisations receive_weight_stream synchronized the device after its try block, so an asynchronous fault in the bucket writes, or in commit's own finalisation, surfaced after commit had declared the version good, and nothing fenced serving. The synchronize is now inside the try, and a fault there aborts. The SHM and IPC loops counted every FP8 requantisation as updated, though the requantiser returns without writing on a shape it cannot shard or a quant type it does not know. A reload of nothing else then reported success and lifted the fence over the old weights. They now branch on its result, as the transactional path already did. The helper's docstring said group_name namespaces the store. The isolation between groups is _new_process_group_helper's, which keys everything under the group name whatever store it is given; the extra prefix is only for a store created here, as init_process_group does, and the trainer's copy does the same. Said so, rather than changing the key layout on one end of the group only. A contract test composes the real mixins, so an advertised name the receiver spells differently cannot drop out of discovery unnoticed. Co-authored-by: Cursor <cursoragent@cursor.com>
…tream's state alone A non-transactional reload lifted the fence once it finished, and finishing says nothing about what it covered: an empty update_weights, or an SHM or IPC reload whose last bucket held one weight, put serving back over parameters the failed stream left as they were. While fenced, writes on every path are now counted the way a transaction counts them, from the moment the fence went up, and the fence lifts only once every parameter has been rewritten -- commit's own test. The SHM and IPC loops now share _apply_named_tensors with the direct and RDMA paths, which is where that counting lives, instead of keeping copies of it; and a fused expert buffer is credited expert by expert, after its relayout, so experts sent across several reloads add up. The receiver aborted on any failure, a begin that refused the stream included. That abort ends whatever reload is open, so a stream refused because another caller's reload was open tore that reload down, and a replayed version fenced a runner whose committed weights were fine. It now aborts only a reload it opened. A frame of an unknown command failed at its header, leaving the trainer and every other rank in the two broadcasts that follow every frame but the end marker. Those are now received first, by the sizes the header carries, and the frame fails the stream at the end marker like any bad bucket. Only a negative size, which no buffer can be posted for, still fails at once. Co-authored-by: Cursor <cursoragent@cursor.com>
main (#2419) narrowed the except around torch.cuda.ipc_collect() to AttributeError and RuntimeError. On a CPU-only torch, which is what the non-GPU CI job installs, ipc_collect raises AssertionError ("Torch not compiled with CUDA enabled") instead, so every IPC test that reaches a last bucket would fail there. Releasing the sender's mapping is not what those tests check, so it is stubbed out. Co-authored-by: Cursor <cursoragent@cursor.com>
…iver off a fence The receive loop held a bucket until its names were reassigned, so the previous one was still resident while the next was allocated -- two buckets at the peak -- and the last one stayed resident through commit. Draining after a failure held the failed bucket the same way, through the traceback of the error kept for the end marker, while every later bucket was allocated; running out of memory there would take this rank out of the very broadcasts the drain exists to keep it in. Buckets are now released at the top of each pass, and the kept error's frames are cleared. verify_full_load=False let any caller commit a reload without the coverage check. It stays -- LumenRL passes it to this backend and to vLLM alike -- but now waives only what it says: an update that resends part of the model, over a committed version that keeps serving the rest. Over a fence there is no such version, so there the check stands whatever was asked. An unverified commit says so in a warning, and the manifest and the receive stats carry how many parameters it did not rewrite. The duplicate-name report counted each name against the whole list; it is a Counter now. The sender checks read LumenRL's source and skipped wherever LUMENRL_ROOT was unset, CI included. They now read a verbatim snapshot of the sender's sending half, kept in tests/fixtures, and also check that each tensor is described by the keys the decoder reads; with LUMENRL_ROOT set, the snapshot is held to the sender itself. Co-authored-by: Cursor <cursoragent@cursor.com>
81342a9 to
91c9d29
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unbounded allocation can desynchronize RCCL participants, and hardware/protocol compatibility validation remains incomplete.
Review effort: Lite
Findings: 5
Open (11)
Bound header sizes and coordinate aborts before allocation · New RDMA load allows bypassing full-coverage verification Previous bucket payload remains alive during next allocation Successful non-commit reloads leave the weight update fence active Concurrent RPC replies are misrouted between request waiters Add multi-process RCCL smoke coverage for RDMA path · New Enforce sender compatibility checks in CI · New Duplicate-name validation is quadratic for large buckets FP8 support is advertised based on method presence RDMA lifecycle methods are missing from capability discovery Sender compatibility tests are skipped without CI fixture
…oad; test a real group A header whose sizes this rank cannot allocate for -- [BUCKET, 1, 2**63-1, v], say -- raised from torch.empty while the trainer and every peer were already in that frame's broadcasts. A broadcast is received whole, so no drain can follow it. Sizes are now held to explicit ceilings, and on a GPU to the memory free for them, before anything is allocated; a frame that fails either, or whose allocation fails anyway, raises RDMAStreamOutOfStep naming the header, and receive_weights_rdma tears down this rank's end of the group so nothing later rides it out of step. The sender and any peer still in the broadcast are released by the group's timeout. The full-load check can no longer be waived. A partial commit served the rest of the model from another version, the state the transaction exists to rule out: commit_weight_update takes no waiver, and receive_weights_rdma still accepts verify_full_load -- LumenRL drives the vLLM worker with the same arguments -- but ignores False, with a warning. The receiver's tests stubbed the group and dist.broadcast throughout. A new module runs the trainer and the receivers as separate processes over a group init_independent_process_group builds: over gloo everywhere, CI included -- a clean stream, then one in which a rank fails and drains -- and over RCCL, with LumenRL's own sender, wherever there are two GPUs. The sender contract now runs in CI against LumenRL itself: a workflow checks LumenRL out and runs the contract tests with LUMENRL_ROOT set, on any change to the receiver, its tests or the snapshot, and daily. Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
One critical and four moderate unresolved findings remain.
Review effort: Lite
Findings: 3
Open (4)
Resolved since last review (8)
Bound header sizes and coordinate aborts before allocation RDMA load allows bypassing full-coverage verification Previous bucket payload remains alive during next allocation Enforce sender compatibility checks in CI Add multi-process RCCL smoke coverage for RDMA path Duplicate-name validation is quadratic for large buckets FP8 support is advertised based on method presence Sender compatibility tests are skipped without CI fixture
update_weights, update_weights_shm and update_weights_ipc ran through call_func(..., wait_out=True), which returns rank 0's result alone: a rank that rejected a tensor while rank 0 succeeded reported success, and the caller went on with a partly updated model -- when that rank's raise had not ended its worker loop outright. All three now go through the generic collective_rpc path, where every rank answers on its own channel and a failure is caught where it happens; the reply is rank 0's count only if every rank succeeded, and otherwise names each rank that failed. Co-authored-by: Cursor <cursoragent@cursor.com>
…e group on any break-off A begin refused on one TP rank left that rank serving its old weights while its peers committed the same stream: one engine holding two versions. Commit is now prepare and finish. Each rank prepares and synchronizes, the engine's TP ranks vote over their CPU group, and a rank finishes only if every rank can. If the vote fails and any rank began the stream, every rank ends fenced -- one that refused through the new fence_weight_update, which leaves a reload already open there to its caller. Every rank votes exactly once per stream, on its way out too, so a peer in the vote is never left waiting. Only the rank whose own frame could not be received tore its end of the group down. Its peers failed in the broadcast it never joined, kept the group, and a later stream rode it out of step. Any failure before the end marker now raises RDMAStreamOutOfStep, so every rank that cannot finish the stream leaves the group. The multi-process test's receivers now stand for one engine's TP ranks and vote over a gloo group of their own: a clean stream commits on both, and a rank that fails a bucket, or refuses the stream, leaves neither serving it. Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical correctness, memory-safety, and CI path issues remain.
Review effort: Lite
Findings: 5
Open (5)
LUMENRL_ROOT path is duplicated in contract test workflow · New All-rank wrapper reports success despite skipped weight updates · New tolist() causes excessive memory use for maximum metadata frames · New TP-only vote omits PCP workers from transaction-wide agreement · New Successful non-commit reloads leave the weight update fence active
| persist-credentials: false | ||
| - name: Run the sender contract tests | ||
| env: | ||
| LUMENRL_ROOT: ${{ github.workspace }}/lumenrl |
| failed = [r for r in replies if not r.ok] | ||
| if failed: | ||
| error = "; ".join(f"TP rank {r.tp_rank}: {r.error}" for r in failed) | ||
| logger.error( | ||
| f"{self.label}: {cmd} failed on {len(failed)} of {len(replies)} " | ||
| f"TP rank(s): {error}" | ||
| ) | ||
| return {"cmd": cmd, "error": error} | ||
| result = replies[0].value |
| application, and copying here would double the peak footprint of a transfer | ||
| already sized in tens of GB. | ||
| """ | ||
| metadata = json.loads(bytes(metadata_tensor.cpu().tolist()).decode("utf-8")) |
| def _vote_with_tp_peers(self, ready: bool, began: bool) -> tuple[bool, bool]: | ||
| """The stream's decision across this engine's TP ranks, every one of | ||
| which takes it: one rank's yes is not enough to serve it.""" | ||
| if int(getattr(self, "world_size", 1) or 1) <= 1: | ||
| return ready, began | ||
| from aiter.dist.parallel_state import get_tp_group | ||
|
|
||
| tp = get_tp_group() | ||
| if tp.world_size <= 1: | ||
| return ready, began | ||
| return vote_on_stream(tp.cpu_group, ready, began) |



Motivation
Stacked on #2297
ATOM rollout has no direct weight path from a trainer: weights arrive via
safetensors on shared storage or a ZMQ CUDA-IPC hop. This lets each TP worker
join an independent process group alongside the trainer, so weights land
straight in resident GPU memory.
Technical Details
The wire format is frozen by the sending side and matched byte for byte: a
4x int64 header
[command, metadata_bytes, payload_bytes, version], a uint8JSON metadata list, then a uint8 payload each entry views into. Decoded tensors
are views over the payload rather than copies, since a transfer already sized in
tens of GB cannot afford to double its peak footprint. A test reads the sender's
own constants from source, so drift on either side fails here rather than
decoding tens of GB into garbage.
Two things make this more than a memcpy.
Fused parameters:
qkv_projandgate_up_projcombine several checkpointtensors, and one fused parameter's shards can arrive in different buckets, so
state must survive across them.
update_weights' application loop is extractedinto
_apply_named_tensorsand shared, leaving finalisation -- expert relayout,KV clear, accumulator teardown -- with the caller. A bucketed stream does that
once at commit; per bucket it would relayout a half-built parameter.
In-place application: a stream that fails halfway leaves the model a mix of two
versions, and inference would keep serving, quietly wrong. So a stream is a
transaction -- begin/apply/commit/abort -- with monotonic versions so a replayed
or out-of-order stream is refused, and commit-time coverage verification against
named_parameters().forwardnow callsassert_weight_update_ready(), whichrefuses to serve mid-reload or after a partial one. The cost of a false stop is
a raised error; the cost of not stopping is rollouts from a half-updated model,
surfacing much later as unexplained divergence.
Coverage accounting credits
weight_scalevalues that requantisation writes asa side effect. They are derived, never sent, so without that they look
permanently missing and verification would reject every stream.
init_independent_process_groupis vendored intoatom/utilsrather thanimported from the RL framework, so a rollout container carrying only ATOM can
still build the group.
Note the coverage check currently gates the RDMA path only; extending it to the
shm and IPC transports is follow-up work.
Test Plan
11 new CPU-only tests covering the wire-format round trip, view-not-copy, and
rejection of every corrupt-metadata shape, plus a test pinning the sender's
constants. Re-ran #2028's weight-sync tests, which are the real check on the
shared-loop extraction.
Test Result
black --check .clean;ruffno findings on touched filesNot yet validated on hardware. The 9-rank run needs two GPU nodes
simultaneously and has not been schedulable yet.
Submission Checklist