Skip to content

UTS: add translator-skill guide and uts-to-lang-skill-creator skill - #558

Open
sacOO7 wants to merge 22 commits into
mainfrom
uts/translator-skills-docs
Open

sacOO7 wants to merge 22 commits into
mainfrom
uts/translator-skills-docs

Conversation

@sacOO7

@sacOO7 sacOO7 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR adds two things that help an SDK team build a uts-to-<lang> translator skill and the harness (UTS test infrastructure) it runs on. A translator skill turns UTS pseudocode specs into native tests for one SDK (uts-to-python, uts-to-csharp, uts-to-go, …).

  • A guide, uts/docs/writing-uts-spec-translator-skills.md. It says what a translator skill and its harness must contain, and why, using MUST / SHOULD / MAY.
  • An agent skill, uts/skills/uts-to-lang-skill-creator/. An agent follows it inside an SDK repo to create that repo's uts-to-<lang> skill and harness, or to audit and upgrade an existing one. It is packaged as an Agent Skill, so the same files load in Claude Code and in OpenAI Codex.

Two translator skills exist today: uts-to-swift in ably-cocoa and uts-to-kotlin in ably-java. Until now, what they learned lived only in those repos. This PR moves it into the spec repo, next to the UTS docs. The guide and the skill are self-contained. The two existing skills are named only as a guarded, last-resort reference.

Why

A translator skill makes translating a spec close to mechanical:

  • Easy to use: one command (/uts-to-<lang> <module-dir> in Claude Code, $uts-to-<lang> <module-dir> in Codex) and four questions.
  • Full context on every run: the model gets the spec, the translation rules, the harness, the module notes and the resolver output each time.
  • Reproducible: when a translation is wrong or the spec changes, you fix the skill, the notes or the harness, then regenerate the test. You don't hand-maintain generated tests.

What's added or changed

Path Reader Purpose
uts/docs/writing-uts-spec-translator-skills.md SDK engineers What a translator skill and its harness must contain, and why
uts/skills/uts-to-lang-skill-creator/ An agent (Claude Code or Codex) working in an SDK repo How to build or upgrade them: stop points, templates, scripts and acceptance checks
uts/README.md Everyone Adds the guide and the skill to the docs tree and the Guides list, with install steps. Also fixes the tree and spec counts, which were missing objects/ and the REST proxy tier
uts/docs/writing-derived-tests.md Everyone One paragraph in the Overview linking to the guide. That document remains the authority on translation and evaluation semantics

The guide

  1. Purpose and scope. What a translator skill is; what it owns (harness mechanics) and what it defers to writing-derived-tests.md (translation semantics); and the model tier (an Opus-class model is required to create the skill and harness, and recommended for translation runs). Translator skills, and the skill creator, are for Ably Pub/Sub SDK repositories.

  2. The harness. The harness and the skill are one deliverable. This section covers:

    • SDK test hooks;
    • a shared test library implementing the helper specs (mock_http, mock_websocket, mock_vcdiff, standard_test_pool);
    • sandbox provisioning, the uts-proxy, and build and CI wiring;
    • the SDK's concurrency and time model;
    • permanent, CI-run smoke tests and helper self-tests, with a minimum set per tier.

    Scope follows the SDK's capabilities: a REST-only or realtime-only SDK gets a partial skill.

  3. Skill anatomy. The layout and the mapping file, and a cross-tool format: portable frontmatter, script paths relative to the skill directory, no reliance on $ARGUMENTS, and one source installed for both .claude/skills/ and .agents/skills/. Also outlines for SKILL.md, the module notes and a generated test file.

  4. Workflow. Selection, then a harness preflight that stops on red, then read, generate, compile, run and audit for each spec.

  5. Translation rules: traceability, fidelity, structural variants, teardown, integration hygiene, time and waits, assertions, internal access.

  6. Pseudocode construct catalogue. Every construct the corpus uses, including undocumented ones, with a column to fill in for your language.

  7. Evaluation and deviations. The three acceptable end states, harness stand-ins, deviations.md, inputs a language can't express (including capability-inapplicable tests), and stopping rules.

  8. Deterministic tooling. Contracts for the resolver, the audit and the corpus scanner.

  9. Keeping in sync with spec changes. Recording the spec SHA, a re-sync mode, renamed or merged Test IDs, and upgrading an existing skill.

  10. Verification and CI.

  11. Final report format.

  12. Lessons learned.

  13. Checklist for a new skill and harness: 50 items.

  • Appendix. The existing skills, patterns to avoid, and the rules for reading them as a last-resort reference.

Normative change for reviewers. §3.3 says name and description MUST follow the portable format. As a result, uts-to-swift and uts-to-kotlin don't conform on two or three items: the description text, script paths relative to the skill directory, and an .agents/skills link.

The skill

uts-to-lang-skill-creator/
├── SKILL.md            Step 0 (Orient), ground rules and stop points, run modes, phases at a glance, scripts
├── references/         Orient (section 13), phases 1–6, Upgrade/Fix (diff-driven; full gap audit),
│                       LiveObjects support, glossary, failure modes
├── assets/
│   ├── eligibility.json        the whitelist of Ably Pub/Sub SDK repositories
│   ├── capability-names.json   SDK client aliases, door factories, earlier-revision LiveObjects names
│   └── templates/              repo profile, harness design, design record, skill-gap audit,
│                               acceptance checklist, final report
└── scripts/            10 read-only Python 3.8+ scripts (no network)

Step 0, Orient. Every run starts here. It is read-only, and orient.py runs the other scripts and prints one State summary:

  • Spec clone. spec_clone_info.py resolves and pins the local ably/specification clone. A bad explicit path is an error, never replaced by another clone.
  • Eligibility, against a whitelist. check_repo_eligibility.py accepts a repository only if it is ably/<name> with <name> on the whitelist in assets/eligibility.json. The list holds current and legacy names (for example ably-pubsub-js and ably-js) and planned rename targets. Every other repository is rejected with a fixed message. Adding a new SDK repository means adding it to that file. A fork-only clone, or one with no GitHub remote, is asked about. A secondary gate checks that the checkout defines a REST or Realtime client.
  • Capabilities and scope. detect_capabilities.py finds the REST and Realtime clients, and for split SDKs the server and device doors. The scope follows them, so a REST-only or realtime-only SDK gets a partial skill.
  • Existing skills. inspect_existing_skill.py finds any uts-to-* skill, where it came from, its harness and its UTS-tagged tests. The repo is then classified as S0–S4.
  • LiveObjects. detect_liveobjects.py gives a one-line verdict, used later.

The detectors read client, channel, connection, presence and LiveObjects names from the spec clone's IDL at run time (spec_names.py). Only names the IDL can't give are kept in capability-names.json.

Run modes.

  • Create (S0, S1), in seven phases:
    1. understand the repo;
    2. design the harness;
    3. build and verify it;
    4. record the design decisions;
    5. generate the skill files;
    6. validate with pilot translations;
    7. write the final report.
  • Upgrade/Fix (S2, S3), in one of two sub-modes:
    • diff-driven: for a skill this procedure built, follow what changed since its recorded run;
    • full gap audit: for any other skill, audit it against every checklist item, then add, skip or defer per item.

LiveObjects decision. STOP-14 always asks whether to add uts/objects support, and the agent recommends an answer from detect_liveobjects.py's evidence. The options are:

  • full support (translate-only runs by default while the SDK's implementation is incomplete);
  • a placeholder, until the public API exists;
  • no objects support.

A full objects tier whose smoke test fails only because the SDK doesn't implement the feature yet is recorded as SDK-blocked, not as a harness failure.

The numbers.

  • 10 scripts.
  • 17 stop points (STOP-1 to STOP-17), where the agent asks the user.
  • 31 recorded decisions (D-01 to D-31).
  • An acceptance checklist of 50 items, mapped one-to-one to guide §13.

The skill links to the guide for each requirement instead of restating it. The existing skills may be read only as a last resort, under the guide's rules: read-only, recorded, and never copied.

Compaction. After a context compaction, Claude Code keeps only the start of an invoked skill (about 5,000 tokens), and Codex may keep none of it. So SKILL.md puts the ground rules and the stop table first, tells the agent to re-read the whole file after a compaction or on resuming, and scripts/check_layout.py checks that both stay inside that window.

Install (also in uts/README.md)

SPEC=~/src/specification            # your local clone of ably/specification
mkdir -p ~/.claude/skills ~/.agents/skills
ln -sfn "$SPEC/uts/skills/uts-to-lang-skill-creator" ~/.claude/skills/uts-to-lang-skill-creator   # Claude Code
ln -sfn "$SPEC/uts/skills/uts-to-lang-skill-creator" ~/.agents/skills/uts-to-lang-skill-creator   # Codex

Then, in the SDK repo, run /uts-to-lang-skill-creator <spec-clone-path> (Claude Code) or $uts-to-lang-skill-creator <spec-clone-path> (Codex). You can also ask for a uts-to-<lang> skill in your own words. Without installing anything, any agent can be told: "Read <spec-clone>/uts/skills/uts-to-lang-skill-creator/SKILL.md and follow it."

Where the guide diverges from existing UTS docs

Each case is labelled in the guide as a divergence, with its reason:

  • AWAIT_STATE: subscribe, then check, so a transition between the two isn't missed.
  • Draining before negative assertions: done explicitly, because work chained across queues needs more than one yield.
  • Features specs: read from the local spec clone at the recorded SHA, not from GitHub main.
  • Missing SDK APIs: recorded as an uncompilable stub.

Follow-ups (not in this PR)

How this was verified

Checked against the branch head:

  • Links: every relative link and #anchor in the changed Markdown files resolves.

  • IDs: STOP, D, P and G IDs and the S0–S4 classes are contiguous.

  • Checklist: guide §13 and the acceptance checklist have 50 items each, in the same 7 groups and order.

  • Frontmatter: the description is at most 1,024 characters, with no < or >.

  • Scripts: all 10 compile and parse as Python 3.8.

    • Both JSON data files are valid, and a malformed one gives a clear error naming the file and key.
    • The eligibility check was tested with legacy, current, SSH and trailing-slash remotes; with non-whitelisted repos; and with fork-only, no-remote and non-GitHub clones.
    • A bad --spec-clone is reported, not silently replaced.
    • A PHP repo with // UTS: tags is classed S1.
    • Orient run read-only on an existing SDK checkout classes it S1 and shows the clone actually used.
  • Layout: check_layout.py passes: the stop table and the ground rules end within the compaction window.

  • Lint: editorconfig-checker and git diff --check are clean on the changed files.

  • Spec counts in uts/README.md:

    Tier rest realtime objects
    unit 41 54 15
    integration (direct) 11 13 3
    integration (proxy) 1 7 1

    Plus 4 helper specs, for a total of 150.

Review guide

  • The guide: read the intro, §1 "Model tier" and §2 (especially 2.7). Then read §3.3 and skim §6.
  • The skill: read SKILL.md (Step 0, ground rules, stop points, run modes), then references/orient.md, then one phase end to end.
  • Scope: no spec or test changes, only docs, uts/README.md and the skill. The skill's scripts are read-only helpers.

sacOO7 added 2 commits October 8, 2026 12:37
Adds a guide for SDK engineers on what a per-language Claude Code
skill that translates UTS specs into native tests must contain: harness
prerequisites, skill anatomy, the two-phase workflow, translation
rules, a pseudocode construct catalogue, evaluation and deviations,
resolver and audit tooling, re-sync with spec changes, verification,
report format, lessons learned and a checklist.
Adds a step-by-step procedure an LLM agent follows to create a
uts-to-<lang> skill for an SDK repo: study the repo's existing code,
tests and test support; design, build and verify the UTS harness; then
generate the skill and validate it with pilot translations. Links both
docs from uts/README.md.
sacOO7 added 2 commits October 8, 2026 14:01
Renames generating-a-translator-skill.md to uts-to-lang-skill-creator.md,
matching the naming of skill-creator skills, and updates its title and
the links to it from the guide and uts/README.md.
Treats the harness (UTS test infrastructure) and the skill as one
deliverable: the harness is designed from the repo's existing test
support, and permanent, CI-run per-tier smoke tests and helper
self-tests prove it before the skill is generated. The skill gets a
Harness reference section, the mapping and resolver report the harness,
and a preflight step runs the tier's harness tests before translating.

Also:
- allows reading the existing Swift and Kotlin skills and harnesses,
  pinned to reviewed commits, as a guarded last resort when the guide
  and UTS docs don't answer a question; never copied, always recorded
- recommends Opus-class models: required for creating the skill and
  harness, recommended for translation runs, and recorded in reports
- tightens procedure ordering, collected-test counting, known-gaps
  matching and other points found in review, and corrects claims
  about the existing skills

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The audit, re-sync, source-pinning, scanner, and CI-readiness instructions contain correctness gaps.

5 open findings
What changed in this PR

Adds comprehensive guidance for creating language-specific UTS translator skills and their test harnesses.

Changes:

  • Defines translator-skill architecture, workflow, auditing, and maintenance.
  • Adds an agent procedure for building and validating skills and harnesses.
  • Links both guides from the UTS README.
File Description
uts/​README.md Adds translator-skill documentation links.
uts/​docs/​translator-skills/​writing-translator-skills.md Defines translator-skill and harness requirements.
uts/​docs/​translator-skills/​uts-to-lang-skill-creator.md Provides the step-by-step creation procedure.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment thread uts/docs/translator-skills/uts-to-lang-skill-creator.md Outdated
Comment thread uts/docs/translator-skills/uts-to-lang-skill-creator.md Outdated
Comment thread uts/docs/translator-skills/writing-translator-skills.md Outdated
Comment thread uts/docs/writing-uts-spec-translator-skills.md
Comment thread uts/docs/translator-skills/writing-translator-skills.md Outdated
@sacOO7

sacOO7 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Multiple moderate documentation and procedure issues remain unresolved.

6 open findings

🧠 Review effort: Lite


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment thread uts/docs/translator-skills/writing-translator-skills.md Outdated
- count expected-failure forms (FAILS WITH, THROWS, EXPECT THROW,
  AWAIT_ERROR) as assertions in the audit contract, so surplus waits
  can't mask a dropped assertion
- report two-letter keywords (IS, IF, IN, OR, AS) in the corpus scanner
- treat any change to a spec file as changed during re-sync, and diff
  against the working tree, including untracked specs
- stop and ask before translating a locally modified spec, and record
  its blob hash when the user proceeds
- state that a tier whose harness tests aren't wired into CI keeps the
  skill non-conforming until CI runs them
Replaces the single uts-to-lang-skill-creator.md procedure doc with an
agent skill at uts/skills/uts-to-lang-skill-creator/ that loads in both
Claude Code and Codex: a short SKILL.md (portable frontmatter, ground
rules and stop points first), per-phase references, record templates,
and runnable scripts (corpus scanner, spec-clone locator, repo survey).

Also:
- the guide recommends a skill format that works in both tools:
  portable frontmatter, script paths relative to the skill directory,
  and install paths for Claude Code and Codex
- correct the guide's description of __PASSTHROUGH__, which uts-proxy
  passes through unchanged
- link the translator skill guide from writing-derived-tests.md
- update the uts/README.md tree and spec counts (adding objects, the
  REST proxy tier and skills), and document installing the skill
The guide is now the only file left in uts/docs/translator-skills/, so
move it next to the other writing-* guides it builds on
(writing-test-specs.md and writing-derived-tests.md) as
uts/docs/writing-uts-spec-translator-skills.md, and retitle it
"Writing UTS Spec Translator Skills".

Update every relative link to and from the guide (the skill, the README
and writing-derived-tests.md), the plain-text paths in the skill
(SKILL.md, the maintenance diff, the checklist template and
spec_clone_info.py), and the uts/README.md tree and Guides entry.
No other content changes.
Remove notes about when and against which commits the guide was
written (dates, reviewed SHAs, "as of" stamps, commit ids in examples)
from the guide and the uts-to-lang-skill-creator skill. They read as
noise and go stale.

The reference-implementation links now point at main. The last-resort
rules no longer refer to pinned commits: whatever is taken from the
existing skills is checked against Patterns to avoid, which describes
them when the guide was written. Local clones of those repos are read
at their current HEAD. No requirements, IDs or checklist items change.
The spec counts are a snapshot; point readers at the find command for
current numbers instead of stamping a date.
The skill creator now starts every run with a read-only Orient step:
- check the repo is an Ably Pub/Sub SDK (owner ably; ably-<lang> or
  ably-pubsub-<lang>, with deny and allow lists in assets/eligibility.json)
- detect REST and Realtime capabilities, door entry points and
  LiveObjects support, using class names read from the spec clone at
  run time, plus SDK-specific aliases in assets/capability-names.json
- classify the repo state and always ask for the mode: create a new
  skill, or upgrade/fix an existing uts-to-* skill (diff-driven, full
  gap audit, or regenerate)

Scope follows the detected capabilities: REST-only and realtime-only
SDKs get a partial skill that a later upgrade can extend, and tests
that need the absent client are marked capability-inapplicable.

Also adds the existing-skill gap audit with per-item add/skip/defer,
an always-asked LiveObjects decision with an SDK-blocked translate-only
path, and the matching guide and template updates.
sacOO7 added 11 commits October 10, 2026 16:13
Shorten the capability-names.json description and comments, describe the
_description and _comment conventions once in the guide and section 10,
and refer to a reviewer rather than a lead reviewer.
Replace the name pattern, deny-list and allow-list in eligibility.json with
an explicit whitelist of Ably Pub/Sub SDK repositories, current and legacy
names, plus the planned rename targets. Any other repository is rejected
with a message saying that supporting it requires updating the skill
creator. The data file is checked on load: every name it maps must be on
the whitelist. Fork, no-remote and non-GitHub handling and the definition
gate are unchanged. The docs describe the whitelist instead of the name
rule.
…ipts

An explicit --spec-clone (or UTS_SPEC_CLONE) that isn't in a spec clone
was silently replaced by another clone. spec_names.py now walks a path up
to the clone root and reports NOT_A_SPEC_CLONE instead of falling back,
and orient.py resolves the clone with spec_clone_info.py first and passes
that path to every script, so the State summary shows the clone used.

inspect_existing_skill.py had its own source-extension list without .php,
so UTS tags in PHP tests were missed and the repo was classed S0. It now
uses the list from detect_liveobjects.py, like detect_capabilities.py.
…n.md

Section 10 is the diff-driven sub-case of Upgrade/Fix, not a separate
mode: retitle it and update every link. Ship the unreleased skill as
version 1.0.0.
…audit

Section 11 is the other sub-case of Upgrade/Fix; its title said
"Upgrade mode", a name also used for all of Upgrade/Fix.
- Every script handles --help (docstring to stdout, exit 0) and reports
  usage errors on stderr with exit 2, including missing flag values,
  unknown options and stray arguments; usage strings match the code.
- spec_clone_info.py accepts --spec-clone and reports a warnings list;
  spec_names.py reports ok: true.
- orient.py reports a failing child script's exit code and stderr.
- survey_repo.py no longer parses git error text as data and runs git
  with --no-optional-locks.
- The list of agent skill directories is defined once; dead names are
  removed; repo paths expand ~ before they are checked.
- Module docstrings are cut to purpose, usage, output and exit codes,
  pointing to the reference sections for the rules.
- capability-names.json has a one-line description.
Use Upgrade/Fix consistently (diff-driven or full gap audit), drop the
hard-coded corpus test lists, the pre-merge legacy clauses and stale
ID ranges, fix the Done-when placement, keep one STOP-17 trigger list
in section 13.8, add capabilities and scope to the Orient summaries,
and add glossary entries for terms used without definition.
Move Run modes and Phases at a glance after Step 0, link to the guide's
reference-implementation rules and model tier instead of restating
them, trim the Scripts table and Reference index, and move the
data-file upkeep note to a Maintaining this skill section. The
description names the whitelist.
Remove repeated statements of the same rule in favour of links, the
(new) markers, most inline bold and an em-dash pair; align requirement
casing with the section 13 checklist; make the section 2.7 tier
subsections headings; stop pinning the uts-proxy name limit to one
version as an absolute claim.
- orient.py passes an explicit spec clone that isn't a git checkout to
  every script instead of dropping it, and its State summary names it.
- Document the stop order when the spec clone path is bad, the second
  STOP-16 reject message, which scripts Orient runs, the older record
  forms the inspector still reads, and copying the checklist column in
  the gap audit.
- Fix a glossary naming claim and a dangling pronoun in the guide.
@sacOO7 sacOO7 changed the title UTS: add guide and agent procedure for building uts-to-<lang> translator skills UTS: add translator-skill guide and uts-to-lang-skill-creator skill Oct 10, 2026
… compaction window

After a context compaction, Claude Code re-attaches only the start of an
invoked skill (about 20,000 characters of the text after the frontmatter),
and the stop table and ground rules ended past it.

- Add a rule under the title: after a compaction, re-read the whole
  SKILL.md before the next action; mirror it in the design-record
  template.
- Condense Step 0; its detail stays in references/orient.md, which now
  also holds the Phase 1 record-copying step.
- Move Run modes and Phases at a glance after section 2, and the
  reference-implementation rules from 2.1 to their own section, leaving
  a pointer. No anchor changes.
- Add scripts/check_layout.py, which checks that the stop table and
  section 2 end within 16,000 and 18,500 characters, and document it in
  Maintaining this skill.

This branch was successfully deployed

1 active deployment
staging/pull/558 — 9d67e217 Deployed Oct 10, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants