A test written after a refactor describes what the refactor happened to do rather than what it had to do. Confirmed exit 1, naming all six skills over the ceiling with no references/ dir, before the first split ran.
A SKILL.md body has no partial load: the moment a skill triggers, the whole body is paid. Six of them — three that rewrite refs and end in a squash merge — went from 9,565 words of always-loaded context to 3,536, with every safety guard held in the core.
ship is the short one at 571 — review forced it lower; see Review follow-ups.
Implemented from docs/plans/split-git-social-skills.md, fanned out across six parallel subagents. All four tests pass, and 15 of the plan's 16 acceptance criteria are met — the sixteenth is behavioural verification, which has not run. The plan's behavioural verification is not run, so status: stayed in-progress.
The unit of optimisation is the core, not the plugin directory. Content merely relocated under kit/plugins/ is not “moved out of context”; content moved behind a link the model may never open is worse than either.
Never merge on anything but green, no-verify, Cap autofix at 3 attempts per failing check. Procedure is the commands and tables that implement it: the gh api graphql review-thread query, the CI failure classification table, the branch-name type-inference table, the stash-pop recovery script.
social-media-tools/skills/share-react/references/props-extraction.md — skill-local, linked as Read `references/x.md` (bundled with this skill).git-agent/skills/create-issue/references/ — the same shape on the git-agent side.Both are distinct from plugin-level social-media-tools/references/ — eight files that eleven skills read. Those were left byte-untouched; every new file is skill-local, and the objective test freezes the plugin-level link counts at 7 / 8 / 11 so a future edit cannot quietly rewire shared infrastructure.
Inline guard prose bottomed out near 700 words — 22 negative imperatives plus 12 step headings do not fit alongside per-step text. The core now leads with a ## Guardrails block holding every hard stop, followed by all step headings (0, 1, 2, 2.5, 3, 4, 5, 6, 6a–6d, 7, 8) in unchanged order, each reduced to a pointer. Read that file first — it is the template the other five follow.
tests/plugins/test-skill-split-git-social.sh is the objective gate, modelled on test-remaining-skill-splits.sh from PR #487. Six checks: word ceiling, reference placement, link integrity in both directions, descriptions pinned to literal strings, guard retention per owning core, and the frozen plugin-level counts.
test-ship-self-review.sh was retargeted, not relaxed. Its 22 checks span ship/SKILL.md and agents/agent-ship.md; Step 4.5's content moved, so checks 5–7 now read ship/references/self-review.md while the policy checks (2, 3, 4, 8, 9, 10) stay on the core. A new check 4.5 asserts the core actually links the reference — without it, the retarget would pass against an orphan.
A test written after a refactor describes what the refactor happened to do rather than what it had to do. Confirmed exit 1, naming all six skills over the ceiling with no references/ dir, before the first split ran.
wc -w.Borrowed from test-remaining-skill-splits.sh. These bodies are full of em dashes, → and ≤; in the C locale a standalone — is not a word, in C.UTF-8 it is. That is a ~20-word swing per file — the difference between passing on a dev machine and failing on a CI runner, which is the exact drift the test exists to prevent.
The description: is the sole trigger surface. Literal pinning works with no git available and prints a want-vs-got diff on failure. git diff main is kept as a separate acceptance check rather than as the mechanism.
The plan pins 2,448 / 1,863 / 1,515 / 1,414 / 1,284 / 1,234, but commits 745584e and ce69bc8 had already trimmed these same bodies to 2,406 / 1,840 / 1,476 / 1,391 / 1,261 / 1,191. Step 1's verify was unmeetable as written; the ceiling is the real invariant.
The splits are independent — disjoint file sets, no shared state. Each agent received the guard phrases it had to retain verbatim, the exact reference filenames, and its own verification commands. Main-loop context stayed at 41 tool calls because no 1,200–2,400-word skill body was ever read into it.
status: in-progress with 15 of 16 criteria met.The plan's behavioural verification requires opening a real PR, pushing to the remote, and an irreversible merge. Marking completed would assert verification that did not happen.
share-explanation needed 5 against the plan's 3; share-session needed 3 against 2. The fixed floor of a core — frontmatter, phase-index table, plan-mode guard, scrub gate, ~12 phase headings — is already ~450 words, so the 600-word ceiling and the plan's own Files list were mutually unsatisfiable. Chose the ceiling, because that is what CI enforces.
To revisit: raise the ceiling to ~750 and the spec's file list becomes achievable.
ship left at 599 of 600 — superseded, now 571Trimming for headroom was considered and rejected for now: that file carries 22 test assertions and every edit risks one of them. Shipping verified-green beat shipping tidier-but-re-verified.
To revisit: trim Step 3's rules list or the duplicated closing STOP block, then re-run test-ship-self-review.sh.
opus-5Six agents at 100k–176k tokens each. The mechanical splits would plausibly have run on a cheaper tier; the guard-retention requirement argued against it for ship-autonomous, and uniformity won over per-agent tuning.
To revisit: pass model per agent and keep the top tier only for the guard-heavy targets.
TEMPLATES_DIR dedup left out of scopeThe same three-line find/CLAUDE_PLUGIN_ROOT bootstrap appears verbatim in eleven social-media-tools skills, ~70 words each. It touches eight skills this work does not otherwise open; the spec routes it to Next Steps and that held.
git checkout -- on uncommitted work restores the pre-change file, not the pre-mutation one.This cost real work. The tautology checks mutate a file, assert the test fails, then restore. Restoring with git checkout -- <path> reverted ship/SKILL.md and share-session/SKILL.md to their pre-split 1,191- and 1,391-word state — the finished cores had never been committed, so HEAD held no copy. Both had to be rebuilt from their surviving reference files.
git checkout -- <tracked> <untracked> errors on the unmatched pathspec and restores neither file. That is why ship-autonomous came back at 583 words rather than 595: it had kept the tautology mutation (one guard line deleted), and self-review.md had kept its deleted Responsive bullet. The failure is quiet — the error names the pathspec, not the fact that the whole operation was abandoned.
Working method: cp to the scratchpad, mutate, test, cp back. All four checks ran clean on the second pass.
Removing lines matching Responsive from a wrapped Markdown bullet left breakpoint, srcset, width… dangling under the previous item. Same shape in the ship-autonomous guardrails, where with --match-head-commit. orphaned under the review bullet. Neither looked like corruption.
session_usage.py reports 161k for this session; the eight subagents reported roughly 1.05M more, in separate transcripts. Fan-out moved cost off the main-loop accounting rather than reducing it — worth knowing before reading any session-usage comparison as a cost measurement.
ship's fourth pre-flight hard stop — CLI not available or not authenticated … STOP — out of the core and into references/platform-clis.md: the precise failure mode this work exists to prevent. The test could not have caught it, because its guard list had no phrase covering CLI auth. Both are fixed; guard assertions went 11 → 12 and are now tallied rather than hard-coded.
| Gate | Result |
|---|---|
| test-skill-split-git-social.sh (new) | exit 0 6 checks |
| test-ship-self-review.sh (retargeted) | exit 0 all 22 checks |
| test-description-budget.sh | exit 0 |
| test-no-orphan-plugin-dirs.sh | exit 0 |
| BASE_REF=main node scripts/check-plugin-versions.mjs | exit 0 |
| Acceptance criteria | 15 / 16 |
| Behavioural check (plan §Verification) | not run |
Run separately so a single lenient grep cannot hide behind a passing sibling:
ship/SKILL.md → exit 1, FAIL: ship(999).Never merge on anything but green line into references/merge-gate.md → exit 1, lost guard. This is the realistic failure mode, and the one a presence-anywhere check would miss.share-session's description: → exit 1 with a want/got diff.Responsive bullet from ship/references/self-review.md → test-ship-self-review.sh exit 1.Clean exit 0 after each restore.
No skill was executed. Word counts prove nothing about whether these skills still branch, ship, and publish cards. The plan's behavioural check needs claude --plugin-dir sessions, a real branch and PR, ship-autonomous against a deliberately failing lint script, and three Playwright PNG renders. It is the single outstanding gate and the reason the plan is not marked complete. A structural read-through of ship-autonomous/SKILL.md stood in as a partial substitute — guards up front, step order intact, every pointer landing at the right reference — which is weaker evidence than a run.
share-session 597, ship-autonomous 597, share-explanation 596, share-selection 593 — against a ceiling of 600. (ship escaped to 571 only because review forced a guard back into it.) This is not a ceiling with an upgrade path, it is a cliff: routine wording edits will break the build. Either budget a real margin per core or raise the ceiling.
Blocks marking the plan complete. Needs a scratch branch and permission to open and close a PR.
The eleven-skill TEMPLATES_DIR dedup, and git-agent/skills/merge/SKILL.md — 1,156 words, no references, ends in an irreversible squash merge. Both carry self-contained paste-ready prompts in the plan.
## Guardrails block holds all 22 negative imperatives.--no-track, and the no-retry/no-force rule.no-verify prohibition, the CLI-auth hard stop restored during review, and Step 4.5's four policy lines.gh/glab detection and auth messages; the four regression checks and amend procedure; the PR/MR body template.plugin.json gained a version field.status: in-progress, modified: 2026-07-30; HTML re-rendered by scripts/build-plan-html.mjs.