Engineering recap · agentics · 2026-07-30

Six monolithic skill bodies split into cores plus references

A SKILL.md body has no partial load: the moment a skill triggers, the whole body is paid. Six of them — three that rewrite refs and end in a squash merge — went from 9,565 words of always-loaded context to 3,536, with every safety guard held in the core.

commit e5fccc7 33 files git-agent 4.8.0 social-media-tools 2.20.0 spec: docs/plans/split-git-social-skills.md

At a glance

63%context cut
20references added
15/16criteria met
4/4tests passed
3open items
skill corebeforeafterproportion
ship-autonomous2,406597
share-explanation1,840596
branch-agent1,476582
share-session1,391597
share-selection1,261593
ship1,191571
9,565 → 3,536 words. The “after” bars are near-identical because four of six cores land within 7 words of the 600-word gate. ship is the short one at 571 — review forced it lower; see Review follow-ups.

Implemented from docs/plans/split-git-social-skills.md, fanned out across six parallel subagents. All four tests pass, and 15 of the plan's 16 acceptance criteria are met — the sixteenth is behavioural verification, which has not run. The plan's behavioural verification is not run, so status: stayed in-progress.

Architecture and code paths

The unit of optimisation is the core, not the plugin directory. Content merely relocated under kit/plugins/ is not “moved out of context”; content moved behind a link the model may never open is worse than either.

The load-bearing rule: guards stay, procedure moves. A guard is the statement — Never merge on anything but green, no-verify, Cap autofix at 3 attempts per failing check. Procedure is the commands and tables that implement it: the gh api graphql review-thread query, the CI failure classification table, the branch-name type-inference table, the stash-pop recovery script.

model opens on demand

skill triggers

SKILL.md core~595 wordsALWAYS paid

Guardrailsevery hard stop

step headingsorder preserved

one-line pointer

references/*.mdcommands, tables, queries

decides whetheranything happens at all

Look at the dashed edge: it is the only thing that is not paid on every trigger. Everything a guard decides sits to its left.

Layout follows two existing conventions

Both are distinct from plugin-level social-media-tools/references/ — eight files that eleven skills read. Those were left byte-untouched; every new file is skill-local, and the objective test freezes the plugin-level link counts at 7 / 8 / 11 so a future edit cannot quietly rewire shared infrastructure.

ship-autonomous needed a structural change

Inline guard prose bottomed out near 700 words — 22 negative imperatives plus 12 step headings do not fit alongside per-step text. The core now leads with a ## Guardrails block holding every hard stop, followed by all step headings (0, 1, 2, 2.5, 3, 4, 5, 6, 6a–6d, 7, 8) in unchanged order, each reduced to a pointer. Read that file first — it is the template the other five follow.

Test topology

tests/plugins/test-skill-split-git-social.sh is the objective gate, modelled on test-remaining-skill-splits.sh from PR #487. Six checks: word ceiling, reference placement, link integrity in both directions, descriptions pinned to literal strings, guard retention per owning core, and the frozen plugin-level counts.

test-ship-self-review.sh was retargeted, not relaxed. Its 22 checks span ship/SKILL.md and agents/agent-ship.md; Step 4.5's content moved, so checks 5–7 now read ship/references/self-review.md while the policy checks (2, 3, 4, 8, 9, 10) stay on the core. A new check 4.5 asserts the core actually links the reference — without it, the retarget would pass against an orphan.

Decisions

Write the objective test before touching any skill.

A test written after a refactor describes what the refactor happened to do rather than what it had to do. Confirmed exit 1, naming all six skills over the ceiling with no references/ dir, before the first split ran.

Count words in Python, not wc -w.

Borrowed from test-remaining-skill-splits.sh. These bodies are full of em dashes, → and ≤; in the C locale a standalone — is not a word, in C.UTF-8 it is. That is a ~20-word swing per file — the difference between passing on a dev machine and failing on a CI runner, which is the exact drift the test exists to prevent.

Pin descriptions to literal strings in the test, not to a git diff.

The description: is the sole trigger surface. Literal pinning works with no git available and prints a want-vs-got diff on failure. git diff main is kept as a separate acceptance check rather than as the mechanism.

Assert the ceiling, not a delta from the spec's baselines.

The plan pins 2,448 / 1,863 / 1,515 / 1,414 / 1,284 / 1,234, but commits 745584e and ce69bc8 had already trimmed these same bodies to 2,406 / 1,840 / 1,476 / 1,391 / 1,261 / 1,191. Step 1's verify was unmeetable as written; the ceiling is the real invariant.

Fan out one subagent per skill, six concurrent.

The splits are independent — disjoint file sets, no shared state. Each agent received the guard phrases it had to retain verbatim, the exact reference filenames, and its own verification commands. Main-loop context stayed at 41 tool calls because no 1,200–2,400-word skill body was ever read into it.

Leave status: in-progress with 15 of 16 criteria met.

The plan's behavioural verification requires opening a real PR, pushing to the remote, and an irreversible merge. Marking completed would assert verification that did not happen.

Tradeoffs and rejected options

More reference files than the spec budgeted

share-explanation needed 5 against the plan's 3; share-session needed 3 against 2. The fixed floor of a core — frontmatter, phase-index table, plan-mode guard, scrub gate, ~12 phase headings — is already ~450 words, so the 600-word ceiling and the plan's own Files list were mutually unsatisfiable. Chose the ceiling, because that is what CI enforces.

To revisit: raise the ceiling to ~750 and the spec's file list becomes achievable.

ship left at 599 of 600 — superseded, now 571

Trimming for headroom was considered and rejected for now: that file carries 22 test assertions and every edit risks one of them. Shipping verified-green beat shipping tidier-but-re-verified.

To revisit: trim Step 3's rules list or the duplicated closing STOP block, then re-run test-ship-self-review.sh.

Every subagent inherited opus-5

Six agents at 100k–176k tokens each. The mechanical splits would plausibly have run on a cheaper tier; the guard-retention requirement argued against it for ship-autonomous, and uniformity won over per-agent tuning.

To revisit: pass model per agent and keep the top tier only for the guard-heavy targets.

TEMPLATES_DIR dedup left out of scope

The same three-line find/CLAUDE_PLUGIN_ROOT bootstrap appears verbatim in eleven social-media-tools skills, ~70 words each. It touches eight skills this work does not otherwise open; the spec routes it to Next Steps and that held.

Learnings

git checkout -- on uncommitted work restores the pre-change file, not the pre-mutation one.

This cost real work. The tautology checks mutate a file, assert the test fails, then restore. Restoring with git checkout -- <path> reverted ship/SKILL.md and share-session/SKILL.md to their pre-split 1,191- and 1,391-word state — the finished cores had never been committed, so HEAD held no copy. Both had to be rebuilt from their surviving reference files.

Two of the four restores silently did nothing.

git checkout -- <tracked> <untracked> errors on the unmatched pathspec and restores neither file. That is why ship-autonomous came back at 583 words rather than 595: it had kept the tautology mutation (one guard line deleted), and self-review.md had kept its deleted Responsive bullet. The failure is quiet — the error names the pathspec, not the fact that the whole operation was abandoned.

Working method: cp to the scratchpad, mutate, test, cp back. All four checks ran clean on the second pass.

A line-filter delete leaves orphaned continuation lines that read as valid prose.

Removing lines matching Responsive from a wrapped Markdown bullet left breakpoint, srcset, width… dangling under the previous item. Same shape in the ship-autonomous guardrails, where with --match-head-commit. orphaned under the review bullet. Neither looked like corruption.

Subagent token cost does not appear in the session's own ledger.

session_usage.py reports 161k for this session; the eight subagents reported roughly 1.05M more, in separate transcripts. Fan-out moved cost off the main-loop accounting rather than reducing it — worth knowing before reading any session-usage comparison as a cost measurement.

Tests and verification

Corrected after review. This section originally reported a clean pass. Code review then found that the split had moved ship's fourth pre-flight hard stop — CLI not available or not authenticated … STOP — out of the core and into references/platform-clis.md: the precise failure mode this work exists to prevent. The test could not have caught it, because its guard list had no phrase covering CLI auth. Both are fixed; guard assertions went 11 → 12 and are now tallied rather than hard-coded.
GateResult
test-skill-split-git-social.sh (new)exit 0 6 checks
test-ship-self-review.sh (retargeted)exit 0 all 22 checks
test-description-budget.shexit 0
test-no-orphan-plugin-dirs.shexit 0
BASE_REF=main node scripts/check-plugin-versions.mjsexit 0
Acceptance criteria15 / 16
Behavioural check (plan §Verification)not run

Tautology check — four passes, each against a different assertion

Run separately so a single lenient grep cannot hide behind a passing sibling:

  1. +400 filler words in ship/SKILL.md → exit 1, FAIL: ship(999).
  2. Moving — not deleting — the Never merge on anything but green line into references/merge-gate.md → exit 1, lost guard. This is the realistic failure mode, and the one a presence-anywhere check would miss.
  3. One character in share-session's description: → exit 1 with a want/got diff.
  4. Deleting the Responsive bullet from ship/references/self-review.md → test-ship-self-review.sh exit 1.

Clean exit 0 after each restore.

Knowingly untested

No skill was executed. Word counts prove nothing about whether these skills still branch, ship, and publish cards. The plan's behavioural check needs claude --plugin-dir sessions, a real branch and PR, ship-autonomous against a deliberately failing lint script, and three Playwright PNG renders. It is the single outstanding gate and the reason the plan is not marked complete. A structural read-through of ship-autonomous/SKILL.md stood in as a partial substitute — guards up front, step order intact, every pointer landing at the right reference — which is weaker evidence than a run.

Review follow-ups and tech debt

Four of six cores sit within 7 words of the CI gate

share-session 597, ship-autonomous 597, share-explanation 596, share-selection 593 — against a ceiling of 600. (ship escaped to 571 only because review forced a guard back into it.) This is not a ceiling with an upgrade path, it is a cliff: routine wording edits will break the build. Either budget a real margin per core or raise the ceiling.

Behavioural verification outstanding

Blocks marking the plan complete. Needs a scratch branch and permission to open and close a PR.

Next Steps from the spec, unstarted

The eleven-skill TEMPLATES_DIR dedup, and git-agent/skills/merge/SKILL.md — 1,156 words, no references, ends in an irreversible squash merge. Both carry self-contained paste-ready prompts in the plan.

Files touched

git-agent skills — cores rewritten, guards retained in-core

skills/ship-autonomous/SKILL.md
2,406 → 595w. New ## Guardrails block holds all 22 negative imperatives.
skills/ship-autonomous/references/{preflight-and-verify, pr-events, ci-autofix, merge-gate}.md
Step 1/2.5 commands; subscribe-vs-poll, triage and review handling; CI classification table; merge and branch-deletion sequences.
skills/branch-agent/SKILL.md
1,476 → 582w. Keeps the Step 1 guard trio, --no-track, and the no-retry/no-force rule.
skills/branch-agent/references/{branch-naming, stash-and-recovery}.md
Name resolution and type inference; conflict detection and stash-pop recovery.
skills/ship/SKILL.md
1,191 → 571w. Keeps Step 1 guards, the no-verify prohibition, the CLI-auth hard stop restored during review, and Step 4.5's four policy lines.
skills/ship/references/{platform-clis, self-review, pr-body, commit-message}.md
gh/glab detection and auth messages; the four regression checks and amend procedure; the PR/MR body template.

social-media-tools skills — cores rewritten, scrub gate retained in-core

skills/share-explanation/SKILL.md 1,840 → 596w
references/{target-resolution, synthesis-structure, card-population, bootstrap, copy-drafting}.md — five-tier lookup, per-target section structures, escape order and variable tables, Phase 0 bootstrap, platform copy guidance.
skills/share-session/SKILL.md 1,391 → 597w
references/{session-data, card-population, draft-copy}.md — Phase 1a–1e gathering, escape order and card variables, per-platform recap formats.
skills/share-selection/SKILL.md 1,261 → 593w
references/{selection-sources, card-population}.md — source precedence and file guards, template pick and diff/snippet variable tables.

Tests and CI

tests/plugins/test-skill-split-git-social.sh
New. The objective gate: ceiling, references, link integrity both ways, pinned descriptions, guard retention, frozen plugin-level counts.
tests/plugins/test-ship-self-review.sh
Checks 5–7 retargeted to the reference; new check 4.5 asserts the core links it.
.github/workflows/check-plugin-versions.yml
One step after “Test gallery index merge driver”, with a comment on why these six guards are checked at review time.

Metadata

.claude-plugin/marketplace.json
git-agent 4.7.1 → 4.8.0, social-media-tools 2.19.2 → 2.20.0. Neither plugin.json gained a version field.
kit/plugins/{git-agent, social-media-tools}/CHANGELOG.md
v4.8.0 and v2.20.0 entries with before/after counts and the guards-stay-in-core rationale.
docs/plans/split-git-social-skills.md · .html
Fourteen criteria ticked, status: in-progress, modified: 2026-07-30; HTML re-rendered by scripts/build-plan-html.mjs.