Remove process-reminder imperatives from the five most over-constrained SKILL.md files in plan-agent, git-agent, and skill-reviewer, but only after each skill's behavior is captured as a committed, passing baseline test that proves the pruning changed nothing that matters.
Five skills carry 118 hard imperatives across 13,758 words, and most of them restate process the model already infers. We will know this worked when the pruned skills reproduce their recorded structural manifests exactly on the same fixed inputs, and every safety guard we classified as load-bearing is still literally present in its file.
Read and implement all steps in the plan at docs/plans/remove-skill-process-imperatives.md — Earn every NEVER — baseline first, then prune. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/remove-skill-process-imperatives.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
More ways to run this plan — goal & workflow prompts, file path
Achieve this goal: Earn every NEVER — baseline first, then prune. The plan at docs/plans/remove-skill-process-imperatives.md describes one approach — use it as reference, but optimize for the outcome. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/remove-skill-process-imperatives.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
remove-skill-process-imperatives.html
docs/plans/remove-skill-process-imperatives.html
docs/plans/remove-skill-process-imperatives.md
Context
The story behind this plan — what prompted the work and why it matters now.
Anthropic's "The new rules of context engineering for Claude 5 generation models" makes Rule 1 — judgment over rules — the headline: they removed 80%+ of Claude Code's system prompt with no measurable loss. The catch is the part that does not travel: they had evals. Removing constraints with "it still looks fine" as the verification is precisely how a silent regression ships, and the regressions in this repo are not cosmetic. git-agent:ship-autonomous ends in an irreversible squash merge. plan-agent:implementation-plan is forbidden from writing source files. git-agent:branch-agent stashes and pops a user's uncommitted work. If a dropped sentence changes any of those behaviors, the failure mode is a destroyed working tree or an unreviewed merge, discovered later.
A measured count of lines matching NEVER|ALWAYS|MUST|do not|don't (case-insensitive — grep -ciE, matching lines rather than occurrences) across the five heaviest skills:
| Skill | Imperatives | Words |
|---|---|---|
| kit/plugins/plan-agent/skills/build/SKILL.md | 34 | 2,907 |
| kit/plugins/plan-agent/skills/implementation-plan/SKILL.md | 34 | 3,735 |
| kit/plugins/git-agent/skills/ship-autonomous/SKILL.md | 20 | 2,448 |
| kit/plugins/git-agent/skills/branch-agent/SKILL.md | 16 | 1,515 |
| kit/plugins/skill-reviewer/skills/optimizing-skill-frontmatter/SKILL.md | 14 | 3,153 |
That is 118 imperatives across 13,758 words, in three plugins: plan-agent (5.0.0), git-agent (4.7.0), skill-reviewer (2.2.8).
The central discriminator. Keep an imperative only if violating it fails silently and expensively. Everything else is a process reminder a Claude 5 generation model infers from the surrounding step. Concretely — KEEP the safety, scope, and irreversibility guards: ship-autonomous's "Never merge on anything but green" and "never pass that flag on the strength of a merge" (the --delete-branch guard — the source line wraps after "merge", so only that much is greppable as one line), its "Never dismiss a review on your own initiative", branch-agent's "Do not retry. Do not force." and "Do not drop the stash before staging resolved files", implementation-plan's entire ## Scope Constraint — Plans Only block including "This constraint is never lifted here", build's "Never resolve a gate by picking for the user" and its never-hand-edit-the-HTML rule, optimizing-skill-frontmatter's "Never write disable-model-invocation: false", and the blocking security-scrub gate wherever a skill publishes. DROP the process scaffolding: ship-autonomous's "Run Steps 0–5 in strict order", branch-agent's opening "Follow these steps in strict order." (its ## Step 6: Confirm and STOP block already states the same stopping rule), every Step 0 sentence explaining why a mutation needs plan mode exited (build's "This skill writes source files, so it cannot run inside plan mode." and implementation-plan's "writing plan files is a filesystem mutation that cannot proceed inside harness plan mode"), build's Step 2 aside "Spec edits only (see the source-of-truth rule above)" now that the ## Overview source-of-truth paragraph states it, implementation-plan's Step 4 Rename restatement of the kebab-case verb-target convention already given in Step 2's write instruction, and optimizing-skill-frontmatter's "Follow these steps exactly" and "Count carefully".
Sequencing is the safety mechanism, not a nicety. Baselines are captured, committed, and green before a single imperative is removed. A baseline recorded after the fact is a memory, not a control. This is why workflow: false — parallelizing these steps would let a pruning step land before its own baseline exists, which destroys the only evidence the refactor is safe.
Risks, with mitigations. (1) Behavioral baselines need the claude CLI, which GitHub Actions does not have — mitigated by splitting the harness: test-skill-behavior-baselines.sh is a local-only gate that exits 1 rather than skipping when the CLI is missing, and test-imperative-pruning.sh is the CI-wired structural gate. (2) Model nondeterminism could make a baseline flaky and get it disabled — mitigated by asserting only structural facts (a file exists at path X, status: is in-progress, a refusal was emitted, no file was written outside docs/plans/), never prose wording. (3) A prune could silently change a frontmatter description, breaking activation without breaking any test — mitigated by diffing each description line against git show origin/main:<path>.
Abort condition, stated up front. If a pruned skill's baseline test fails and the cause is not obvious within one fix attempt, restore the imperative rather than debugging further. The imperative cost a few tokens; the regression costs a merge. Restoring is a success outcome for this plan, not a failure.
Honest accounting. This is the smallest token return of the context-engineering set — roughly five skills, and the pruned text is a few hundred words per file at most — and by a wide margin the highest regression risk. Dropping this plan entirely is a legitimate outcome. If Step 2 shows the baseline harness costs more to build and maintain than the saving is worth, stop there and record that finding; do not proceed to pruning without the harness.
Files that change
Every file this plan touches, and what happens to each one.
docs/plans/remove-skill-process-imperatives.mdnew this spectests/fixtures/imperative-baselines/keep-phrases.txtnew the committed KEEP classification: one literal guard phrase per line, prefixed with the skill path it must remain intests/fixtures/imperative-baselines/scenarios/new fixed inputs per skill: a known plan spec, a known dirty tree script, a known SKILL.mdtests/fixtures/imperative-baselines/*.expectednew recorded structural manifests, one per target skill- tests/plugins/
test-skill-behavior-baselines.shnew local-only behavioral harness; runs each skill headless against its fixed scenario and diffs against the recorded manifesttest-imperative-pruning.shnew objective test; CI-wired structural gate
kit/plugins/plan-agent/skills/build/SKILL.mdmodified prune process reminders, keep gate guardskit/plugins/plan-agent/skills/implementation-plan/SKILL.mdmodified prune process reminders, keep## Scope Constraint — Plans Onlyintactkit/plugins/git-agent/skills/ship-autonomous/SKILL.mdmodified prune process reminders, keep merge and branch-deletion guardskit/plugins/git-agent/skills/branch-agent/SKILL.mdmodified prune duplicate ordering text, keep stash and no-force guardskit/plugins/skill-reviewer/skills/optimizing-skill-frontmatter/SKILL.mdmodified prune process reminders, keep thedisable-model-invocation: falseprohibition.claude-plugin/marketplace.jsonmodified plan-agent 5.0.0 → 5.1.0, git-agent 4.7.0 → 4.8.0, skill-reviewer 2.2.8 → 2.3.0kit/plugins/plan-agent/CHANGELOG.mdmodified 5.1.0 entrykit/plugins/git-agent/CHANGELOG.mdmodified 4.8.0 entrykit/plugins/skill-reviewer/CHANGELOG.mdmodified 2.3.0 entry.github/workflows/check-plugin-versions.ymlmodified add a step runningbash tests/plugins/test-imperative-pruning.sh
Steps
The step-by-step work, in order — each step says what to do, why it matters, and how to check it worked.
tests/fixtures/imperative-baselines/keep-phrases.txt as <skill path>\t<literal phrase> lines copied verbatim from the source, where each phrase must be a substring of a single source line — these SKILL.md files are hard-wrapped, so a guard whose sentence spans a line break contributes only its longest single-line fragment (for example never pass that flag on the strength of a merge, which wraps before "approval").
grep -F even when the guard is fully intact.while IFS=$'\t' read -r f p; do grep -qF "$p" "$f" || echo "MISSING $p"; done < tests/fixtures/imperative-baselines/keep-phrases.txt prints nothing, and cut -f1 tests/fixtures/imperative-baselines/keep-phrases.txt | sort -u | wc -l prints 5.tests/plugins/test-skill-behavior-baselines.sh plus tests/fixtures/imperative-baselines/scenarios/, giving each target skill one fixed input — a known todo plan spec for build, a known objective for implementation-plan, a known dirty tree in a throwaway git init sandbox for branch-agent and ship-autonomous (dry-run to the pre-flight guard only, never reaching gh pr create), and a known SKILL.md copy for optimizing-skill-frontmatter; the harness runs each skill headless via claude -p --plugin-dir kit/plugins/<plugin> and records only structural facts — files written and their paths, gates that fired, refusals emitted.
bash tests/plugins/test-skill-behavior-baselines.sh --record writes five .expected manifests under tests/fixtures/imperative-baselines/ and a plain re-run exits 0 against the unmodified skills.claude CLI not found — behavioral baselines cannot be skipped when the CLI is absent, and confirm it writes nothing outside its sandbox and leaves no temp directories behind.
PATH=/usr/bin:/bin bash tests/plugins/test-skill-behavior-baselines.sh; echo $? prints 1, and git status --porcelain is clean after a full run.tests/plugins/test-imperative-pruning.sh asserting four things — every keep-phrases.txt entry is literally present in its named file, each of the five description: frontmatter lines is byte-identical to git show origin/main:<path>, all five .expected manifests exist and are non-empty, and the behavioral harness is invoked and must pass when command -v claude succeeds — then wire it into .github/workflows/check-plugin-versions.yml as a named step with a comment explaining why.
bash tests/plugins/test-imperative-pruning.sh exits 0 on the unmodified tree, and grep -q test-imperative-pruning .github/workflows/check-plugin-versions.yml succeeds.kit/plugins/.
git show --stat HEAD | grep -c 'kit/plugins/' prints 0 and bash tests/plugins/test-imperative-pruning.sh exits 0 at that commit.kit/plugins/plan-agent/skills/build/SKILL.md and kit/plugins/plan-agent/skills/implementation-plan/SKILL.md — remove build's Step 0 rationale sentence "This skill writes source files, so it cannot run inside plan mode." and implementation-plan's equivalent "writing plan files is a filesystem mutation that cannot proceed inside harness plan mode", build's Step 2 aside "Spec edits only (see the source-of-truth rule above)" now that the ## Overview source-of-truth paragraph carries it, the per-step re-render reminders that merely re-point at the ## Re-render (subroutine — referenced by every step below) section, and implementation-plan's Step 4 Rename restatement of the kebab-case verb-target convention already given in Step 2's write instruction — leaving ## Scope Constraint — Plans Only and every keep-phrases.txt entry untouched.
bash tests/plugins/test-imperative-pruning.sh exits 0, and bash tests/plugins/test-build-skill.sh still exits 0; if the baseline for either skill fails and one fix attempt does not explain it, restore the removed text and re-run.kit/plugins/git-agent/skills/ship-autonomous/SKILL.md and kit/plugins/git-agent/skills/branch-agent/SKILL.md — remove "Run Steps 0–5 in strict order", branch-agent's opening "Follow these steps in strict order." plus the "STOP immediately after step 6." clause it wraps into, now that the ## Step 6: Confirm and STOP block states the same rule, and both Step 0 mutation rationales — leaving the merge-on-green guard, the --delete-branch guard, the review-dismissal guard, and both stash guards literally intact.
ship-autonomous ends in an irreversible merge and branch-agent handles the user's uncommitted work, so these are the two files where a wrong deletion is unrecoverable.bash tests/plugins/test-imperative-pruning.sh exits 0, bash tests/plugins/test-merge-shorthand.sh exits 0, and grep -cF 'Never merge on anything but green' kit/plugins/git-agent/skills/ship-autonomous/SKILL.md prints 1.kit/plugins/skill-reviewer/skills/optimizing-skill-frontmatter/SKILL.md — remove "Follow these steps exactly", "Count carefully — both limits apply independently", and the Step 0 plan-mode rationale — keeping the disable-model-invocation: false prohibition and the placeholder-value prohibition verbatim.
bash tests/plugins/test-imperative-pruning.sh exits 0 and bash tests/plugins/test-description-budget.sh exits 0.plan-agent to 5.1.0, git-agent to 4.8.0, and skill-reviewer to 2.3.0 in .claude-plugin/marketplace.json, and add a matching entry to each plugin's CHANGELOG.md naming the pruned skills and stating that behavior baselines were recorded before the prune.
kit/plugins/<name>/ requires a version higher than main, and the changelog is where a future reader learns the baselines exist.BASE_REF=main node scripts/check-plugin-versions.mjs exits 0 and each of the three CHANGELOG.md files contains its new version heading. Shipped as plan-agent 7.2.0, git-agent 4.9.0, skill-reviewer 2.4.0 — the targets named above were already met or exceeded on main by the time this ran, so bumping to them would have been a regression; see the Completion Report. The step text is left as originally specified so the deviation stays visible rather than being edited out of the record.Tests
The tests that prove the change does what it promises.
keep-phrases.txt phrase is present in its named SKILL.md, all five description: lines are byte-identical to origin/main, all five .expected manifests exist and are non-empty, and the behavioral harness passes when the claude CLI is available; Run: bash tests/plugins/test-imperative-pruning.shbuild, implementation-plan, ship-autonomous, branch-agent, optimizing-skill-frontmatter run headless against fixed scenario inputs; Key cases: build on a known todo spec writes only that spec's files and sets status: in-progress; implementation-plan writes nothing outside the plans directory (the Scope Constraint); branch-agent on a conflicting dirty tree stashes, branches, and pops without losing a file; ship-autonomous stops at the pre-flight guard on a clean tree and never reaches gh; optimizing-skill-frontmatter never emits disable-model-invocation: false; missing claude CLI exits 1 rather than skipping.build skill's three implementation gates and the merge skill's safety gates; Key cases: both exit 0 unchanged after Steps 6 and 7.Definition of done
The plan counts as done when every statement below is true — check each one off as you verify it.
Final check
One last pass to confirm the whole change works end to end.
Run bash tests/plugins/test-imperative-pruning.sh on the final tree; expected result is exit 0 with a per-skill line for all five targets and a baselines: 5/5 match summary. Then run BASE_REF=main node scripts/check-plugin-versions.mjs; expected result is exit 0 listing three bumped plugins.
Tautology check — the tests must be able to fail. Delete the line **Never merge on anything but green.** from kit/plugins/git-agent/skills/ship-autonomous/SKILL.md and re-run bash tests/plugins/test-imperative-pruning.sh; it must exit 1 naming that phrase. Restore the line with git checkout -- kit/plugins/git-agent/skills/ship-autonomous/SKILL.md and confirm the test returns to exit 0. Repeat the same break-and-revert on a description: line (append one character) and confirm the description-drift assertion exits 1 — a test that only catches the first break is half a test.
Behavioral check — actually run the skills, do not infer from word counts. In a scratch clone, load the three plugins with claude --plugin-dir kit/plugins/plan-agent --plugin-dir kit/plugins/git-agent --plugin-dir kit/plugins/skill-reviewer and exercise each pruned skill by hand against its recorded scenario: invoke /plan-agent:build on the fixture plan and confirm it walks the steps, ticks the spec, re-renders, and stops without committing; invoke /plan-agent:implementation-plan on the fixture objective and confirm it writes only under the plans directory and refuses to touch source; run /git-agent:branch-agent on the conflicting dirty tree and confirm the stash/branch/pop cycle loses nothing (git stash list empty, all modified files present); trigger git-agent:ship-autonomous on a clean tree and confirm it stops at "Nothing to ship" rather than proceeding; run /skill-reviewer:optimizing-skill-frontmatter on a fixture SKILL.md and confirm it never writes disable-model-invocation: false. A reduced word count is not evidence any of this still works.
If any behavioral check diverges from its recorded baseline and one fix attempt does not explain the divergence, restore the removed imperative for that skill, re-run bash tests/plugins/test-imperative-pruning.sh to confirm green, and record the restoration in the plugin's CHANGELOG entry. A partial prune that keeps a guard is the correct outcome, not a failed plan.
Wrapping up
Three gates that must all pass before this plan is marked completed.
Completion Report
- Premise stale, contract intact
- this spec was written before PRs #487/#489 split these skills into cores plus references. Re-measured at implementation time, three of five targets had already shrunk ~80%:
ship-autonomous2,448 to 597 words (20 to 13 imperatives),branch-agent1,515 to 582,optimizing-skill-frontmatter3,153 to 579 with 1 matching line, down from 14. Several named DROP targets were already gone, and the KEEP guardDo not drop the stash before staging resolved fileshad moved intobranch-agent/references/stash-and-recovery.md. The real remaining DROP set was five items, not 118, so the prune is 7 insertions and 14 deletions across five files. The spec's own honest accounting predicted the smallest return of the set; the true figure is smaller still, and that is the finding rather than a shortfall. - Version targets unsatisfiable as written
marketplace.jsonalready readplan-agent 7.0.1,git-agent 4.8.0,skill-reviewer 2.3.0, so the specified 5.1.0 / 4.8.0 / 2.3.0 would have been a version regression thatcheck-plugin-versions.mjsrejects. Shipped as plan-agent 7.2.0, git-agent 4.9.0, skill-reviewer 2.4.0, honouring the criterion's intent of three plugins abovemainwith the guard exiting 0. The plan-agent figure moved twice: 7.1.0 was chosen against amainreading 7.0.1, then #494 landed mid-review and took 7.1.0 itself, so the rebase renumbered this branch to 7.2.0.- Harness defect, inherited stdin
- headless runs inherited the caller's stdin; when the harness itself ran non-interactively that stdin is a pipe nobody closes, and the run wedged after
claudeexited. Fixed by redirecting the run from/dev/null. - Harness defect, watchdog held the manifest pipe
- the timeout watchdog subshell inherited stdout, which is the write end of the
scenario_fn | sortmanifest pipe, and killing the subshell orphaned itssleep, which kept that descriptor open. Every run therefore blocked for the fullRUN_TIMEOUTno matter when the work finished: 900s pinned, versus 84s measured after the fix. A gate that always takes fifteen minutes is a gate that gets switched off, so this was fixed rather than tolerated. - Behaviour proven neutral, not inferred
- baselines were recorded against the unmodified skills, independently reproduced 5/5 before any edit, committed as
ed6b854touching no shipped plugin file, and reproduced 5/5 again after the prune. Reproducing them before editing is what makes a later divergence a regression rather than model nondeterminism. The load-bearing facts held:implementation-planleft the source file its objective named byte-identical with zero writes outside the plans dir;branch-agent's stash and pop returned every file with its contents and an emptygit stash list, adding no commit;buildpromotedstatus:tocompletedand created only the file its plan named;optimizing-skill-frontmatternever emitteddisable-model-invocation: false;ship-autonomousleftHEAD, branch, and working tree unchanged. - Gate proven falsifiable
- deleting the merge-on-green guard makes
tests/plugins/test-imperative-pruning.shexit 1 naming that exact phrase, and appending one character to adescription:line makes it exit 1 reporting the drift; both return to exit 0 ongit checkout --. All 29 KEEP-classified guards are literally present and no description line drifted. - Abort condition never invoked
- no imperative had to be restored.
- Description drift compared against a golden file, not
origin/<base> - Step 4 specified a byte-identical comparison with
git show origin/main:<path>, which is correct for this PR but wrong as a permanent CI gate: it would block every future legitimate description update to these five skills with no way to pass. The assertion now compares againsttests/fixtures/imperative-baselines/descriptions.expected, so a deliberate change updates that file in the same commit and surfaces in review as an explicit act. Accidental drift still fails; both directions are covered by the tautology check. Raised in review as Codex P2. gh_invokedreplaced bygh_mutating_invoked- the original boolean could not distinguish a read-only pre-flight
gh auth statusfromgh pr create, and the recordedship-autonomousvalue ofyeswould have failed a correct run that stopped before reaching anyghcall. Evidence from a kept sandbox showed the run invoked onlygh auth statusand then stopped, refusing to route around the stub — no mutation occurred. The manifest now asserts the fact that matters, with a read-only allowlist and deny-by-default for anything else. Raised in review as Codex P2. BASELINE_ONLYnow rejects unknown scenario names- a typo previously matched nothing, leaving
TOTAL=0, printingbaselines: 0/0 match, and exiting 0 without running a single skill. That is the same silent-green failure the CLI-absent gate exists to prevent. Unknown values now exit 1, and a zero-scenario run refuses to report success. Raised in review as Codex P2.