The skill that ships your work now checks how long the conversation has
gotten before it starts, and offers two cheaper ways to run.
2026-07-28git-agent 4.7.0 → 4.8.0working tree, not yet committed
At a glance
Where this landed
5Changes shipped
5Files touched
6Decisions made
4Open items
ship-autonomous is the git-agent skill that runs the whole delivery
pipeline for you: branch, test, commit, open a pull request, watch the
build, fix the common failures, and merge once you approve. It now opens
by asking whether it should be running in this conversation at all.
The reason is that the pipeline reads every input it needs from
git and GitHub — never from the conversation. All that
accumulated history is cost with no benefit, and the cost repeats every
time the build wakes the session back up.
Nothing is committed yet. Five files sit in the working tree, all
verified: a new 12-check test passes, one deliberate sabotage of the skill
text made the right check fail, and the existing git-agent suite and the
version guard both still pass.
What changed
Five things are different
The pipeline asks before spending your conversation
Affects: anyone who shipsReach it: next invocation from a long session
It used to start working the moment it was invoked. Now its first step
states plainly that it reads nothing from the conversation, then —
only when the session is already long — offers three routes:
clear (you run /clear and re-invoke, losing nothing),
background (hand the work to subagents, each with its own fresh
conversation, and keep this session for other work), or
continue (run here anyway).
The clear route stops rather than pretending
Affects: nobody directly — a correctness guarantee
A skill cannot clear its own context. If it tried and silently failed,
it would run the entire pipeline inside exactly the bloated session you
asked to escape. So choosing clear ends the run and hands back
to you.
A skip condition, so it is not a prompt on every run
Affects: anyone shipping from a fresh session
The guard exists to catch an expensive default, not to add friction. On
a short session, or one started for this ship, the step is skipped
silently and you never see it.
A test that pins the guard's own justification
Affects: the next person to edit this skill12 checks passing
The important check asserts that the skill text literally still says
“No step reads the conversation.” If somebody later
adds a step that does read the transcript and edits that
sentence, the test fires — because clearing context would no
longer be safe.
Bumped in marketplace.json with a matching changelog entry.
Minor rather than patch: no new command or skill was added, but the
skill behaves differently now.
How it works now
The new opening, and the cost it avoids
What to look at: the three arrows out of “Which route?”. Only
continue goes straight on, and the background route still comes
back for the merge approval.
What to look at: the cost is per event, not per run — so it
multiplies by how many times the build fails.
Before and after
Rule by rule
Before
After
The skill started working the moment it was invoked.
It checks the session length first and offers cheaper routes.
Nothing pointed at the background commands, so people did not use them.
The background route names both commands, in order.
Exiting plan mode was Step 0.
It is Step 0.5. Every other step number is unchanged.
The README walkthrough listed 10 steps.
It lists 11, with the guard first.
git-agent was at 4.7.0.
4.8.0.
No test covered the skill's opening.
12 checks, including step ordering and the safety claim.
Decisions
What was chosen, and what was turned down
Add the guard as text in the skill, not as new machinery
One edit to one file, and its content points at the two background commands that already exist.
Rejected — a new agent-ship-autonomous subagent
plus a /ship-autonomous-bg command: two new files, and it
loses the merge approval gate, because a subagent has no user to ask.
A UserPromptSubmit hook: a new Python file that still
cannot clear context — the same redirect with more parts. Relying
on habit alone: nothing in the product enforces it.
The clear route hard-stops
A skill has no way to clear its own context. Continuing after a no-op
would defeat the entire purpose of the step.
Skip the guard on a short session
A prompt on every run is friction paid by everybody to protect against
a case that only some runs hit.
Insert as Step 0 and demote the old Step 0 to Step 0.5
The skill's later steps cross-reference each other by number about a
dozen times — “return to Step 6”,
“Step 8 blocks on this”. The file already used a
half-step (Step 2.5), so this matches its own convention.
Rejected — renumbering every step, for the blast radius
across all those cross-references.
Merging still comes back to the foreground
The background route ends at the merge gate on purpose.
Rejected — letting a subagent merge. The existing
agent-merge already does exactly that when you want it,
and folding it in here would remove a human decision from an
irreversible action.
Minor version bump, not patch
No new command or skill was added, but the skill behaves differently.
Rejected — patch, as understating the change.
Learnings
Three things worth carrying forward
A grep-based test can be tautological, so one was deliberately broken.
The stop-on-clear assertion was the subtle one, so the clause was
deleted from the skill, the suite re-run (check 8 failed, as intended),
and the file restored (all 12 passed again). A passing grep only proves
the text is present; the mutation proves the grep is load-bearing.
The shell tool's working directory persists between calls.
A chmod failed with “No such file or
directory” because an earlier compound command had left the
shell inside kit/plugins/git-agent. Absolute paths avoid
the whole class of problem.
Markdown renumbers ordered lists by position, but the source still matters.
Inserting a step into the README left two literal 3.
entries. It renders correctly and reads wrong, so the remaining items
were re-sequenced with a small script rather than by hand.
Open items
What is not done
Nothing is committed.
Five files are in the working tree. They need a commit, a branch, and a pull request.
The new test is not wired into continuous integration.
The workflows under .github/workflows/ name individual test
scripts one by one — check-plugin-versions.yml and
publish-dist.yml each list a handful — and the new
script is in neither. It runs only when someone runs it. Adding it to
check-plugin-versions.yml beside the other git-agent tests
is the obvious follow-up.
“Is this session long?” is a judgment call, not a measurement.
No tool reports the current context size, so the guard relies on the
model assessing its own conversation. How reliably it does that across
models is untested.
No background version of the full pipeline was built, deliberately.
If one-command unattended shipping is wanted later — and a
report-instead-of-merge ending is acceptable — that is the change
to make.
Files touched
Five files, four areas
The skill
kit/plugins/git-agent/skills/ship-autonomous/SKILL.mdNew Step 0 context guard; the former Step 0 (exit plan mode) is now Step 0.5.
Documentation
kit/plugins/git-agent/README.mdThe ship-autonomous walkthrough now opens with the guard; the numbered list was re-sequenced 1–11.
kit/plugins/git-agent/CHANGELOG.mdv4.8.0 entry covering the guard, the three routes, why no new agent was added, and the test.
Tests
tests/plugins/test-ship-autonomous-context-guard.shNew, 12 checks: step ordering, the no-session-context claim, all three routes, the stop guarantee, the skip condition, tool permissions, and README sync.
The git-agent skill that runs the whole delivery pipeline: branch, test, commit, open a pull request, watch the build, fix common failures, and merge once you approve.
context
Everything said so far in a session. It is re-sent to the model on every turn, so a long session costs more per turn than a short one.
subagent
A separate assistant dispatched to do one job. It starts with an empty conversation and reports back, so its work does not add to yours.
Step 0.5
A half-numbered step. This skill already used Step 2.5; half-steps let a step be inserted without renumbering the ones other steps refer to by number.
gh
GitHub's official command-line tool. The pipeline uses it to open pull requests and read build status.
CI
Continuous integration — the automated build and test run that fires on every push to a pull request.
marketplace.json
The registry at the repository root listing every plugin and its version. Editing a plugin requires bumping its version here.
mutation test
Deliberately breaking the thing under test to confirm the test actually fails. Guards against tests that pass no matter what.
allowed-tools
The line in a skill file listing which tools it may use. Omitting one causes a permission prompt in the middle of a run.