Team recap · git-agent

Context guard for the ship-autonomous pipeline

The skill that ships your work now checks how long the conversation has gotten before it starts, and offers two cheaper ways to run.

2026-07-28 git-agent 4.7.0 → 4.8.0 working tree, not yet committed

At a glance

Where this landed

5Changes shipped
5Files touched
6Decisions made
4Open items

ship-autonomous is the git-agent skill that runs the whole delivery pipeline for you: branch, test, commit, open a pull request, watch the build, fix the common failures, and merge once you approve. It now opens by asking whether it should be running in this conversation at all.

The reason is that the pipeline reads every input it needs from git and GitHub — never from the conversation. All that accumulated history is cost with no benefit, and the cost repeats every time the build wakes the session back up.

Nothing is committed yet. Five files sit in the working tree, all verified: a new 12-check test passes, one deliberate sabotage of the skill text made the right check fail, and the existing git-agent suite and the version guard both still pass.

What changed

Five things are different

The pipeline asks before spending your conversation

Affects: anyone who ships Reach it: next invocation from a long session

It used to start working the moment it was invoked. Now its first step states plainly that it reads nothing from the conversation, then — only when the session is already long — offers three routes: clear (you run /clear and re-invoke, losing nothing), background (hand the work to subagents, each with its own fresh conversation, and keep this session for other work), or continue (run here anyway).

The clear route stops rather than pretending

Affects: nobody directly — a correctness guarantee

A skill cannot clear its own context. If it tried and silently failed, it would run the entire pipeline inside exactly the bloated session you asked to escape. So choosing clear ends the run and hands back to you.

A skip condition, so it is not a prompt on every run

Affects: anyone shipping from a fresh session

The guard exists to catch an expensive default, not to add friction. On a short session, or one started for this ship, the step is skipped silently and you never see it.

A test that pins the guard's own justification

Affects: the next person to edit this skill 12 checks passing

The important check asserts that the skill text literally still says “No step reads the conversation.” If somebody later adds a step that does read the transcript and edits that sentence, the test fires — because clearing context would no longer be safe.

bash tests/plugins/test-ship-autonomous-context-guard.sh

Version and changelog

Affects: anyone installing git-agent 4.7.0 → 4.8.0

Bumped in marketplace.json with a matching changelog entry. Minor rather than patch: no new command or skill was added, but the skill behaves differently now.

How it works now

The new opening, and the cost it avoids

No

Yes, so ask

clear

background

continue

merge still needsa human answer

Invoke ship-autonomous

Step 0: is this sessionalready long?

Step 0.5: exit plan mode

Which route?

STOPyou run /clear, re-invoke

ship-bg,then ship-ci-bg

Steps 1-4: guards, branch,test, commit, open PR

Step 5: watch the pull request

Step 8: merge, on your approval

What to look at: the three arrows out of “Which route?”. Only continue goes straight on, and the background route still comes back for the merge approval.
Pull requestYour sessionPull requestYour sessionwhole conversationcarried on every wakea subagent restarts thiswith an empty conversationopen PR, then end turnbuild event 1 (conversation re-sent)push a fixbuild event 2 (conversation re-sent)push a fixbuild event 3 (conversation re-sent)
What to look at: the cost is per event, not per run — so it multiplies by how many times the build fails.

Before and after

Rule by rule

BeforeAfter
The skill started working the moment it was invoked.It checks the session length first and offers cheaper routes.
Nothing pointed at the background commands, so people did not use them.The background route names both commands, in order.
Exiting plan mode was Step 0.It is Step 0.5. Every other step number is unchanged.
The README walkthrough listed 10 steps.It lists 11, with the guard first.
git-agent was at 4.7.0.4.8.0.
No test covered the skill's opening.12 checks, including step ordering and the safety claim.

Decisions

What was chosen, and what was turned down

Add the guard as text in the skill, not as new machinery

One edit to one file, and its content points at the two background commands that already exist.

Rejected — a new agent-ship-autonomous subagent plus a /ship-autonomous-bg command: two new files, and it loses the merge approval gate, because a subagent has no user to ask. A UserPromptSubmit hook: a new Python file that still cannot clear context — the same redirect with more parts. Relying on habit alone: nothing in the product enforces it.

The clear route hard-stops

A skill has no way to clear its own context. Continuing after a no-op would defeat the entire purpose of the step.

Skip the guard on a short session

A prompt on every run is friction paid by everybody to protect against a case that only some runs hit.

Insert as Step 0 and demote the old Step 0 to Step 0.5

The skill's later steps cross-reference each other by number about a dozen times — “return to Step 6”, “Step 8 blocks on this”. The file already used a half-step (Step 2.5), so this matches its own convention.

Rejected — renumbering every step, for the blast radius across all those cross-references.

Merging still comes back to the foreground

The background route ends at the merge gate on purpose.

Rejected — letting a subagent merge. The existing agent-merge already does exactly that when you want it, and folding it in here would remove a human decision from an irreversible action.

Minor version bump, not patch

No new command or skill was added, but the skill behaves differently.

Rejected — patch, as understating the change.

Learnings

Three things worth carrying forward

A grep-based test can be tautological, so one was deliberately broken.

The stop-on-clear assertion was the subtle one, so the clause was deleted from the skill, the suite re-run (check 8 failed, as intended), and the file restored (all 12 passed again). A passing grep only proves the text is present; the mutation proves the grep is load-bearing.

The shell tool's working directory persists between calls.

A chmod failed with “No such file or directory” because an earlier compound command had left the shell inside kit/plugins/git-agent. Absolute paths avoid the whole class of problem.

Markdown renumbers ordered lists by position, but the source still matters.

Inserting a step into the README left two literal 3. entries. It renders correctly and reads wrong, so the remaining items were re-sequenced with a small script rather than by hand.

Open items

What is not done

Nothing is committed.

Five files are in the working tree. They need a commit, a branch, and a pull request.

The new test is not wired into continuous integration.

The workflows under .github/workflows/ name individual test scripts one by one — check-plugin-versions.yml and publish-dist.yml each list a handful — and the new script is in neither. It runs only when someone runs it. Adding it to check-plugin-versions.yml beside the other git-agent tests is the obvious follow-up.

“Is this session long?” is a judgment call, not a measurement.

No tool reports the current context size, so the guard relies on the model assessing its own conversation. How reliably it does that across models is untested.

No background version of the full pipeline was built, deliberately.

If one-command unattended shipping is wanted later — and a report-instead-of-merge ending is acceptable — that is the change to make.

Files touched

Five files, four areas

The skill

  • kit/plugins/git-agent/skills/ship-autonomous/SKILL.md New Step 0 context guard; the former Step 0 (exit plan mode) is now Step 0.5.

Documentation

  • kit/plugins/git-agent/README.md The ship-autonomous walkthrough now opens with the guard; the numbered list was re-sequenced 1–11.
  • kit/plugins/git-agent/CHANGELOG.md v4.8.0 entry covering the guard, the three routes, why no new agent was added, and the test.

Tests

  • tests/plugins/test-ship-autonomous-context-guard.sh New, 12 checks: step ordering, the no-session-context claim, all three routes, the stop guarantee, the skip condition, tool permissions, and README sync.

Marketplace

  • .claude-plugin/marketplace.json git-agent 4.7.0 → 4.8.0.

Glossary

Terms used on this page

ship-autonomous
The git-agent skill that runs the whole delivery pipeline: branch, test, commit, open a pull request, watch the build, fix common failures, and merge once you approve.
context
Everything said so far in a session. It is re-sent to the model on every turn, so a long session costs more per turn than a short one.
subagent
A separate assistant dispatched to do one job. It starts with an empty conversation and reports back, so its work does not add to yours.
Step 0.5
A half-numbered step. This skill already used Step 2.5; half-steps let a step be inserted without renumbering the ones other steps refer to by number.
gh
GitHub's official command-line tool. The pipeline uses it to open pull requests and read build status.
CI
Continuous integration — the automated build and test run that fires on every push to a pull request.
marketplace.json
The registry at the repository root listing every plugin and its version. Editing a plugin requires bumping its version here.
mutation test
Deliberately breaking the thing under test to confirm the test actually fails. Guards against tests that pass no matter what.
allowed-tools
The line in a skill file listing which tools it may use. Omitting one causes a permission prompt in the middle of a run.