Close the plan-completion gap from both ends: make every prompt that starts an implementation require the plan be checked off, and add a PostToolUse hook that deterministically catches a plan whose steps and criteria are all checked while its status still says otherwise — then unlock finalize-plan so the model can act on that signal.
Plans get implemented and then quietly stay at "todo" forever, so the gallery lies about what shipped. The completion machinery already exists — it is just unreachable unless you happen to be standing in the right skill at the right moment. This fixes both ends: the prompts now ask every agent to check the plan off, and a deterministic detector fires the instant one contradicts itself.
Read and implement all steps in the plan at docs/plans/fix-plan-completion-drift.md — Make a finished plan admit it is finished. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/fix-plan-completion-drift.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
More ways to run this plan — goal & workflow prompts, file path
Achieve this goal: Make a finished plan admit it is finished. The plan at docs/plans/fix-plan-completion-drift.md describes one approach — use it as reference, but optimize for the outcome. Fan out across parallel subagents where that serves the outcome. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/fix-plan-completion-drift.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
Run a workflow to implement the plan at docs/plans/fix-plan-completion-drift.md — Make a finished plan admit it is finished. Brief subagents with the plan file at docs/plans/fix-plan-completion-drift.md. Reserve a final verification phase for the lead agent, not a subagent. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/fix-plan-completion-drift.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
fix-plan-completion-drift.html
docs/plans/fix-plan-completion-drift.html
docs/plans/fix-plan-completion-drift.md
Context
The story behind this plan — what prompted the work and why it matters now.
Plans generated by plan-agent routinely finish implementation and never get
marked completed. The gallery then misrepresents what shipped.
The completion machinery is not missing — it is unreachable. Verified in the
current tree:
- skills/implementation-plan/SKILL.md:423,439,461 — the acceptance-criteria,
end-to-end-verification, and completion-checklist gates are all nested inside
the Implement now branch at :417. No other path reaches them.
- skills/implementation-plan/SKILL.md:400,516 — the sibling branch,Exit — I'll implement later, states: "No implementation, no status change —
the plan stays at todo everywhere."
- skills/finalize-plan/SKILL.md:5 — disable-model-invocation: true bars the
model from ever auto-invoking the skill. Only a manually typed/plan-agent:finalize-plan reaches it.
- hooks.json:3 — all four hooks are PostToolUse on Write|Edit|MultiEdit.
There is no Stop or SubagentStop hook, so nothing observes a session
ending.
- The only "always update status" instruction is ~/.claude/rules/plan-mode.md
§7 — a maintainer-local file that ships to nobody. Marketplace installs
(plan-agent 3.0.0) receive none of it.
Net: the dominant path — generate a plan, exit, implement it in a later
session — hits no gate at all.
### The mechanism, and why this one
Recommended: a drift-detector hook that hands off to the existing
verification skill. A new hooks/detect-plan-completion-drift.py fires on
plan-spec writes and reports one specific contradiction: every step carries[x], every acceptance criterion is - [x], and status is not completed.
It reports; it never writes status. finalize-plan — unlocked in Step 3 — does
the actual verification and status write, behind its own existing
user-confirmation gate.
This is one mechanism in two files: the detector needs a callee it can actually
reach, and today it has none.
The design turns on a distinction that matters. The hook does not judge
whether the work is done — it cannot; acceptance criteria are prose. It detects
that the document contradicts itself, which is mechanical string-matching over
committed state. Resolving the contradiction is delegated to the skill that
already verifies against the codebase and asks the user. So nothing in this
chain asks the model to self-attest to completion.
Against the four failure modes:
| Failure mode | How this handles it |
|---|---|
| Silent skip (today's bug) | Solved for the dominant case. The renderer's implement-prompt already instructs the implementing agent to tick [x] markers and criteria as it goes; what gets forgotten is the final status: flip. That forgetting is the drift state, so the last tick trips the detector. |
| False completion | Structurally prevented. The hook writes nothing and decides nothing. finalize-plan inspects codebase evidence and gates on AskUserQuestion before any status write — behavior this plan does not alter. |
| Noisy prompts | Four conditions must hold simultaneously: the written file is a # Plan: spec inside the resolved plansDirectory, every step is [x], every criterion is - [x], and status != completed. The hook cannot fire in a session with no plan in play, because it only fires on writes to plan specs. |
| Plugin-user gap | Every behavioral change lands in kit/plugins/plan-agent/** and ships via the git-subdir source. Nothing depends on ~/.claude/rules/, repo-local settings, or this repo's layout. |
What this deliberately does not solve: a session that implements a plan and
ticks nothing. With no write to the spec there is no hook event, and the plan
stays silently at todo. Closing that hole completely needs a Stop hook
scanning the transcript to distinguish "implemented this plan" from "happened to
read it" — which trades directly into the noise column and cannot make that
distinction reliably. It is named in Next Steps rather than smuggled in here.
Partially ticked plans (some steps [x], not all) are likewise out of range by
design: that is a genuinely unfinished plan, and status: in-progress is
accurate.
### Feeding the detector: the prompts must ask for ticks
A detector that only fires on ticks is worth nothing if the prompts never ask
for any. Three prompts start an implementation, and only one of them teaches
check-off:
| Prompt | Path | Instructs check-off today? |
|---|---|---|
| Implement | copyCmd → buildImplementPrompt() (plan-shell.mjs:1306) | Yes — five numbered instructions covering [x] step markers, - [x] criteria, status: completed, and re-render. |
| Goal | copyGoal (plan-shell.mjs:1390) + plan-goal meta | No — copies the bare one-liner's textContent. |
| Workflow | copyWorkflow (plan-shell.mjs:1383) + plan-workflow meta | No — copies the bare one-liner's textContent. |
The workflow gap is the worst of the three, because that prompt says "Brief
subagents with the plan file" — so every subagent is briefed with a plan and
never told to record anything. Nothing ticks, no spec write happens, and the
detector in Steps 1–2 never fires. This plan's own rendered HTML has the defect:
its plan-workflow meta tag tells subagents to implement it and says nothing
about completion.
Both the copy handlers and the meta-tag strings need fixing, because they are
different delivery paths: implementation-plan/SKILL.md:482-483 sends the user
the workflow prompt read from the plan-workflow meta tag, not from the copy
button. Fixing only the button would leave the primary workflow path untouched.
Steps 6–8 close this. Together with the detector they narrow the never-ticked
hole to sessions that ignore the prompt entirely — without the Stop hook.
### Risks
- Exit-2 styling. This plugin's hook convention is stderr + sys.exit(2)
(hooks/validate-plan-filename.py:210-211), which the UI renders as a
"blocking error" even though PostToolUse cannot block an already-run tool.
A completion nudge styled as an error is a cosmetic mismatch. Mitigation:
follow the established convention for consistency; the message text carries
the framing. Revisiting the delivery channel repo-wide is a Next Step.
- Unlocking finalize-plan widens its activation surface. Mitigated by its
narrow description and its own confirmation gate. Step 3 tightens the
description so it reads as a completion action rather than an ambient one.
- The goal prompt gets the full check-off, including per-step [x] markers.
Decided during the plan interview, over the narrower alternative of criteria +
status only. The tension it creates is real and worth stating: the goal prompt
exists to tell an agent to optimize for the outcome and treat the plan as
reference, so an agent may legitimately achieve the objective without
following the plan's steps. Ticking those steps would then record work that
never happened — a false record, adjacent to the false-completion failure mode
this plan otherwise guards against. What keeps it honest is that the inherited
instruction is already conditional ("after completing each step, mark it
done"), so a bypassed step stays unticked. The consequence is that a
goal-driven agent which met every criterion by a different route cannot reachstatus: completed, because not every step is [x]. Step 6 resolves that with
the mechanism the section catalog already provides: route the discrepancy to a## Completion Report entry rather than a false tick. Re-open this if goal-run
plans start stalling at in-progress with clean criteria.
- The version bump reads the rule's spirit over its letter. Decided during
the plan interview: .claude/rules/marketplace.md lists "changing activation
behavior" under MAJOR, and this plan changes it. It ships as 3.1.0 anyway,
because the clause targets changes that invalidate existing usage and this one
does not — the command still works, nothing an installed user does breaks. The
cost is that a user skimming version numbers sees a routine minor bump and may
not notice the skill can now fire on its own. Mitigation: Step 9 makes the
CHANGELOG entry state the activation change outright, since it is now the only
place that information surfaces.
Files that change
Every file this plan touches, and what happens to each one.
`kit/plugins/plan-agent/hooks/detect-plan-completion-drift.py`new the detector.`kit/plugins/plan-agent/hooks.json`modified register it onWrite|Edit|MultiEdit.`kit/plugins/plan-agent/skills/finalize-plan/SKILL.md`modified dropdisable-model-invocation, retune the description.`kit/plugins/plan-agent/skills/implementation-plan/SKILL.md`modified document the detector on theExitbranch and in Step 6.`kit/plugins/plan-agent/scripts/lib/plan-shell.mjs`modified generalisebuildImplementPrompt(); rewirecopyGoalandcopyWorkflowto it.`kit/plugins/plan-agent/scripts/build-plan-html.mjs`modified carry the check-off clause into theplan-goalandplan-workflowmeta tags.`tests/plugins/test-build-plan-html.mjs`modified assert all three prompts instruct check-off. Outside the plugin path by the same test convention as below.`kit/plugins/plan-agent/CHANGELOG.md`modified release entry.`.claude-plugin/marketplace.json`modified version bump. Outsidekit/plugins/plan-agent/**by necessity: for relative-path plugins this repo setsversiononly inmarketplace.json, never inplugin.json(per.claude/rules/marketplace.md). There is no in-plugin alternative.`tests/plugins/test-plan-completion-drift.sh`new Tier 1 test. Outside the plugin path by convention: every plugin test in this repo lives undertests/plugins/; no plugin directory contains tests.
Steps
The step-by-step work, in order — each step says what to do, why it matters, and how to check it worked.
kit/plugins/plan-agent/hooks/detect-plan-completion-drift.py, modelling it on hooks/render-plan-html.py — reuse its _project_dir(), _load_settings(), _get_plans_dir() and _is_plan_spec() helpers verbatim (copy, do not import; hooks run as standalone scripts). Read tool_input.file_path from stdin JSON, exit 0 immediately unless the path is a # Plan: spec inside the resolved plans directory. Then parse the spec for three facts: status: in frontmatter, step markers (^\d+\. \[x\] vs ^\d+\. ), and criteria bullets under ## Acceptance Criteria (- [x] vs - [ ]). Fire only when there is at least one step and at least one criterion, all steps are [x], all criteria are - [x], and status is not completed; write a message to stderr naming the spec path and instructing /plan-agent:finalize-plan <path>, then sys.exit(2). Every other case exits 0 silently.
echo '{"tool_input":{"file_path":"docs/plans/fix-plan-completion-drift.md"}}' | python3 kit/plugins/plan-agent/hooks/detect-plan-completion-drift.py; echo "exit=$?" prints exit=0 against this plan while it is unticked.kit/plugins/plan-agent/hooks.json as a fifth PostToolUse entry matching Write|Edit|MultiEdit, invoked as python3 "${CLAUDE_PLUGIN_ROOT}/hooks/detect-plan-completion-drift.py" with "timeout": 5.
${CLAUDE_PLUGIN_ROOT} is what makes the hook resolve for marketplace installs rather than only in this checkout — the plugin-user gap closes here or not at all.python3 -c "import json; h=json.load(open('kit/plugins/plan-agent/hooks.json'))['hooks']['PostToolUse']; print(len(h)); print([e['hooks'][0]['command'] for e in h])" prints 5 and lists the new command with ${CLAUDE_PLUGIN_ROOT} intact.kit/plugins/plan-agent/skills/finalize-plan/SKILL.md, delete the disable-model-invocation: true line (:5) and retune the description (:4) so it reads as a completion action triggered by finishing implementation — keep it under 200 chars per .claude/rules/skill-authoring.md, and drop the trailing "Use via /plan-agent:finalize-plan." now that the command is no longer the only door.
head -8 kit/plugins/plan-agent/skills/finalize-plan/SKILL.md shows no disable-model-invocation key, and grep -c disable-model-invocation kit/plugins/plan-agent/skills/finalize-plan/SKILL.md returns 0.kit/plugins/plan-agent/skills/implementation-plan/SKILL.md, amend the Exit — I'll implement later branch (:516) to state that a later implementing session is covered by the drift detector provided it ticks steps and criteria in the spec, and add one line to Step 6 (:144) recording that the detector watches for the all-ticked/not-completed contradiction. Do not move or duplicate the three gates.
Exit branch is where the dominant failure path begins, and it currently reads as a dead end; a reader there has no idea anything downstream will catch them.grep -n "drift\|detect-plan-completion-drift" kit/plugins/plan-agent/skills/implementation-plan/SKILL.md returns hits in both the Step 6 region and the Exit branch, and grep -c "Acceptance criteria gate" kit/plugins/plan-agent/skills/implementation-plan/SKILL.md still returns 1.tests/plugins/test-plan-completion-drift.sh following the shape of tests/plugins/test-finalize-all-flag.sh. Build spec fixtures in a temp dir under a fake plans directory and assert the hook's exit code for each: (a) all steps [x] + all criteria - [x] + status: in-progress → exit 2 with the spec path on stderr; (b) same but status: completed → exit 0 silent; (c) one step unticked → exit 0; (d) one criterion - [ ] → exit 0; (e) a non-plan .md in the plans dir → exit 0; (f) a # Plan: spec outside the plans dir → exit 0.
bash tests/plugins/test-plan-completion-drift.sh exits 0 and reports all six cases passing.kit/plugins/plan-agent/scripts/lib/plan-shell.mjs, generalise buildImplementPrompt() (:1306) into a shared builder that takes the source command element id (implement-cmd / goal-cmd / workflow-cmd), keeping its existing five instructions and live-DOM status line verbatim. Rewire copyGoal (:1390) and copyWorkflow (:1383) to call it instead of copying raw textContent. Keep copyCmd's output byte-identical to today's. Add one instruction that applies to every variant: if the objective was achieved without following a step, leave that step unticked and record it as a ## Completion Report bullet in the spec rather than ticking it.
copyGoal/copyWorkflow produce text containing [x], status: completed, and the Completion Report clause, and that copyCmd's output still matches a pre-change capture except for that one added clause.kit/plugins/plan-agent/scripts/build-plan-html.mjs, extend the goal (:176) and workflow (:180) strings so the check-off requirement survives into the plan-goal and plan-workflow meta tags — append a clause instructing the agent to tick [x] step markers and - [x] criteria in the spec and set status: completed when done. For the workflow string, place it so it applies to the briefed subagents, not just the orchestrator.
implementation-plan/SKILL.md:482-483 hands the user the workflow prompt read from the meta tag, so a copy-handler-only fix would miss the primary workflow path entirely.grep -o '<meta name="plan-workflow"[^>]*>' docs/plans/fix-plan-completion-drift.html contains both [x] and status: completed; same for plan-goal.tests/plugins/test-build-plan-html.mjs. Assert that the rendered plan-goal and plan-workflow meta tags each contain the check-off clause, that the shared builder is wired to all three copy handlers (copyCmd, copyGoal, copyWorkflow all reference it), and that a plan with workflow: false still emits a plan-goal tag carrying the clause.
node tests/plugins/test-build-plan-html.mjs exits 0 with the new assertions reported.3.1.0 entry to kit/plugins/plan-agent/CHANGELOG.md above the 3.0.0 entry, dated 2026-07-15, with ### Added (the drift-detector hook) and ### Changed (finalize-plan is now model-invocable — state explicitly that the /plan-agent:finalize-plan command is unchanged and existing invocations keep working, and that the skill may now activate on its own when a plan finishes).
head -3 kit/plugins/plan-agent/CHANGELOG.md shows the ## 3.1.0 heading, and grep -n "finalize-plan" kit/plugins/plan-agent/CHANGELOG.md | head -1 lands inside the new entry.plan-agent's version in .claude-plugin/marketplace.json from 3.0.0 to 3.1.0.
.claude/rules/marketplace.md) plus a widened activation surface that breaks nothing. /plan-agent:finalize-plan keeps working identically, so no installed user has to change anything. This is a deliberate reading of the rule's spirit over its letter: the rule lists "changing activation behavior" under MAJOR, but that clause is aimed at changes that invalidate existing usage, and nothing here does.python3 -c "import json; print([p['version'] for p in json.load(open('.claude-plugin/marketplace.json'))['plugins'] if p['name']=='plan-agent'])" prints ['3.1.0'], and the file still parses.Tests
The tests that prove the change does what it promises.
tests/plugins/test-plan-completion-drift.sh; Type: smoke; Asserts: feeding the hook a spec with all steps [x], all criteria - [x] and status: in-progress exits 2 and names the spec path, while the five non-drift fixtures each exit 0 silently; Run: bash tests/plugins/test-plan-completion-drift.shtests/plugins/test-plan-completion-drift.sh; Targets: _is_plan_spec() and the status/step/criteria parse in detect-plan-completion-drift.py; Key cases: non-plan .md inside the plans dir, # Plan: spec outside the plans dir, spec with zero criteria, spec with a # schema: frontmatter comment ahead of the title headingtests/plugins/test-plan-completion-drift.sh; Targets: hooks.json + detect-plan-completion-drift.py; Key cases: the fifth PostToolUse entry exists, its command interpolates ${CLAUDE_PLUGIN_ROOT}, and the referenced script path exists relative to the plugin roottests/plugins/test-build-plan-html.mjs; Targets: build-plan-html.mjs prompt strings + plan-shell.mjs copy handlers; Key cases: rendered plan-goal and plan-workflow meta tags each contain [x] and status: completed; copyCmd/copyGoal/copyWorkflow all route through the shared builder; a workflow: false plan still emits a compliant plan-goal tagDefinition of done
The plan counts as done when every statement below is true — check each one off as you verify it.
Final check
One last pass to confirm the whole change works end to end.
Walk the bug end-to-end as a plugin user would hit it, in a scratch copy of a
plan spec under the plans directory:
1. Take any plan spec with status: in-progress and at least one unticked step.
Tick the remaining steps and criteria one at a time via Edit. Nothing fires
until the final tick — confirming the hook is silent on partial progress.
2. On the edit that completes the last step or criterion, the hook fires: the
session surfaces a message naming that spec and pointing at/plan-agent:finalize-plan. This is the moment that silently did nothing
before.
3. Confirm the model can now act on it — that finalize-plan resolves without
the user typing the command, and that it still runs its own verification and
asks before writing status. A completion that was never verified must still
not happen.
4. Re-run the edit once status: completed is set. The hook stays silent —
the contradiction is gone, so the nudge is gone.
5. Edit an unrelated .md file and a source file in the same session. No hook
output. This is the noise contract holding.
Then verify the prompts actually feed the detector — the two halves are useless
apart:
6. Re-render this plan and read its own plan-goal and plan-workflow meta
tags. Both must now instruct [x] ticks, - [x] criteria, andstatus: completed. Before this change the workflow tag briefed subagents to
implement the plan and said nothing about recording it.
7. Copy each of the three prompts from the rendered page. All three must carry
the check-off loop; before this change only the implement prompt did.
8. Close the loop end-to-end: paste the workflow prompt into a fresh session and
confirm the resulting run ticks the spec as it goes, which is what trips the
detector from step 2 above. A prompt that instructs check-off and a hook that
fires on check-off are one mechanism — this step is the only place that shows
them working as one.
Wrapping up
Three gates that must all pass before this plan is marked completed.
Completion Report
No items to report — all requirements met.