Replace build's grep drift check with a deterministic render check

Medium completed
2026-08-14 agentics feature Medium effort

Add a --check mode to plan-agent-render that proves a plan's HTML is current and that a completed spec is internally consistent, then rewire build Step 5.3 and finalize-plan to run it instead of describing DOM invariants the model can only reach with Grep.

At a glance

Plan build's final gate tells the model to confirm five DOM invariants named as CSS selectors, with no way to evaluate them — so it greps the HTML and gets false drift. This adds plan-agent-render --check, which byte-compares the rendered file against a fresh in-memory render and asserts a completed spec is internally consistent, and rewires the gate to call it.

Implement Read and implement all steps in the plan at docs/plans/replace-grep-drift-check.md — Replace build's grep drift check with a deterministic render check. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/replace-grep-drift-check.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
More ways to run this plan — goal & workflow prompts, file path
Pursue as goal — optimize for the outcome
Achieve this goal: Replace build's grep drift check with a deterministic render check. The plan at docs/plans/replace-grep-drift-check.md describes one approach — use it as reference, but optimize for the outcome. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/replace-grep-drift-check.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
File replace-grep-drift-check.html
Path docs/plans/replace-grep-drift-check.html
Spec docs/plans/replace-grep-drift-check.md
Definition of done 12 / 12 done

Context

The story behind this plan — what prompted the work and why it matters now.

The 2026-08-14 usage report names grep-as-verification as the single largest
source of wasted passes: "Verification greps counted CSS selectors instead of
rendered elements, producing false drift signals and forcing a second
verification pass"
and *"Grep patterns didn't match the renderer's actual
markup, so Claude needed repeated passes just to confirm the spec's
status/checkbox state."*

The instruction causes the behaviour.
skills/build/references/completion-gates.md Step 5.3 says to confirm the
HTML matches the spec, and then names the evidence as five CSS selectors —
.step-card elements completed, criteria inputs checked, three status
representations, cc1–cc3, completion-checklist carrying all-complete. It
prescribes no mechanism. A model handed selector names and no tool reaches for
Grep, which searches the source markup rather than evaluating anything, so a
selector defined in the stylesheet counts as a match and a class the renderer
emits conditionally does not. Both failure directions were observed.

The check is aimed at the wrong artifact. The HTML is derived from the
spec by a deterministic function this repo owns. Two separate properties got
conflated into one DOM inspection:

1. Is the HTML current? — a build-freshness question. If the file on disk was
produced from this spec, it is byte-identical to a fresh render. A stale
file means a re-render was skipped, the write partially failed, or someone
hand-edited the HTML. All three are caught by one comparison.
2. Is a completed spec internally consistent? — a Markdown question. Every
step marked [x], every criterion - [x], status: completed. Nothing
here needs the HTML at all; the render is what displays the answer, not
what determines it.

Neither property requires reading rendered markup, so the whole DOM-inspection
framing goes away rather than getting a better tool.

The parser is already written and already proven. scripts/lib/plan-spec.mjs
carries parseSpecMarkdown and extractSections, and the round-trip property
extractSections(render(spec)) deep-equals the spec's sections is already
enforced by tests/plugins/test-build-plan-html.mjs. --check composes
existing, tested pieces.

The renderer lives at the repo root; the plugin carries a copy. scripts/
is canonical and kit/plugins/plan-agent/scripts/ is a bundled duplicate —
byte-identical today, and held that way by an assertion in
tests/plugins/test-build-plan-html.mjs whose failure message is *"drifted
from scripts/… — re-copy it"*. The unit tests import the root module
(../../scripts/build-plan-html.mjs), so an edit confined to the plugin copy
would both fail the parity assertion and leave every test blind to --check.
Four files are mirrored this way: build-plan-html.mjs,
extract-plan-spec.mjs, lib/plan-spec.mjs, and lib/plan-shell.mjs. Every
step below edits the root and re-copies, in that order.

Determinism holds at a fixed output path — measured, not assumed. The
renderer keeps plan-created stable across re-renders by reading the value
back from the existing HTML, and created comes from frontmatter. Rendering
docs/plans/replace-grep-drift-check.md twice over the same output path was
verified byte-identical on 2026-08-14; rendering it to two different paths
differs on exactly two lines, plan-file and plan-path, which
build-plan-html.mjs:537-538 derives from the output path and
lib/plan-shell.mjs:1884-1885 emits as meta tags.

That distinction is load-bearing. Those two fields are semantic, not noise: a
plan's recorded path is part of what the file asserts about itself. Normalizing
them away to make a cross-path comparison succeed would let --check pass an
HTML file copied in from another location — a stale state it is specifically
meant to catch. Step 1 therefore fixes the output path and compares there,
and any normalization is a last resort that must name the field it drops.

Scope. This does not touch the renderer's output, the plan CSS, or the
gallery. --check writes nothing.

Files that change

Every file this plan touches, and what happens to each one.

agentics/
  • scripts/build-plan-html.mjs modified canonical source; --check mode: render to memory, compare, assert spec consistency, print a table
  • kit/plugins/plan-agent/scripts/build-plan-html.mjs modified byte-identical re-copy of the above, held by the parity assertion
  • kit/plugins/plan-agent/skills/build/references/completion-gates.md modified Step 5.3 becomes one command; the selector list is deleted
  • kit/plugins/plan-agent/skills/finalize-plan/references/write-completions.md modified same gate, kept consistent per completion-gates.md's own instruction
  • kit/plugins/plan-agent/
    • README.md modified document --check
    • CHANGELOG.md modified 9.4.0 entry
  • .claude-plugin/marketplace.json modified plan-agent 9.3.0 to 9.4.0
  • tests/plugins/test-build-plan-html.mjs modified determinism and check-mode assertions

Steps

The step-by-step work, in order — each step says what to do, why it matters, and how to check it worked.

1
done Establish that rendering an unchanged spec is byte-deterministic at a fixed output path: render a committed plan spec, keep a copy of the bytes, render again over that same path, and diff the copy against the result. Compare at one path only — never across two — because plan-file and plan-path are derived from the output path and are meant to differ when it does.
Why
--check is a byte comparison, so a genuinely volatile field — a timestamp, a generated id, a locale-dependent date — would make it fail on every correct plan; and normalizing the path fields to force a cross-path diff to zero would make --check blind to an HTML file copied in from somewhere else, which is one of the stale states it exists to catch.
Verify
the same-path diff is empty; if it is not, name the volatile field in this plan's Context and carry it as a normalization the comparison applies before diffing.
2
done Add --check to the canonical scripts/build-plan-html.mjs: parse the spec, render to a string, compare against the existing output file, and report the first differing line with its line number and 40 characters of context on each side. Missing output file is a FAIL naming the render command, not a crash.
Why
freshness is the property the old Step 5.3 was actually reaching for, and a first-difference report is what makes the failure actionable — a bare "files differ" sends the model back to grepping; the root file is the one the tests import, so implementing anywhere else leaves the change untested.
Verify
--check on a freshly rendered plan exits 0 and prints html PASS; the same plan with one character edited into its HTML exits non-zero and prints the edited line's number.
3
done Extend --check with the spec-consistency assertions, evaluated on the parsed Markdown and skipped entirely unless status: completed: every numbered step carries [x], every ## Acceptance Criteria bullet is - [x].
Why
this is the half of Step 5.3 that is not a freshness question, and evaluating it on the spec is what removes the last reason to open the HTML.
Verify
a spec with status: completed and one - [ ] criterion exits non-zero and names that criterion's text; the same spec at status: in-progress exits 0 with the consistency rows reported as skipped.
4
done Make --check print a fixed PASS/FAIL table — one row per property (html, steps, criteria), a summary line, and exit 0 only when every row passes.
Why
the model needs to know which property broke to fix the right file, and a stable table is also what the test asserts against.
Verify
the table's row labels appear in the same order for a passing plan, a stale-HTML plan, and an inconsistent-spec plan.
5
done Re-copy the edited renderer to kit/plugins/plan-agent/scripts/build-plan-html.mjs so the bundled copy is byte-identical again, and confirm the other three mirrored files (extract-plan-spec.mjs, lib/plan-spec.mjs, lib/plan-shell.mjs) were not touched.
Why
the parity assertion fails the moment the two diverge, and its failure message — "drifted from scripts/… — re-copy it" — is the whole contract; skipping this makes every later step's test run red for an unrelated reason.
Verify
cmp scripts/build-plan-html.mjs kit/plugins/plan-agent/scripts/build-plan-html.mjs exits 0, and plan-agent-render --check invoked by bare name (which resolves through the plugin's bin/) behaves identically to the root module.
6
done Rewrite skills/build/references/completion-gates.md Step 5.3 as a single command — plan-agent-render "<stem>.md" -o "<stem>.html" --check — stating that a non-zero exit names the property and that the fix is always in the spec, never the HTML. Delete the five-selector paragraph. Keep sub-step 4's "fix the spec, never the HTML" rule and its no-promoting-status clause verbatim.
Why
the selector list is the instruction that produced the defect, and leaving it beside the new command lets a future run fall back to it.
Verify
grep -c 'step-card' kit/plugins/plan-agent/skills/build/references/completion-gates.md returns 0, and the file names the --check invocation exactly once.
7
done Apply the same replacement to skills/finalize-plan/references/write-completions.md, which runs the same completion rules for plans implemented outside build.
Why
completion-gates.md ends by instructing that the two stay consistent, so a one-sided change is a defect the next reader inherits.
Verify
both files reference --check and neither mentions .step-card.
8
done Extend tests/plugins/test-build-plan-html.mjs with the determinism case, the stale-HTML case, the missing-HTML case, the completed-with-unchecked-criterion case, and the in-progress skip case — importing from ../../scripts/build-plan-html.mjs as the file already does — while keeping every existing assertion green, the byte-identical parity assertion included.
Why
the gate is only worth what its failure modes are worth, each of these is a way --check could silently pass, and the parity assertion is what catches a Step 5 re-copy that was skipped.
Verify
node tests/plugins/test-build-plan-html.mjs reports zero failures with a higher assertion count than before the change, including the plugin-bundled renderer copies are byte-identical to the repo-root sources check.
9
done Bump plan-agent to 9.4.0 in .claude-plugin/marketplace.json, add the CHANGELOG entry, and document --check in the README's renderer section.
Why
the CI guard fails any PR whose touched plugin does not exceed the base branch version, and an undocumented flag is a flag the next session re-invents. Deviation from the plan as authored: it specified 9.3.0, which was correct when origin/main carried 9.2.0; PR #554 has since shipped 9.3.0, so 9.3.0 no longer exceeds the base and the guard rejects it as "version not bumped" (measured against the guard's own findViolations).
Verify
git fetch origin && BASE_REF=main node scripts/check-plugin-versions.mjs exits 0.

Tests

The tests that prove the change does what it promises.

Tier 1 — This plan changes application code
Objective: a plan whose HTML is stale, or whose completed spec has an unchecked criterion, fails the gate without any HTML being searched. File: tests/plugins/test-build-plan-html.mjs; Type: smoke; Asserts: --check exits 0 on a freshly rendered consistent plan, and non-zero with the offending property named for a stale render and for an unchecked criterion under status: completed; Run: node tests/plugins/test-build-plan-html.mjs
Unit: render determinism. File: tests/plugins/test-build-plan-html.mjs; Targets: renderPlanHtml; Key cases: two renders of one unchanged spec over the same output path are byte-identical, a re-render over an existing output file preserves plan-created, and rendering to a different path changes plan-file/plan-path and nothing else — pinning them as derived-from-path rather than volatile
Unit: check-mode reporting. File: tests/plugins/test-build-plan-html.mjs; Targets: the --check code path; Key cases: missing output file reports FAIL naming the render command rather than throwing, the first differing line is reported with its line number, and the table's rows appear in fixed order
Unit: consistency assertions gate on status. File: tests/plugins/test-build-plan-html.mjs; Targets: the spec-consistency branch; Key cases: status: in-progress with unchecked criteria exits 0, status: completed with an unchecked step exits non-zero, and a plan with no ## Acceptance Criteria section does not crash
Integration: existing render behaviour survives. File: tests/plugins/test-build-plan-html.mjs; Targets: the whole renderer; Key cases: every current assertion, including the extractSections(render(spec)) round-trip property

Definition of done

The plan counts as done when every statement below is true — check each one off as you verify it.

Final check

One last pass to confirm the whole change works end to end.

Run node tests/plugins/test-build-plan-html.mjs and confirm zero failures.
Then run git fetch origin && BASE_REF=main node scripts/check-plugin-versions.mjs
and confirm exit 0.

End-to-end, take a committed plan from docs/plans/ that is status: completed
and copy its .md and .html into the scratchpad. Run --check on the copy
and confirm exit 0. Append a stray character to the copied HTML and confirm the
check exits non-zero naming that line. Restore the HTML, flip one acceptance
criterion in the copied spec back to - [ ], and confirm the check exits
non-zero quoting that criterion — while the HTML row still passes, proving the
two properties are reported independently.

Finally, run build against a small plan end to end and confirm Step 5 issues
exactly one --check invocation and no Grep call against the plan HTML.

Wrapping up

Three gates that must all pass before this plan is marked completed.

Required

Completion Report

Step 9 version, authored as 9.3.0
origin/main moved to plan-agent 9.3.0 in PR #554 after this plan was written, so 9.3.0 no longer exceeds the base. Shipped 9.4.0 instead. Measured against the guard's own findViolations: 9.3.0 returns version not bumped, 9.4.0 passes.
Step 7 verify line, "neither mentions .step-card"
not met, and deliberately so. write-completions.md keeps two .step-card mentions: one describing what the renderer derives (an argument against HTML surgery) and one in legacy mode, the attribute-surgery path for plans that have no .md spec, where the selector is an edit target rather than drift evidence. Deleting them would break legacy finalization, which this plan's Scope does not cover. The acceptance criterion as written — neither file "names a CSS selector as evidence" — is met.
Verification section, e2e case 3 expectation that the html row still passes
not achievable as written. Flipping an acceptance criterion in the spec changes the rendered progress bar (10 / 10 done to 9 / 10 done), so the on-disk HTML is genuinely stale and the html row correctly fails alongside criteria. Independent reporting was verified instead on a re-rendered fixture, where html PASS appears beside steps FAIL and criteria FAIL.

Next steps

Follow-up ideas that came up along the way — none of them are required to finish this plan.

Wire the check into the plan-write hook

Paste this prompt into Claude to execute this follow-up:

In the agentics repo, kit/plugins/plan-agent/hooks/render-plan-html.py
re-renders a plan's HTML on every spec write. Now that
scripts/build-plan-html.mjs has a --check mode, evaluate whether the hook
should run it after rendering and surface a failure to the session, or
whether that duplicates the build skill's Step 5 gate for no benefit. Note
that plugin hooks do not register in the Claude Code desktop app, so the
hook cannot be the only place the check runs. If you recommend adding it,
bump the plan-agent version in .claude-plugin/marketplace.json, add a
CHANGELOG entry, and extend tests/plugins/test-build-plan-html.mjs. Verify
with `node tests/plugins/test-build-plan-html.mjs` reporting zero failures.