Team recap · 1 August 2026

The tool now stops and asks before it spends

Six changes to the plan-agent plugin. Two of them close a class of bug where the tool quietly decided something on your behalf, then spent real time acting on it.

Branch: claude/build-proposal-artifact-76a34c Pull request #508 Version 7.6.0 → 7.10.2

At a glance

6
Changes shipped
9
Files touched
6
Decisions recorded
7
Open items
24–36 → 0

Tool calls the proposal builder made before showing you anything, measured on the same request against the same 995-file repository. It used to restate your idea and start researching in one breath; now it stops and waits.

The session started from one request — always offer to publish the finished proposal — and the rest came out of actually running the tool and watching it misbehave. Two separate instructions turned out to be broken in ways that reading them would never have revealed. Alongside the fixes, the plugin's two largest instruction files were cut down so routine work stops paying to load detail it never reads.

What changed

Six changes, grouped by who notices them.

Affects everyone running build-proposal

It confirms your ask before researching it

The bug was real and observed. Given this request:

should we add a shared telemetry and usage-analytics layer across all 13 plugins so we can see which skills and commands actually fire in real sessions

the tool restated it as the same thing plus a goal nobody had mentioned:

… and let that data drive keep/merge/cut decisions.

It then researched against its own invention — a dispatched search agent, twenty-odd file searches, all before the human saw a word. Two rules now prevent this. Restate, do not enrich forbids adding a motive you never stated: if one seems missing, that is a clarifying question, not a blank to fill. A new checkpoint then presents the objective and waits for Looks right or Refine it, capped at two rounds.

Affects everyone finishing a proposal

Publishing is always offered, never assumed

Every finished proposal now gets one question: publish this as a shareable page, yes or no? Nothing goes out without an explicit yes.

The subtlety is what cannot turn the question off. A blanket “no further questions” covers the proposal's own decisions, not this one — publishing is the only step in the workflow you cannot undo by editing a file, so it keeps its own confirmation.

Affects everyone running write-prompt

You can name the prompt type directly

The first word of your request now sets the type, if it is one of system, task, creative, or analytical:

claude "/plan-agent:write-prompt creative a bedtime story about a lighthouse keeper"

The convention already existed, but as a private arrangement between two tools, documented for exactly one type. It is now a real feature for all of them, advertised in the command's own argument hint.

This matters because the type is not cosmetic — it picks both the questions you get asked and the techniques applied to the draft. Guess wrong and you answer a batch of wrong questions before the mismatch is visible. When no type is named and the classifier is confident, a new checkpoint asks before any of that starts.

Affects deep runs — you wait less

Research actually runs in parallel now

The instruction said to run the codebase search and the web research together. In practice the tool ran them one after the other, with 21 sequential file searches wedged in between.

The old wording prescribed one specific technique. The new wording permits either route to concurrency and instead forbids the single setting that broke it.

Affects teammates indirectly

The two biggest instruction files were cut down

An instruction file is paid in full every time its tool runs — there is no partial load — and a typical run reads maybe a quarter of it. Both of the plugin's largest files now keep only the rules in the core and push the mechanics into reference files that load on demand.

FileWasNowPlus
build-proposal380 lines349 lines1 reference
write-prompt434 lines343 lines3 references

How it works now

Three places where the shape of the work actually moved.

AFTER

Refine it - max 2 rounds

Looks right

Your request

Restate the objectiveno invented motives

Is this theobjective you meant?

Search agent + web research

BEFORE

Your request

Restate the objectiveand dispatch researchin the same message

Search agent + 20 file searches

First chance to correct it:24-36 tool calls in

The checkpoint sits immediately before the most expensive part of the run. Look at where the first tool call happens relative to the first question — that gap is the whole fix. Quick answers skip the checkpoint entirely, since they research nothing.

AFTER - concurrent

turn 1code sweepbackgrounded

turn 6web researchsweep still running

BEFORE - sequential

turn 5code sweepblocking

turns 6-2621 searches

turn 27web research

Same request, same repository. Read the turn numbers: the code sweep no longer blocks the web research, and the twenty-one searches that used to sit between them are gone.

yes

no

no

yes - NEW

/write-prompt ...

First word namesa type?

Type is settledno question asked

Is the classifierconfident?

Clarify menufour options

Looks right?Change the type?

Type-specific interview,then the draft

Four ways in, one exit. The branch marked NEW is the change: a confident classification used to be announced in passing and acted on. When another tool supplies the answers, every question here is skipped.

Before and after

BehaviourBeforeAfter
Restating your objectiveRestated and researched in one messageRestated, then waits for you
Inventing a motive you didn't stateHappened, uncheckedExplicitly forbidden
Tool calls before the first checkpoint24–360
Publishing a finished proposalSometimes offeredAlways offered, never assumed
Saying “no further questions”Suppressed the publish offer tooCovers proposal decisions only
Naming a prompt typeInferred from your prose; explicit for internal use onlyFirst word sets it, all four public types
A confidently wrong typeAnnounced in passing, then acted onConfirmed before the interview starts
Code sweep vs. web researchturn 5, then turn 27turn 1 and turn 6
Cost of a simple runWhole instruction file, every timeCore only; detail loads on demand

Decisions

Each with the option that was weighed and rejected, so the next person doesn't re-litigate it.

Check the output of framing, not whether the input looked vague

The existing rule fired when the input seemed underspecified. The observed failure had perfectly clear input that got confidently embellished, so the rule never fired. Checking the restatement instead catches both cases with one question, and needs no judgment call about how vague is vague enough.

RejectedTightening the existing “ask if underspecified” rule. It would not have fired on the case that actually broke.

The two checkpoints behave differently on purpose

When the interactive question tool is unavailable, one of them blocks and asks in plain text; the other proceeds and states its assumptions as a table you can correct in one reply. This looks inconsistent and is deliberate: whether blocking is right depends on what the next step costs. Blocking before a research sweep saves real money. Blocking before another question just deadlocks, because that question needs the same unavailable tool.

RejectedOne uniform rule for both. It was the first draft, and running it proved it wrong.

Forbid the setting that broke it, rather than prescribe the fix

The parallel-research instruction originally named one technique. The re-run reached concurrency by a different route and still passed. Prescribing only the original route would have been a rule the tool doesn't follow, describing a fix that wasn't the fix.

Publishing keeps its own yes

Everything else in the workflow writes files a person can edit or delete. Publishing sends a page outward. That asymmetry is why one blanket “stop asking me things” covers the rest of the run but not this.

The test suite set the boundary for what could move

Each test extracts a single step's body and asserts against it — by design, so that a rule cannot pass by living in the wrong step. Anything a test pins had to stay in the core. This is the existing principle — guards in the core, mechanics behind links — encoded as continuous integration, and it is why the directory-resolver script could move out but the precedence rules it implements could not.

One layer stayed in the core against that rule

In write-prompt, the layer that carries a proposal's evidence stays put while its seven siblings moved out. It passes tables and appendices through verbatim rather than shaping tone — behind a link it is a rule that may never load, and the prompt silently loses the evidence it exists to carry.

RejectedMoving all eight layers for consistency.

Learnings

Running the change caught two real bugs. Reading the diff would have caught neither.

The first publish-offer wording was broken

“Ask every run” read as an ordinary interview question, so a user saying “no further questions” suppressed it. The instruction was loaded and read — it just lost to a blanket directive. Only driving the tool end-to-end surfaced that.

The first version of the type checkpoint was unreachable

It said: if the question tool is unavailable, ask in plain text and wait. Two automated runs ignored it and proceeded — and they were right to. The next step needs the same tool, so blocking would strand the run waiting for something that can never arrive. The instruction now documents the behaviour the runs actually demonstrated.

An instruction describing behaviour the tool won't take is worse than no instruction. It reads as a guarantee during review and silently isn't one at runtime.

  • The file-trimming audit indicted the same session that prompted it — the two largest files in the plugin were the two this session had just added about 90 lines to.
  • A tool's own tests can be the safest refactoring guide available. Without them, more would have moved and the tools would have been quietly weakened.

Open items

Enough context on each to pick it up cold.

Files touched

Instruction files

kit/plugins/plan-agent/skills/build-proposal/SKILL.md The objective checkpoint, the no-enrichment rule, the publish offer, the concurrency fix. Trimmed 380 → 349 lines.
kit/plugins/plan-agent/skills/write-prompt/SKILL.md Leading type token, type checkpoint, non-interactive fallback. Trimmed 434 → 343 lines.

New reference files — loaded on demand

build-proposal/references/artifact-resolution.md The directory resolver, which runs once and only on runs that write something.
write-prompt/references/interview-questions.md The four type-specific question sets.
write-prompt/references/structuring-and-drafting.md Seven generic layers and the drafting rules.
write-prompt/references/saving-prompts.md Directory precedence and filename derivation.

Metadata

.claude-plugin/marketplace.json Version 7.6.0 → 7.10.2.
kit/plugins/plan-agent/CHANGELOG.md Six entries, one per change.
kit/plugins/plan-agent/README.md Documents the core-plus-references shape.

Also in this branch: a version collision, resolved

The main branch shipped a gallery redesign mid-session and claimed version 7.7.0, which this branch had also used. The branch's six entries were renumbered to sit above it — 7.7.0–7.9.2 became 7.8.0–7.10.2 — and the manifest bumped to match. Two different changes sharing one version number is a defect, not a cosmetic clash, so renumbering is the fix rather than picking a winner.

Glossary

Every term on this page a new teammate would otherwise have to ask about.

Skill
A Markdown instruction file that Claude Code loads to perform a task. The unit being edited throughout this session.
SKILL.md
The file itself. Its entire contents load every time the skill runs; there is no partial load.
Progressive disclosure
Splitting a skill into a small always-loaded core plus reference files that load only when a given run needs them.
Core and references
The two halves of that split.
Checkpoint (or gate)
A point where the tool stops and waits for a human answer before continuing.
Tier
How deep a proposal run goes. The shallowest answers directly and writes nothing; the deepest is a full research sweep.
Fan-out
Dispatching several research tasks at once.
Turn
One message in the exchange. Turn numbers are how the concurrency fix was measured.
Headless run
Claude Code invoked non-interactively. It cannot answer interactive questions, which is why the fallback paths matter.
--answers-gathered
A flag one tool passes another to say “decisions are already made, skip the interview.”
Merge driver
A script that auto-resolves a specific file's merge conflicts. This repository has one that keeps the higher version number.