Dynamic depth and execution mode for write-prompt

Medium todo
2026-07-01 agentics feature Medium effort

Teach write-prompt to right-size its output — plain-prose prompts for simple asks, full XML scaffolding for complex ones — and to classify how each prompt should run: one-shot, loop, workflow, or goal, delivered ready to paste in that mode.

Implement Read and implement all steps in the plan at docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md — Dynamic depth and execution mode for write-prompt. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
More ways to run this plan — goal & workflow prompts, file path
Pursue as goal — optimize for the outcome, in parallel
Achieve this goal: Dynamic depth and execution mode for write-prompt. The plan at docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md describes one approach — use it as reference, but optimize for the outcome. Fan out across parallel subagents where that serves the outcome. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
Run as workflow — launch parallel subagents
Run a workflow to implement the plan at docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md — Dynamic depth and execution mode for write-prompt. Brief subagents with the plan file at docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md. Reserve a final verification phase for the lead agent, not a subagent. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
File add-dynamic-depth-and-mode-to-refine-prompt.html
Path docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.html
Spec docs/plans/add-dynamic-depth-and-mode-to-refine-prompt.md
Definition of done 0 / 7 done

Context

The story behind this plan — what prompted the work and why it matters now.

The write-prompt skill ( kit/plugins/plan-agent/skills/write-prompt/SKILL.md ) runs a fixed pipeline: classify into one of four prompt types, run a batched interview, apply the full XML technique matrix for that type, and fill a template from references/ . There is no depth axis anywhere — a trivial ask like "summarize this doc in 3 bullets" gets <context> , <example> , and <thinking> scaffolding it does not need, because the technique matrix is keyed only on type , never on complexity . The skill has an escalation valve ("go deeper?") but no de-escalation valve.

Separately, the skill never considers how the finished prompt will run. A prompt meant to poll CI every ten minutes needs an interval and a /loop wrapper; a repo-wide audit prompt benefits from workflow orchestration; an autonomous objective needs a stated completion condition. Today all of these are delivered as bare one-shot text. An investigation on 2026-07-01 confirmed both gaps are fixable with edits to a single SKILL.md — the four templates stay untouched. Interview decisions locked in: the lite path always asks exactly one batched clarifying question; mode is judged by run-intent (not raw keywords) and announced only when non-one-shot, with a symmetric "run as one-shot" override; depth and mode are independent axes; existing saved prompts are not backfilled (that is a follow-up).

Files that change

Every file this plan touches, and what happens to each one.

agentics/
  • .claude-plugin/marketplace.json modified plan-agent 3.0.0 → 3.1.0
  • kit/plugins/plan-agent/
    • CHANGELOG.md modified 3.1.0 entry for depth + mode
    • README.md modified document lite/full and execution modes
  • kit/plugins/plan-agent/skills/write-prompt/SKILL.md modified depth axis, lite path, mode phase, delivery + save
  • tests/plugins/test-write-prompt-dynamic.sh new smoke test for dynamic-output markers

Steps

The step-by-step work, in order — each step says what to do, why it matters, and how to check it worked.

1
todo Add the depth axis to Phase 1 (Classify) in kit/plugins/plan-agent/skills/write-prompt/SKILL.md — after the type table, add a Depth rubric classifying lite vs full (lite: single clear action, no documents/examples/persona/edge-case handling needed, obvious output shape; full: everything else, or the user asks). Extend the one-line announcement to include the depth verdict and the escape hatch: "say 'go full' for the structured version" (and "go lite" for the reverse).
Why
The pipeline is keyed only on prompt type, so trivial asks get the full XML treatment; a depth verdict at classification time is the single routing decision everything downstream hangs off, and the announced escape hatch makes a misclassification cost the user two words.
Verify
Re-read Phase 1 of SKILL.md: it names both axes (type + depth), lists the lite/full signals, and the announcement example shows the depth verdict with the "go full" / "go lite" escape hatch.
2
todo Add the lite path through Phases 2–4 — a short "Lite path" subsection stating: when depth is lite , ask exactly one batched AskUserQuestion confirming intent and output shape (never the full type-specific battery, never a second round), skip Phase 3 XML layers and Phase 4 templates entirely, and draft a 1–3 paragraph plain-prose prompt applying only clarity/directness, positive framing, and lead-with-the-most-important-instruction. Phases 5–7 run unchanged; the existing "go deeper?" offer remains available after a lite draft.
Why
Skipping structure is not skipping quality — the lite path keeps the core Anthropic writing rules while dropping the scaffolding that buries a simple instruction, and the single mandatory question (per the align decision) guards against drafting from a misread intent.
Verify
Read SKILL.md end-to-end: the lite path bypasses templates and the type-specific interview batteries while asking exactly one question; the full-path phase text is untouched; the escalation offer ("go deeper?") still appears after lite delivery.
3
todo Add execution-mode classification as Phase 1.5 — classify one-shot | loop | workflow | goal with the rule stated explicitly: judge by how the user intends the prompt to run , never by keywords appearing in task content (a "for loop" refactor is one-shot). Include the signal table (recurring/monitoring intent → loop; fan-out/audit-everything intent → workflow; autonomous/until-done/background intent → goal; default one-shot). Depth and mode are independent axes — a lite prompt can be a loop. For loop add an interval question and for goal a stop-condition question to the Phase 2 batch (or to the single lite question). Announce the verdict only when non-one-shot, ending with the override: "say 'run as one-shot' to skip the wrapping". When both the Phase 1 depth verdict and the Phase 1.5 mode verdict are non-default (e.g. lite + loop), combine them into the single Phase 1.5 announcement rather than emitting two separate one-liners — one combined sentence naming depth, mode, and both escape hatches.
Why
A prompt that runs repeatedly or autonomously needs an interval or completion condition baked in before drafting — bolting the mode on after Phase 4 produces a one-shot prompt wearing a loop costume; intent-level judgment kills the keyword false-positive class, and silence on one-shot keeps the common case noise-free. Without an explicit merge rule, a lite+loop prompt would surface two back-to-back announcements, which reads as noisier than the single-axis case the design is meant to stay quiet for.
Verify
SKILL.md shows Phase 1.5 between Classify and Interview with the four-mode table, the run-intent (not keyword) rule, the depth-independence note, the per-mode extra interview questions, the announce-only-when-non-one-shot behaviour with the "run as one-shot" override, and an explicit example showing the combined depth+mode announcement when both are non-default.
4
todo Extend delivery and save with mode output — Phase 5 (Recommend) gains a one-line mode verdict alongside tool recommendations; Phase 6 (Deliver) adds mode-specific wrapping to the format block: loop → the prompt pre-wrapped as /loop <interval> <prompt> ; workflow → prefixed with "Run a workflow to …"; goal → framed for a background agent with its completion condition stated; one-shot → unchanged. Phase 7 (Save) adds mode: and depth: fields to the saved file's YAML frontmatter (new files only — no backfill of existing docs/prompts/ files).
Why
The mode verdict is only useful if the delivered artifact is ready to paste in that mode, and mode/depth in frontmatter makes the saved prompt library self-describing and filterable later without a migration.
Verify
Phase 6's format block shows all four mode variants (with the loop example pre-wrapped in /loop ); Phase 7's file-content template lists mode: and depth: in the frontmatter alongside type , intent , techniques , created .
5
todo Write the smoke test tests/plugins/test-write-prompt-dynamic.sh — a bash script following the pattern of tests/plugins/test-goal-prompt.sh , asserting SKILL.md contains: the depth rubric markers ( lite and full in Phase 1), the escape-hatch phrase, the Phase 1.5 mode table with all four modes and the run-intent rule, the /loop wrapping instruction in Phase 6, and the mode: / depth: fields in the Phase 7 frontmatter template.
Why
These are behaviour-bearing markdown instructions with no runtime of their own — a grep-based smoke test is the cheapest regression guard against a future edit silently dropping the routing (matching the align decision that a live eval is out of scope).
Verify
bash tests/plugins/test-write-prompt-dynamic.sh exits 0 against the edited SKILL.md, and exits non-zero when pointed at the pre-change SKILL.md (spot-check one assertion by temporarily reverting a marker).
6
todo Bump the version and update docs — set plan-agent to 3.1.0 in .claude-plugin/marketplace.json (minor: new behaviour, no invocation change), add a 3.1.0 entry to kit/plugins/plan-agent/CHANGELOG.md , and refresh the write-prompt section of kit/plugins/plan-agent/README.md to describe depth-aware output and execution-mode classification. Leave the CLAUDE.md plugin-table line untouched — its description remains accurate.
Why
The marketplace version must exceed the value on main for the change to ship, and README/CHANGELOG are the project's contract for plugin changes; the CLAUDE.md line does not mention output shape, so editing it would be churn.
Verify
marketplace.json parses (the settings hook auto-validates on edit) and reads 3.1.0 for plan-agent; CHANGELOG.md has the 3.1.0 entry; README's write-prompt section names both behaviours; git diff shows no CLAUDE.md change.

Tests

The tests that prove the change does what it promises.

Tier 2 — Non-code plan (markdown skill instructions, metadata, and a test script; no application source files)
Objective write-prompt SKILL.md encodes dynamic depth and execution-mode routing File: tests/plugins/test-write-prompt-dynamic.sh Type: smoke test (grep assertions, matching the test-goal-prompt.sh pattern) Asserts: the shipped SKILL.md actually carries the plan's objective — the Phase 1 depth rubric ( lite / full ) with the "go full" escape hatch, the Phase 1.5 mode table (one-shot / loop / workflow / goal) with the run-intent rule and "run as one-shot" override, the /loop pre-wrapping in Phase 6, and the mode: / depth: fields in the Phase 7 save template. Run: bash tests/plugins/test-write-prompt-dynamic.sh

Definition of done

The plan counts as done when every statement below is true — check each one off as you verify it.

Final check

One last pass to confirm the whole change works end to end.

Run bash tests/plugins/test-write-prompt-dynamic.sh and confirm exit 0. Then dry-run the routing by reading the edited SKILL.md top to bottom against two simulated invocations: /plan-agent:write-prompt summarize this doc in 3 bullets must route lite/one-shot (one clarifying question, prose draft, no template read, no mode announcement), and /plan-agent:write-prompt watch CI every 10 minutes and report failures must route loop (mode announced with the "run as one-shot" override, interval question added, delivery pre-wrapped in /loop , frontmatter carrying mode: loop ). Finally confirm .claude-plugin/marketplace.json parses with plan-agent at 3.1.0 (higher than main's 3.0.0), the CHANGELOG entry exists, and the four template files under references/ are byte-identical to before.

Wrapping up

Three gates that must all pass before this plan is marked completed.

Required

Completion Report

No items to report — all requirements met.

Next steps

Follow-up ideas that came up along the way — none of them are required to finish this plan.

Backfill mode and depth frontmatter into existing saved prompts

Paste this prompt into Claude to execute this follow-up:

Scan every saved prompt file under docs/prompts/ in the agentics repo. For each file whose YAML frontmatter lacks a `mode:` or `depth:` field, infer the values from the prompt content — mode is one-shot, loop, workflow, or goal, judged by how the prompt is intended to run (not by keywords in its task content); depth is lite for plain-prose prompts and full for XML-structured ones. Add only the missing fields, change nothing else in the frontmatter or body, and report a table of files changed with the values assigned plus any files where the inference was ambiguous.
Prompt library gallery with mode and depth filters Wish List

Speculative / blue-sky idea — not on the critical path. Paste into Claude when ready to explore:

Paste this prompt into Claude to execute this follow-up:

Run a workflow to build a filterable HTML gallery for the saved prompts in docs/prompts/ in the agentics repo, mirroring the plans gallery pattern (docs/plans/index.html built by the plan-agent plans-library skill): parse each prompt file's YAML frontmatter (type, mode, depth, created), render cards with colour-coded chips, add filter chip rows for type, mode, and depth, and wire a PostToolUse hook that rebuilds the gallery whenever a prompt file changes, following the rebuild-plans-index hook pattern in kit/plugins/plan-agent. Brief subagents with the existing gallery implementation as the reference.