Team recap

Context engineering plans for the marketplace plugins

2026-07-27 branch: claude/context-engineering-review-2ecf6b commit: 999ac6d

At a glance

7Plans shipped
17Files touched
6Decisions
4Open items
0Plugins changed

Anthropic published new context-engineering guidance for Claude 5 models. We measured this repo against all six of its rules, found roughly 33,000 words of removable context cost across the plugins, and turned the findings into seven committed implementation plans.

Nothing has been implemented yet — every plan is status: todo. The numbers below are measured from the actual plugin tree, not estimated.

What changed

A measured context audit of all 13 plugins

Affects: plugin maintainers

We now have real numbers instead of impressions about where context is being wasted, and each number names the files responsible.

What we measuredResult
Skill definition files59 files / 81,367 words
Skills over 1,200 words shipping as one file17 files / ~28,000 words
Files repeating the same plan-mode preamble52 files / ~2,750 words
Identical lines shared by two recap commands168 lines / 1,568 words
Context loaded at the start of every session3,590 words
Hard constraints in one skill alone34 in 2,907 words

Findings live in the Context section of each plan under docs/plans/

Seven implementation plans, committed

Affects: whoever picks this up next

The audit is now executable. Each plan carries steps with per-step verification, falsifiable acceptance criteria, a runnable objective test, and an end-to-end verification section.

docs/plans/*.md (specs) + matching .html (rendered) — commit 999ac6d

A stale project rule flagged

Affects: contributors — not yet fixed

The marketplace contribution rules tell contributors "There is no CI version guard." That is false — the guard exists and runs on every pull request. Filed as a separate task rather than folded into this branch.

.claude/rules/marketplace.md line 49

How it works now

After: core plus references

common

rare

Skill triggers

Load core only (under 600 words)

Which path?

Done, nothing more loaded

Load one reference file, only when needed

Today: monolithic

Skill triggers

Load entire skill file (2,900 words)

Every branch loaded, even ones never taken

What splitting a skill actually changes. Today a skill's whole body loads the moment it triggers — there is no partial load. The concrete case: one skill carries a 60-line "author a plan first" branch that fires only when no plan is named, yet costs full tokens on every single invocation.

1. Replace repo plugin table (0 plugins bumped)

2. Extract recap core (1 plugin)

3. Remove plan-mode boilerplate (10 plugins)

4. Split plan-agent skills (1 plugin)

5. Split git-agent and social (2 plugins)

6. Split remaining skills (5 plugins)

7. Remove process imperatives (3 plugins, HIGH RISK)

Why the plans are sequenced this way. Order follows blast radius, not topic — this repo requires a version bump for every plugin touched, so a plan spanning ten plugins produces ten bumps in one commit. Plan 4 runs before 5 and 6 to prove the splitting pattern on one plugin first; plan 7 runs last because it removes safety constraints.

Before and after

RuleHow the plugins are written todayWhat the plans change it to
Where detailed guidance livesInside the skill body, loaded in full on every triggerIn reference files, loaded only when that step runs
Repeated instructionsThe same plan-mode preamble in 52 filesOne canonical line, documented once as the pattern
Three recap commandsEach restates the whole shared workflowOne shared core plus three short framing briefs
Repo-level contextA 13-row plugin table with paragraph-long rowsOne line per plugin, pointing at the generated table
Hard constraints34 in a single skill, mixing safety with process remindersSafety guards kept; process reminders dropped
Proof a change is safe"Looks fine"A committed baseline test that fails if behavior moves

Decisions

Split the work by blast radius, not by topic

The obvious grouping was one plan per recommendation.

Rejected because — this repo requires a version bump for every plugin touched, so topic-grouping produces unreviewable commits. Blast radius happens to correlate with risk here, so the cheap mechanical work also lands first.

Write all five recommendations as plans, including the risky one

Ships with a mandatory baseline-capture step and an explicit abort condition: if a pruned skill's baseline test fails and the cause is not obvious in one attempt, restore the constraint rather than debug.

Rejected alternative — dropping the imperative-pruning plan as poor risk-to-return. It was offered and declined.

Split the skill-splitting work into three plans by plugin

Three independently revertible units, grouped into near-equal thirds of 10,776 / 9,758 / 10,545 words.

Rejected alternative — one plan with a step per plugin. It avoids repeating the rationale three times, but produces a single all-or-nothing revert unit.

Reduce the plan-mode guard, do not remove it

The tutorial text goes; the one-line guard stays and is tested for.

Rejected reading — two of the article's rules point at deleting the preamble entirely. Rejected for write-heavy skills: a skill that starts writing inside plan mode violates a standing user preference, and that failure is silent.

Keep the smallest plan small rather than padding it

Plan 1 has 5 steps against 9 for plans 4 through 7. The work genuinely is small, and the repo's own right-sizing guidance says a chore should not carry a context essay it does not need.

Use a multi-agent workflow for the last four plans

Requested explicitly. Eight agents ran — one author plus one adversarial verifier per plan — in about nine minutes with no errors. The verify stage earned its cost.

Learnings

A verification command can be real, reproducible, and still meaningless

A first-draft plan checked a directory count with a command that emits directory headers and blank lines alongside the entries. It printed 14 for a directory holding 4 items. The number is stable, reproducible, and entirely unrelated to what the step claimed to measure — it would have passed review and execution.

This is the strongest argument for the baseline-first sequencing in plan 7: a plan full of plausible-looking verification is exactly the condition under which removing safety constraints goes wrong quietly.

Adversarial verification caught defects, not polish

Across four plans the verifiers found a manifest path that does not exist, a misquoted project rule, a claim that a change "breaks nine assertions" in a test where only three checks touch the moved content, and line citations past the end of a 204-line file.

Our conventions are enforced by tooling, not by prose

A plan file named with the wrong verb was rejected instantly by an automated filename check. That is the article's tool-design rule in miniature — a well-designed interface constrains behavior better than an instruction does. It also suggests pruning constraints is safer than it looks, but only where a hook or test already covers the same ground.

Tried and abandoned: a naive "do all cited paths exist?" check

It flagged 60 missing files across four plans. All but one were files the plans propose to create, correctly declared as new in their file lists; the last was the checker's own pattern matching a fragment of a longer valid path. The check was the bug.

Open items

  • Nothing is implemented. All seven plans are status: todo. The commit adds planning documents only — no plugin, skill, or config file changed.
  • Plan 7 may be worth dropping. It has the smallest return (5 skills) and the highest regression risk of the set. If the baseline harness proves more expensive than the saving it protects, abandoning it is a legitimate outcome — this is stated inside the plan itself.
  • Plans 4 through 6 run about 3x the length of plans 1 through 3 (3,228–4,048 words against 1,035–1,367). Partly justified by covering more skills, but they are also more granular than the repo's right-sizing guidance calls for. Not a defect; worth knowing before reading them side by side.
  • The stale CI-guard claim in the marketplace rules is unfixed, filed as a separate task rather than folded into this branch.

Files touched

Seventeen files. Fourteen are the plan specs and their rendered pages under docs/plans/; the rest are this recap, its session record, and the two rebuilt gallery indexes. No plugin source changed.

Plan specWhat it does
replace-claude-md-plugin-table.mdCut the repo's always-loaded plugin table
extract-recap-command-core.mdDeduplicate the three recap commands
remove-exitplanmode-boilerplate.mdReduce the repeated plan-mode preamble across 52 files
split-plan-agent-skills.mdSplit plan-agent's five monolithic skills
split-git-social-skills.mdSplit six skills in git-agent and social-media-tools
split-remaining-plugin-skills.mdSplit six skills across five remaining plugins
remove-skill-process-imperatives.mdPrune process constraints, baselines captured first

Each spec has a matching .html rendered by the repo's plan generator. The spec is the source of truth and the file to edit; the rendered page is generated and never hand-edited.

Glossary

Progressive disclosure
Loading detail only at the moment it is needed, rather than everything up front. The central idea behind the splitting plans.
Skill file
The instruction file defining a skill. Its entire contents load whenever the skill activates — there is no partial load.
Monolithic skill
A skill shipping as a single file with no supporting reference files, so every branch costs tokens on every invocation.
Blast radius
How many plugins a change touches, and therefore how many version bumps and how large a revert it implies.
Version bump
Raising a plugin's version number in the marketplace registry so installers receive the change. Required for any plugin edit.
Always-loaded context
Files read at the start of every session, paid for whether or not they turn out to be relevant.
Tautology check
Deliberately breaking the thing a test protects, confirming the test fails, then reverting. Proves the test detects regressions rather than passing unconditionally.
Objective test
A runnable check that asserts a plan's stated goal was actually accomplished, distinct from the prose verification written into the plan.
Republish key
A URL stored in a session record so re-running a command updates the existing published page instead of minting a new link.