Session Report · agentics

Per-skill model pinning across the plugin marketplace

How one question — “can my skills switch to Fable when invoked?” — became a verified, tiered model-routing policy across two plugins, a marketplace description rewrite, and a merged-ready PR, with every decision grounded in official docs or repo convention.

Date
2026-07-13
Branch
claude/skills-dynamic-model-switch-21faf0
Pull request
shawn-sandy/agentics #395
Status
Mergeable · smoke tests 15/15
16
Files changed
2
Plugins re-versioned
13
Model pins set or fixed
80%
Description size cut
01

Verify the mechanism before recommending

The request was to add a /model Fable step to skills. /model is an interactive CLI command that a skill body cannot invoke, so the first move was to establish what Claude Code actually supports — from documentation, not memory.

A claude-code-guide agent checked the official docs, and the two load-bearing pages were then fetched directly to confirm the wording first-hand:

model — Model to use when this skill is active. The override applies for the rest of the current turn and is not saved to settings; the session model resumes on your next prompt. Accepts the same values as /model, or inherit to keep the active model. A value excluded by your organization’s availableModels allowlist is not used and the session keeps its current model.

— Extend Claude with skills, frontmatter reference
Rationale

Three properties of this field shaped everything downstream. It is turn-scoped — no skill can permanently change the session model, so pinning is safe to apply liberally. It fails silently — an unrecognized value or an org allowlist exclusion falls back to the session model with no error, which makes exact spelling a correctness concern, not a style one. And it applies to the whole turn, so a pinned skill also governs whatever follows it in that turn — an argument for pinning down only on paths that never edit code.

One claim from the guide agent (a v2.1.196+ version requirement) could not be confirmed on the docs page itself and was reported to the user as unverified rather than repeated as fact.

02

The tiering heuristic

With the mechanism verified, model assignment followed one rule applied per component:

Principle

Pin cheap on the frequent, mechanical paths; pin high on the synthesis-heavy paths; never pin down the path that edits code. Commit messages happen ten times a day and a small model is indistinguishable there; autonomous CI autofix happens rarely and model quality is the entire safety margin.

TierAssigned toTest for membership
claude-fable-5Deep synthesisMulti-step reasoning where output quality compounds: full plan authoring, research-and-decide proposal loops
opusSynthesis and proseCombining many inputs into one judgment, or single-shot generation where quality shows: review synthesis, prompt crafting, prototypes
sonnetStructured judgmentOne narrow lens or outward-facing structured prose: single-dimension reviewers, PR bodies, issue tickets
haikuRigid-format, high-frequencyOutput shape is fixed and errors are cheap: branch names, conventional commit messages
inheritTrivial or dangerousEither too mechanical to matter (galleries, dispatchers) or too risky to downgrade (CI autofix that edits code)
03

plan-agent — 2.22.0 → 2.22.1

Two skills already carried pins (implementation-plan: fable, build-proposal: opus); the seven reviewer agents were already on sonnet. The pass corrected, extended, and aligned:

ComponentBeforeAfterWhy
implementation-planfableclaude-fable-5Full model ID replaces an alias the docs don’t list — silent fallback made the typo undetectable
build-proposalopusclaude-fable-5Same class of work as implementation-plan; same tier
review-plan—opusSynthesizes seven reviewers’ findings and applies them — the synthesis outranks the reviews
agent-review-plansonnetopusBackground twin of review-plan; foreground and background paths should behave identically
refine-prompt—opusPrompt quality is the whole deliverable
prototype—opusSingle-shot HTML generation where design quality shows
finalize-plan—sonnetMethodical evidence-checking, not creative work
plan-reviewer-* ×7sonnetsonnetOne narrow lens each, run seven-wide — the right cost point
plans-library / plans-open / setup-sites—inheritMechanical scan/open/scaffold; pinning adds nothing
04

git-agent — 3.12.0 → 4.0.1

ComponentBeforeAfterWhy
branch-agentHaikuhaikuCase fix — the docs list lowercase aliases, and an unrecognized value falls back silently
commit-agent + agent-commit— / sonnethaikuRigid conventional-commit format at the highest invocation frequency — the biggest cost win in either plugin
pr-agent—sonnetOutward-facing branch summaries; matches agent-pr’s existing pin
create-issue—sonnetStructuring a plan or session into a ticket is synthesis, not plumbing
ship / ship-autonomous—inheritDeliberate non-pin — see below
The deliberate non-decision

ship-autonomous was the one place a pin was explicitly refused. Its autofix step applies real code edits to fix failing CI, up to three attempts per check. Because a skill’s pin governs the whole turn, a cheap pin on the ship pipeline would also govern autonomous debugging — the single spot in either plugin where a weak model does damage. Restraint was recorded in the changelog as a decision, so a future pass doesn’t “complete” the pinning by mistake.

05

Marketplace description rewrite

plan-agent’s marketplace.json description had grown to 4,633 characters — each release had appended its changelog entry into it. It was rewritten to 924 characters.

Rationale

The constraint that shaped the rewrite: one clause per skill, so marketplace search still hits every component name, while per-flag behavior, menu mechanics, and release-note material moved to where they already lived — the README and CHANGELOG. The root cause (descriptions accreting changelog content) was flagged as a candidate rule for .claude/rules/marketplace.md rather than silently fixed once, since git-agent and social-media-tools show the same pattern.

06

Review triage: fix the real findings, ignore the noise

After the PR opened, three categories of bot feedback arrived and were handled differently on purpose.

Acted on — two Codex findings (both P2, both real)

Both threads got a one-line reply naming the fix commit and were resolved. The full smoke suite passed 15/15 before pushing.

Deliberately ignored — quota notices

Copilot posted three identical “quota limit reached” comments and CodeRabbit one rate-limit notice. None contained findings. Per the repo’s review-bot-loops rule, no commits were pushed to re-trigger them and no reply threads were opened — automated reviewers have no memory between runs, and each polish round can cost more than the change it reviews.

Rationale

The distinction applied throughout: a review fire (a bot re-running because a push happened) is not a review concern (a new blocking finding). Only the latter earns a commit.

07

Merge conflict: main moved underneath the PR

While the PR was open, main took git-agent to 4.0.0 (create-issue auto-activation, a breaking change) and added a new team-defaults plugin. Three files conflicted; each resolution preserved both sides’ intent:

08

Timeline

09

Principles this session ran on