Guide · Claude Code review tooling

What /code-review ultra 382 does, and what it did

A walkthrough of the ultra review command and the max-effort pipeline that reviewed pull request #382 in this repo — ten independent finder angles, an adversarial verification pass, a gap sweep, and eleven findings, every one of them fixed on the branch.

11finder agents
27raw candidates
3refuted & dropped
11findings shipped
11 / 11fixed on branch

The command

One line, two very different execution paths.

/code-review ultra 382

Ultra is the top effort tier of the /code-review command. Its native form launches a multi-agent cloud review: the branch (or a GitHub PR, when you pass a number like 382) is shipped to a fleet of reviewer agents running remotely. It is user-triggered and billed — Claude cannot launch the cloud run on your behalf from inside a session. The older /ultrareview name is a deprecated alias for the same thing. It needs a git repository; the no-argument form bundles your local branch and doesn't even need a GitHub remote.

When the cloud run can't start — as in this session — the command falls back to a local max-effort review that reproduces the same philosophy with the Agent tool: many independent readers, then adversarial verification, then a sweep for what everyone missed. That fallback is what actually reviewed PR #382, and it's what the rest of this guide dissects.

The one design idea to remember: at max effort the review optimizes for recall — catching every real bug — not precision. A missed bug ships; a false positive costs one verification. So the pipeline generates aggressively and filters afterwards, instead of asking one careful reader to be right the first time.

The pipeline

Four phases: gather, fan out, verify, sweep.

PHASE 0Gather the diff853 lines · 7 files

The review scope is pinned down first: git diff origin/main...HEAD, exported to one file every agent reads. For PR #382 that was the humanized plan-skeleton change — an HTML template, a skill spec, an extractor library, a new smoke test, and version metadata.

PHASE 1Fan out 10 finder anglesup to 8 candidates each

Ten agents read the same diff in parallel, each through a different lens (listed below). They are deliberately blind to each other — if two angles flag the same line for different reasons, both records survive. Duplicates are a feature: five independent angles converging on one line is itself evidence.

PHASE 2Dedupe, then verify each survivor1 vote · 3 states

Candidates pointing at the same mechanism are merged, keeping the most concrete failure scenario. Each survivor then gets one verification pass against the actual code — not the diff, the files — and lands in one of three states: CONFIRMED, PLAUSIBLE, or REFUTED. Only REFUTED kills a finding at this effort level.

PHASE 3Sweep for gaps1 fresh reviewer

A final agent gets the verified list and one job: find defects not on it. Fresh eyes plus the list of what's already known is a strong prior for what the first pass tends to miss. On #382 the sweeper found four new issues — including the best copy bug of the whole review.

The ten angles, and what each one caught on #382

The three verdicts

CONFIRMED

The verifier can name the inputs or state that trigger the bug and quote the failing line. Ships in the report.

PLAUSIBLE

The mechanism is real but the trigger is uncertain — timing, environment, a generator's future behavior. Ships too: this is recall mode.

REFUTED

The code doesn't say that, or a guard elsewhere covers it — proven with a quoted line. Dropped.

Verification is where the pipeline earns its keep. Three plausible-sounding candidates on #382 died here: a claimed false-pass in the new test (refuted — the test's chunking bounds the search correctly), a CSS coupling complaint (refuted — the plan explicitly mandated that selector reuse), and a "missing sections" claim against the markdown skeleton (refuted — the gap predates the PR). All three would have read as legitimate findings in a single-pass review.

The findings ledger — PR #382

Eleven survivors, ranked most-severe first, and what happened to each.

#FindingOutcome
1Intro-stripping regex matched only one byte-exact tag form, so any attribute drift in a generated plan would leak presentation copy into the extracted spec — the exact pollution the change existed to prevent. Flagged independently by five angles.scripts/lib/plan-spec.mjsFixed
2The skill spec still prescribed the old <summary>Verify</summary> markup its own skeleton had just humanized — a contract contradicting itself.implementation-plan/SKILL.mdFixed
3No defense against the one placement error the spec warns about most: an "At a glance" block nested inside the objective would silently pollute every extracted spec.scripts/lib/plan-spec.mjsFixed
4The progress bar still said "Acceptance criteria" while the section it tracks was renamed "Definition of done" — one checkbox list, two names.reference/SKELETON.htmlFixed
5"At a glance" rendered as the page's first <h2> while the Objective had no heading at all — inverting the outline for screen-reader users navigating by heading.reference/SKELETON.htmlFixed
6For completed plans the collapsed drawer held only the File/Path rows, yet its label gave no hint a file path lived inside; File/Path also vanished from print.reference/SKELETON.htmlPartial
7Half-applied rename: sidebar links "Files" and "Diagram" no longer matched their renamed headings — a live instance of one heading existing as four uncoordinated copies.reference/SKELETON.htmlFixed + guard
8Changelog described a "finalize" prompt that doesn't exist and overclaimed structural parity between the two skeleton formats.plan-agent/CHANGELOG.mdFixed
9Leftover CSS after the drawer regroup: unreachable print rules, inert declarations, and a bordered box nested inside a bordered box.reference/SKELETON.htmlPartial
10A new label style duplicated an existing rule in six of seven declarations, letting two adjacent labels drift apart on any future tweak.reference/SKELETON.htmlFixed
11One placement rule stated in full three times across the skill spec, with wording already diverging between copies.implementation-plan/SKILL.mdFixed

The two partials are deliberate, not unfinished: finding 6's print behavior stays (a closed <details> can't be revealed by print CSS, and the path already prints in the footer), and finding 9 keeps two "dead" print rules because an existing contract test pins one of them byte-for-byte — annotated in the CSS so the next editor knows why.

Fixes worth a closer look

Running it yourself

Three effort tiers, one command.

# cloud multi-agent review of the current branch
/code-review ultra

# cloud review of a specific GitHub PR
/code-review ultra 382

# lighter local tiers — fewer, higher-confidence findings
/code-review            # default effort
/code-review high       # broader coverage, may include uncertain findings

The ultra cloud run must be started by you (it's billed); if it can't launch, the session runs the local max-effort pipeline described above. --fix applies findings to the working tree; --comment posts them as inline PR comments.

Rule of thumb for choosing a tier: default and high for everyday diffs, where a false positive costs more than it's worth; ultra (or the local max-effort fallback) when the change guards a contract other systems depend on — exactly the situation on #382, where one template feeds galleries, hooks, extractors, and a test suite that all grep it byte-for-byte.