Guide · Claude Code review tooling
/code-review ultra 382 does, and what it didA walkthrough of the ultra review command and the max-effort pipeline that reviewed pull request #382 in this repo — ten independent finder angles, an adversarial verification pass, a gap sweep, and eleven findings, every one of them fixed on the branch.
One line, two very different execution paths.
/code-review ultra 382
Ultra is the top effort tier of the /code-review command. Its native form launches a multi-agent cloud review: the branch (or a GitHub PR, when you pass a number like 382) is shipped to a fleet of reviewer agents running remotely. It is user-triggered and billed — Claude cannot launch the cloud run on your behalf from inside a session. The older /ultrareview name is a deprecated alias for the same thing. It needs a git repository; the no-argument form bundles your local branch and doesn't even need a GitHub remote.
When the cloud run can't start — as in this session — the command falls back to a local max-effort review that reproduces the same philosophy with the Agent tool: many independent readers, then adversarial verification, then a sweep for what everyone missed. That fallback is what actually reviewed PR #382, and it's what the rest of this guide dissects.
Four phases: gather, fan out, verify, sweep.
The review scope is pinned down first: git diff origin/main...HEAD, exported to one file every agent reads. For PR #382 that was the humanized plan-skeleton change — an HTML template, a skill spec, an extractor library, a new smoke test, and version metadata.
Ten agents read the same diff in parallel, each through a different lens (listed below). They are deliberately blind to each other — if two angles flag the same line for different reasons, both records survive. Duplicates are a feature: five independent angles converging on one line is itself evidence.
Candidates pointing at the same mechanism are merged, keeping the most concrete failure scenario. Each survivor then gets one verification pass against the actual code — not the diff, the files — and lands in one of three states: CONFIRMED, PLAUSIBLE, or REFUTED. Only REFUTED kills a finding at this effort level.
A final agent gets the verified list and one job: find defects not on it. Fresh eyes plus the list of what's already known is a strong prior for what the first pass tends to miss. On #382 the sweeper found four new issues — including the best copy bug of the whole review.
set -e abort path in the new bash test and the heading-outline problem.The verifier can name the inputs or state that trigger the bug and quote the failing line. Ships in the report.
The mechanism is real but the trigger is uncertain — timing, environment, a generator's future behavior. Ships too: this is recall mode.
The code doesn't say that, or a guard elsewhere covers it — proven with a quoted line. Dropped.
Verification is where the pipeline earns its keep. Three plausible-sounding candidates on #382 died here: a claimed false-pass in the new test (refuted — the test's chunking bounds the search correctly), a CSS coupling complaint (refuted — the plan explicitly mandated that selector reuse), and a "missing sections" claim against the markdown skeleton (refuted — the gap predates the PR). All three would have read as legitimate findings in a single-pass review.
Eleven survivors, ranked most-severe first, and what happened to each.
| # | Finding | Outcome |
|---|---|---|
| 1 | Intro-stripping regex matched only one byte-exact tag form, so any attribute drift in a generated plan would leak presentation copy into the extracted spec — the exact pollution the change existed to prevent. Flagged independently by five angles.scripts/lib/plan-spec.mjs | Fixed |
| 2 | The skill spec still prescribed the old <summary>Verify</summary> markup its own skeleton had just humanized — a contract contradicting itself.implementation-plan/SKILL.md | Fixed |
| 3 | No defense against the one placement error the spec warns about most: an "At a glance" block nested inside the objective would silently pollute every extracted spec.scripts/lib/plan-spec.mjs | Fixed |
| 4 | The progress bar still said "Acceptance criteria" while the section it tracks was renamed "Definition of done" — one checkbox list, two names.reference/SKELETON.html | Fixed |
| 5 | "At a glance" rendered as the page's first <h2> while the Objective had no heading at all — inverting the outline for screen-reader users navigating by heading.reference/SKELETON.html | Fixed |
| 6 | For completed plans the collapsed drawer held only the File/Path rows, yet its label gave no hint a file path lived inside; File/Path also vanished from print.reference/SKELETON.html | Partial |
| 7 | Half-applied rename: sidebar links "Files" and "Diagram" no longer matched their renamed headings — a live instance of one heading existing as four uncoordinated copies.reference/SKELETON.html | Fixed + guard |
| 8 | Changelog described a "finalize" prompt that doesn't exist and overclaimed structural parity between the two skeleton formats.plan-agent/CHANGELOG.md | Fixed |
| 9 | Leftover CSS after the drawer regroup: unreachable print rules, inert declarations, and a bordered box nested inside a bordered box.reference/SKELETON.html | Partial |
| 10 | A new label style duplicated an existing rule in six of seven declarations, letting two adjacent labels drift apart on any future tweak.reference/SKELETON.html | Fixed |
| 11 | One placement rule stated in full three times across the skill spec, with wording already diverging between copies.implementation-plan/SKILL.md | Fixed |
The two partials are deliberate, not unfinished: finding 6's print behavior stays (a closed <details> can't be revealed by print CSS, and the path already prints in the footer), and finding 9 keeps two "dead" print rules because an existing contract test pins one of them byte-for-byte — annotated in the CSS so the next editor knows why.
test-goal-prompt.sh. The cleanup adapted to reality instead of breaking a passing test.Three effort tiers, one command.
# cloud multi-agent review of the current branch
/code-review ultra
# cloud review of a specific GitHub PR
/code-review ultra 382
# lighter local tiers — fewer, higher-confidence findings
/code-review # default effort
/code-review high # broader coverage, may include uncertain findings
The ultra cloud run must be started by you (it's billed); if it can't launch, the session runs the local max-effort pipeline described above. --fix applies findings to the working tree; --comment posts them as inline PR comments.
Rule of thumb for choosing a tier: default and high for everyday diffs, where a false positive costs more than it's worth; ultra (or the local max-effort fallback) when the change guards a contract other systems depend on — exactly the situation on #382, where one template feeds galleries, hooks, extractors, and a test suite that all grep it byte-for-byte.