Vocalize plans with a native Listen control

High todo
2026-06-29 agentics feature High effort

Give every generated plan a native browser Listen control, so the reader can hear a spoken summary (the Objective plus each step’s title) or the whole plan read aloud — powered entirely by the browser’s built-in speechSynthesis , with no dependency, no CDN, no network, and no audio files.

Implement Read and implement all steps in the plan at docs/plans/add-plan-tts-playback.md — Vocalize plans with a native Listen control. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-plan-tts-playback.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
More ways to run this plan — goal & workflow prompts, file path
Pursue as goal — optimize for the outcome, in parallel
Achieve this goal: Vocalize plans with a native Listen control. The plan at docs/plans/add-plan-tts-playback.md describes one approach — use it as reference, but optimize for the outcome. Fan out across parallel subagents where that serves the outcome. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-plan-tts-playback.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
Run as workflow — launch parallel subagents
Run a workflow to implement the plan at docs/plans/add-plan-tts-playback.md — Vocalize plans with a native Listen control. Brief subagents with the plan file at docs/plans/add-plan-tts-playback.md. Reserve a final verification phase for the lead agent, not a subagent. Verify against the plan's Tests, Verification, and Acceptance Criteria before reporting done. If everything passed, mark completion in docs/plans/add-plan-tts-playback.md — tick each step's [x] marker and each criterion's - [x], set status: completed — and re-render the HTML from the spec. If any check failed, leave status: in-progress and say which.
File add-plan-tts-playback.html
Path docs/plans/add-plan-tts-playback.html
Spec docs/plans/add-plan-tts-playback.md
Definition of done 0 / 7 done

Context

The story behind this plan — what prompted the work and why it matters now.

Plan output is a single self-contained .html file — “no external CSS, no CDN links, no external scripts”. The goal here is text→speech : let a reader press a button and listen to the plan, either a short summary or the full document. The browser already does this natively. window.speechSynthesis (the synthesis half of the Web Speech API) is supported in Chrome, Edge, Safari and Firefox, uses the OS/browser voices (mostly on-device), and needs no library, no CDN, and no network — so it honors the self-contained-HTML rule for free. The plan’s text is already in the DOM, so a Listen button simply gathers the relevant text and calls speechSynthesis.speak() ; pause() , resume() , and cancel() give Play / Pause / Stop with no extra machinery. Per the chosen approach, the “summary” is derived from the DOM (Objective + step titles) — zero new content and no change to how plans are authored; an authored narrative summary is a follow-up in Next Steps.

Files that change

Every file this plan touches, and what happens to each one.

agentics/
  • kit/plugins/plan-agent/
    • CHANGELOG.md modified dated 2.9.0 entry
    • README.md modified document the Listen feature
  • kit/plugins/plan-agent/skills/implementation-plan/SKILL.md modified document the always-present Listen control
  • kit/plugins/plan-agent/skills/implementation-plan/reference/SKELETON.html modified Listen control UI + CSS + inline TTS script
  • .claude-plugin/marketplace.json modified bump plan-agent 2.8.4 → 2.9.0
  • tests/plugins/test-tts-playback.sh new smoke test for the Listen control
  • path/to/file.ext modified what changes

Steps

The step-by-step work, in order — each step says what to do, why it matters, and how to check it worked.

1
todo Add the Listen control markup to SKELETON.html — In the header action row (next to Save as PDF ), add a .plan-tts group with two real <button> s — &ldquo;Listen to summary&rdquo; and &ldquo;Listen to full plan&rdquo; — plus a Stop button and an aria-live status span. Hide the whole group in @media print .
Why
Gives the reader an obvious entry point that sits with the other document-level actions and never bleeds into the printed/PDF output.
Verify
SKELETON.html has a .plan-tts group with both Listen buttons, a Stop button, and an aria-live status; the group is display:none under @media print .
2
todo Add the inline TTS script using window.speechSynthesis — Add a self-contained <script> that builds the summary text (the Objective text + every .step-chip-text ) and the full-plan text (walk the main sections in document order), then speaks via speechSynthesis.speak(new SpeechSynthesisUtterance(text)) . Wire Stop to cancel() and cancel on beforeunload .
Why
Native API means zero dependencies and self-contained playback; Play/Pause/Stop come straight from the API.
Verify
Opening a plan and clicking &ldquo;Listen to summary&rdquo; speaks the Objective followed by the step titles; &ldquo;Listen to full plan&rdquo; reads the sections; Stop halts speech immediately.
3
todo Scope the spoken text to human-facing prose only — When gathering text, include Objective, Context, each step&rsquo;s action / why / verify, Tests prose, Acceptance Criteria, Verification, and Next-Steps labels — and exclude the nav, the implement / goal / workflow prompt rows, copy buttons, raw code blocks meant for pasting, and the completion-checklist plumbing.
Why
Reading prompt blobs, JSON, and nav aloud is noise; the listener wants the plan&rsquo;s meaning, not its machinery.
Verify
Full-plan playback does not read the implement/goal/workflow prompts, the nav list, or <pre> / <code> copy blocks.
4
todo Add CSS for the control and a visible &ldquo;speaking&rdquo; state — Style the .plan-tts buttons to match the existing header buttons, and add an active/speaking treatment (e.g. the Stop button highlights and a small pulse) gated behind prefers-reduced-motion .
Why
The listener needs to see that playback is active and exactly how to stop it.
Verify
While speaking, the control shows an active/Stop state; with prefers-reduced-motion: reduce the pulse animation is disabled.
5
todo Make the control accessible and announced — Buttons are keyboard-operable real <button> elements with aria-label s; an aria-live="polite" region announces state changes (&ldquo;Reading summary&hellip;&rdquo;, &ldquo;Reading full plan&hellip;&rdquo;, &ldquo;Stopped&rdquo;). TTS playback must not interfere with a screen reader the user may already be running.
Why
It is an audio feature — it has to be operable without a mouse and announce what it is doing.
Verify
Tabbing to the control and pressing Enter starts/stops playback; the aria-live region announces start and stop.
6
todo Feature-detect and degrade gracefully — Guard the whole control behind if (!('speechSynthesis' in window)) hide the .plan-tts group , so unsupported browsers simply never see it. Handle the async voice list ( getVoices() may be empty until the voiceschanged event) by speaking with the system default voice.
Why
No broken buttons in browsers without support, and no race with the asynchronously-loaded voice list.
Verify
With speechSynthesis stubbed out / unavailable, the .plan-tts group is hidden and the rest of the plan is unaffected.
7
todo Document the Listen control in SKILL.md — Add it to HTML Output Requirements as an always-present header control (mirroring the Save as PDF entry): the markup, the speechSynthesis feature-detect/hide rule, what &ldquo;summary&rdquo; vs &ldquo;full plan&rdquo; read, and a &ldquo;do not remove when filling placeholders&rdquo; note.
Why
Every plan must ship the control, so the skill has to keep the markup + script the same way it keeps the Save-as-PDF button.
Verify
HTML Output Requirements lists the Listen control next to Save as PDF , including the feature-detect rule and the summary-vs-full behavior.
8
todo Bump the version and update docs — Bump plan-agent 2.8.4 &rarr; 2.9.0 (minor / feature) in .claude-plugin/marketplace.json , add a dated CHANGELOG.md entry, and document the Listen feature in the plugin README.md .
Why
Project convention: every plugin change bumps the marketplace version and updates the changelog and README in the same PR.
Verify
marketplace.json shows plan-agent at 2.9.0 , CHANGELOG.md has a 2026-06-29 entry, and README.md documents the Listen control.

Tests

The tests that prove the change does what it promises.

Tier 1 — Code-touching plan
Objective A reader can hear the plan (summary + full) via the native Listen control File: tests/plugins/test-tts-playback.sh Type: smoke test Asserts: SKELETON.html ships the .plan-tts Listen control (summary + full-plan buttons) wired to window.speechSynthesis , feature-detected and print-suppressed, and SKILL.md documents it as an always-present control — i.e. the reader can listen to the plan. Run: bash tests/plugins/test-tts-playback.sh
Unit Control is feature-detected and self-contained File: tests/plugins/test-tts-playback.sh Targets: the inline TTS script in SKELETON.html Key cases: grep confirms a 'speechSynthesis' in window guard and that the control introduces no src= /CDN/external script.
E2E Summary then full-plan playback in a browser File: tests/plugins/test-tts-playback.sh Targets: a generated plan opened in a headless browser Key cases: with speechSynthesis.speak stubbed, &ldquo;summary&rdquo; enqueues Objective + step titles, &ldquo;full plan&rdquo; enqueues the sections (no nav/prompts/code), and Stop calls cancel() .

Definition of done

The plan counts as done when every statement below is true — check each one off as you verify it.

Final check

One last pass to confirm the whole change works end to end.

Run bash tests/plugins/test-tts-playback.sh — it must pass, confirming the .plan-tts control and its speechSynthesis wiring exist and are documented. Then open a freshly generated plan in Chrome and Firefox : confirm the Listen control appears, &ldquo;summary&rdquo; speaks the Objective + step titles, &ldquo;full plan&rdquo; reads the sections in order (skipping nav, prompts, and code blocks), Pause and Stop work, and the control hides when speechSynthesis is unavailable and in print preview. Finally confirm marketplace.json shows plan-agent at 2.9.0 and the JSON still validates.

Wrapping up

Three gates that must all pass before this plan is marked completed.

Required

Completion Report

No items to report — all requirements met.

Next steps

Follow-up ideas that came up along the way — none of them are required to finish this plan.

Add an authored narrative summary for richer listening

Paste this prompt into Claude to execute this follow-up:

Extend plan-agent so each generated plan carries a 2-3 sentence narrative summary in a <meta name="plan-summary"> tag plus a visually-hidden element, authored at plan-creation time from the objective and steps. Update the SKELETON.html Listen control so 'Listen to summary' reads this authored summary when present and falls back to Objective + step titles when absent. Keep it native speechSynthesis with no dependencies. Document it in the plugin README and bump the plan-agent version.
Neural / cloud voices for nicer narration Wish List

Speculative / blue-sky idea — not on the critical path. Paste into Claude when ready to explore:

Paste this prompt into Claude to execute this follow-up:

Explore offering higher-quality neural TTS voices (a cloud TTS API or a WASM on-device model) as an optional upgrade over the browser's built-in speechSynthesis voices for plan playback. Weigh it against plan-agent's self-contained, no-CDN, no-dependency, no-network constraint -- a cloud call or external model would break it. Report whether any approach preserves self-containment; if not, recommend keeping native voices. Do not implement until the tradeoff is decided.