Cadence
A voice-first AI agent Graham built around Claude. He talks to it all day: through a stenographer's mask in public, headphones at home. It acts on his machines, phone, browser, calendar, email, and creative pipelines, and talks back in short spoken turns.
1. The voice loop
Details
- End of turn is a 4-second silence detected by a streaming Silero VAD. He chose 4 s on purpose: room to think mid-sentence without being cut off.
- Barge-in: if he talks over the reply, the brain's turn is interrupted in-band without losing context. A backchannel filter ignores "mm-hm" and "yeah".
- Brain: a warm, long-lived
claude -pprocess in stream-JSON mode (Claude Opus 5.5, high effort). One process per session keeps context across turns. - Replies are chunked by sentence so speech starts before the full answer is written; markdown is stripped so nothing reads "asterisk asterisk."
- Mind vs. mouth: the transcript shows tool calls and work in progress; only the conversational reply is spoken.
- Voices: several TTS voices, auditioned and rated by Graham. Voice changes are a tool the brain can call.
- Fallback: if the GPU sidecar is down, the browser's built-in Web Speech takes over.
2. Why a stenographer's mask
A court-reporter stenomask seals around the mouth, so a microphone inside hears him clearly while people nearby barely hear anything. That makes it possible to talk to an AI in a cafe or a store without broadcasting the conversation. Audio comes back through an earbud. working
3. Network topology
- Everything runs on Graham's own hardware. Only the language model itself is a cloud API (Claude via his subscription).
- Machines talk over key-based SSH. GPU memory is the scarce resource, so heavy jobs are scheduled around it: the voice was moved to the small 2060 so the 3090 stays free for film and image generation.
- The phone is controllable too: an
adbbridge over the VPN lets Cadence open apps, start navigation, and read the screen.
4. Tools (what the brain can do)
| Tool / MCP server | What it does |
|---|---|
| Shell + files | Full Claude Code toolset on the main box: read, write, run, git, SSH to other machines. |
| Browser | Navigate, click, type, read, and screenshot in a real Chrome. Used for shopping, forms, research. |
| Co-drive browser working | A shared browser (neko, on 5.8) that both Graham and Cadence see live. Cadence's actions show up as its own drawn cursor, so he can watch and take over. |
| Vision | A local vision model looks at camera frames and keeps a timestamped journal. "What did you see at 4 pm?" pulls up that frame. parked |
| Image generation | Local diffusion on the 3090, including edits of camera frames; results shown in the transcript and a gallery. |
| Audiate (music) | Records his viola, checks intonation and fifths, compares takes, and plays rendered compositions back in his ear. |
| Control | Switch voices, restart the app, steer the co-drive. |
| Email + calendar | Command-line Gmail and Google Calendar. It confirms before sending anything outward. |
| Web search | For facts that change: hours, prices, news. |
| Local fallback model | An open-weights model served on the 3090 that takes over if the cloud model refuses a benign request. prototype |
5. Thinking in parallel
- Subagents: for anything long, the voice brain spins off a background agent (often a bigger creative model) and keeps talking. It's told when the agent finishes and relays a two-sentence summary.
- Detached jobs: renders that take hours run as standalone headless Claude processes with a written brief and a PROGRESS.md, so they survive the chat session ending or a usage limit and can be resumed.
- Cron pipelines: e.g. the fairy-tale film pipeline builds one film, publishes it, and starts the next on its own.
- Safety model: reversible actions run freely and are shown; irreversible or outward-facing ones (sending email, purchases, posting) get a quick spoken confirm.
6. Memory and specs
- Persistent memory: one small markdown file per fact (preferences, project state, decisions, lessons), plus an index loaded into every new session. Memories link to each other and get corrected when they turn out wrong.
- Spec-driven development: new ideas start with a "grill-me" interview, one spoken question at a time, until both sides share a design. Only then does it become a numbered spec (28 so far). The specs are the source of truth, and code follows them.
- Conversation style is itself a spec, grounded in voice-UX research: spoken attention fades after ~8–10 s, so 1–3 sentences, one question max, lead with the answer, no lists read aloud.
7. Creative pipelines
String quartets for a live sight-reading night
Graham and Cadence compose pieces for Classical Jams Austin, where strong sight-readers play a new piece each night. The flow: talk through the idea → LilyPond engraving → full score + parts PDFs → playability and sight-reading checks → rendered with real orchestral samples (sfizz + VPO on the lab box, never General MIDI) → played back in his ear for a verdict. Several pieces are ready for performance.
FilmForge
Narrated Grimm fairy tales: a TTS narrator reads the text, keyframes are generated per shot, animated with an open video model, kept character-consistent, scored, and published to videos.grahampaasch.com, a self-hosted video site. Being upgraded right now (see below).
Design rules Graham applies to everything
- Voice-first, no buttons: any feature has to work by talking.
- Handcuffs test: can Cadence do it on its own, with no human needed to trigger it?
- Touch-it test: is it open, ours, and modifiable? Third-party SaaS in the critical path is a liability.
8. A real afternoon (Oct 10, 2026)
- The printer was out of black ink. Cadence dug through his email, found an unused $50 Walmart gift card, and he picked up the cartridge for free, with $1 left over.
- It logs his daily routine anchors (wake, first contact, meals, sleep) as they come up in conversation. He couldn't remember when he was at Walmart, so it worked the time out from the conversation timestamps.
- He'd watched the latest generated film, Cinderella, and asked why the characters' shoes kept catching fire. A background agent traced it: the animation prompts said "the fire flickers" when there was no fire in the picture, so the video model set the brightest object alight, the golden slipper. It also found too few shots per minute of story and every shot framed like a portrait. A fix pass and an "A+ quality" research pass are running now, and the next film is held until they land.
- He asked for a fifth quartet: a string-quartet reimagining of Vaughan Williams' Fantasia on a Theme by Thomas Tallis, under 6 minutes, built around the two moments that move him most. That's composing in a detached job right now.
- He took a wrong turn on the way here. With the phone link down on cell data, it talked him back onto Lamar.
- And this page, written while he sat with a friend who wanted to know how it all works.
9. Honest status
Daily-driver solid: the voice loop, tools, memory, subagents, music pipeline, and co-drive browser. Rough edges: Bluetooth and the phone link are flaky, GPU memory juggling is constant, the film pipeline is mid-upgrade, and the proactive camera coach is parked. It's one person's system, held together by specs, tests, and a lot of conversation.
Generated by Cadence, Oct 10 2026. ← grahampaasch.com