Cadence

A voice-first AI agent Graham built around Claude. He talks to it all day: through a stenographer's mask in public, headphones at home. It acts on his machines, phone, browser, calendar, email, and creative pipelines, and talks back in short spoken turns.

1. The voice loop

Stenomask micPrivo, voice muffled Phone browsermic → 16 kHz PCM VPN → pfSenseinto home LAN Speech-to-textGPU sidecar + VAD Claude brainclaude -p + tools Sentence chunkersanitize for speech Text-to-speechKokoro, GPU Stream back24 kHz, gapless Ear+ live transcript Full duplex: the mic stays live while it talks, and Graham can interrupt (barge-in) mid-sentence.

Details

2. Why a stenographer's mask

A court-reporter stenomask seals around the mouth, so a microphone inside hears him clearly while people nearby barely hear anything. That makes it possible to talk to an AI in a cafe or a store without broadcasting the conversation. Audio comes back through an earbud. working

3. Network topology

Phone (Android)cell data, VPN client Laptop / any browsersame VPN pfSense firewallVPN endpoint, routing Main boxRTX 3090: brain host, vision,image + video generation Voice boxRTX 2060: speech-to-text +TTS; media archive Lab box24 cores, 94 GB RAM: musicrendering, co-drive browser, DB Home camerasTapo, RTSP → vision

4. Tools (what the brain can do)

Tool / MCP serverWhat it does
Shell + filesFull Claude Code toolset on the main box: read, write, run, git, SSH to other machines.
BrowserNavigate, click, type, read, and screenshot in a real Chrome. Used for shopping, forms, research.
Co-drive browser workingA shared browser (neko, on 5.8) that both Graham and Cadence see live. Cadence's actions show up as its own drawn cursor, so he can watch and take over.
VisionA local vision model looks at camera frames and keeps a timestamped journal. "What did you see at 4 pm?" pulls up that frame. parked
Image generationLocal diffusion on the 3090, including edits of camera frames; results shown in the transcript and a gallery.
Audiate (music)Records his viola, checks intonation and fifths, compares takes, and plays rendered compositions back in his ear.
ControlSwitch voices, restart the app, steer the co-drive.
Email + calendarCommand-line Gmail and Google Calendar. It confirms before sending anything outward.
Web searchFor facts that change: hours, prices, news.
Local fallback modelAn open-weights model served on the 3090 that takes over if the cloud model refuses a benign request. prototype

5. Thinking in parallel

6. Memory and specs

7. Creative pipelines

String quartets for a live sight-reading night

Graham and Cadence compose pieces for Classical Jams Austin, where strong sight-readers play a new piece each night. The flow: talk through the idea → LilyPond engraving → full score + parts PDFs → playability and sight-reading checks → rendered with real orchestral samples (sfizz + VPO on the lab box, never General MIDI) → played back in his ear for a verdict. Several pieces are ready for performance.

FilmForge

Narrated Grimm fairy tales: a TTS narrator reads the text, keyframes are generated per shot, animated with an open video model, kept character-consistent, scored, and published to videos.grahampaasch.com, a self-hosted video site. Being upgraded right now (see below).

Design rules Graham applies to everything

8. A real afternoon (Oct 10, 2026)

9. Honest status

Daily-driver solid: the voice loop, tools, memory, subagents, music pipeline, and co-drive browser. Rough edges: Bluetooth and the phone link are flaky, GPU memory juggling is constant, the film pipeline is mid-upgrade, and the proactive camera coach is parked. It's one person's system, held together by specs, tests, and a lot of conversation.

Generated by Cadence, Oct 10 2026. ← grahampaasch.com