Running three coding agents at once, by voice
Running one coding agent is a conversation. Running three is a job, and the job is mostly waiting: watching for the one that stopped to ask something, noticing the one that finished twenty minutes ago, and remembering what you asked the third one for in the first place.
Most voice tooling for coding agents solves the input half - you speak, it types. That is genuinely useful, and it is also the half that was never the bottleneck once you had more than one agent running. The bottleneck is that each agent is a window you have to look at to know anything about it.
The actual cost of a second agent
The first agent costs you nothing extra: you are looking at it while it works, because there is nothing else to look at. The second one starts the real tax, and it is not the typing. It is that the two of them now finish, stall and ask questions on their own schedules, and none of those events reach you unless you go and look.
So the loop becomes:
- Send an instruction to agent A.
- Switch to agent B, read enough of its output to remember where it was.
- Send an instruction to B.
- Switch back to A - which has been sitting on a yes/no question for four minutes.
Every one of those switches is a context reload, and the four idle minutes are pure loss. With three agents it stops being a nuisance and becomes the actual shape of the work. You are no longer programming; you are dispatching.
Step one: one place that knows about every session
The fix starts by making "which agent needs me" a thing you can see without switching to anything. ByteVoice keeps a small pill on the desktop, above whatever you are working in, carrying the state of every session it can see - Claude Code, Codex, Cursor, Gemini, Grok, Kimi, or a plain terminal.
Three colours carry the whole state machine, and they are worth learning because everything else here depends on reading them at a glance:
- Blue - working. The agent is mid-turn. Nothing is being asked of you.
- Yellow - needs you. It stopped. A permission prompt, a question, a decision. This is the only state that is costing you time right now.
- Green - done. The turn finished and there is a summary waiting.
The point is not the colours; it is that "is anyone blocked on me?" stops being a question you answer by visiting three windows. You answer it with a glance at one pill, and if nothing is yellow, you keep doing what you were doing.
Step two: get told, instead of looking
A glance is still an interruption if you have to keep taking it. The next step is not having to look at all: when a turn finishes or an agent stops to ask something, the summary comes to you.
Claude · Done
Added retry with backoff to the login request, wrapped the failure paths in clearer error toasts, and covered timeouts and 500s with new tests - all 42 tests pass.
Claude · needs you
The retry logic is in and unit tests pass. Before I push - should I also run the full integration suite, or is the login flow enough for now?
Free your hands - get notified when your agent finishes or has a question
Two details matter more than the notification itself. The first is that it can be read aloud, which is what makes it work while you are looking at something else entirely - the summary arrives in the one channel a second screen does not compete for. The second is that when an agent asks a question with defined options, those options arrive as buttons. You answer the question from the notification; you never open the session.
Permission prompts work the same way. When an agent wants to run something and your permission mode is set to ask, the request surfaces here, and "allow" or "deny" spoken aloud resolves it. The tool call that would have blocked for four minutes blocks for four seconds.
Step three: give your voice an address
Now the input side, which is where the multi-agent case diverges from plain dictation. Dictation types wherever your cursor is. That is exactly right for one agent and exactly wrong for three, because your cursor is usually in an editor, and the agent you want to answer is in a window you are not looking at.
Pin this voice shortcut
Press the pin below, then 🌐 Fn always sends your voice directly to this agent, even when ByteVoice is in the background.
Pinning a session is the difference between a microphone and a telephone. Once one is pinned, pressing fn anywhere - in your editor, in a browser, with nothing focused at all - sends what you say to that agent. You do not switch to it, click into its prompt, or even see it.
And because speech reaching something that edits code and runs commands deserves a check, nothing sends immediately. The transcript appears first with a short countdown; you can edit it, or press esc to cancel. Say the wrong thing and the fix costs a keystroke rather than a revert.
What a morning looks like
Concretely, with three agents on three different problems:
- Claude Code is pinned, working through error handling on the login flow. You are reading a diff in your editor.
- It stops to ask whether to also run the integration suite. The notification reads that question aloud. You say "run the full suite" without leaving the diff - or press the button in the notification, if your hands were already there.
- Codex, which has been migrating billing webhooks, goes green. Its summary is read aloud. Nothing is needed from you, so nothing interrupts you beyond one sentence.
- You think of a follow-up for Codex while Claude is still mid-turn. You switch the pin, press fn, say it. It queues behind the current turn instead of racing it, and goes out when that turn ends.
The switches did not disappear - there are still three problems in flight, and that is inherently more to hold in your head than one. What disappeared is the polling: the repeated, unprompted checking of windows that had nothing new in them. That is most of the tax, and it is the part that scales worst with the number of agents.
Where this does not help
Worth being straight about the limits. If you run exactly one agent and watch it work, most of this is machinery you do not need - dictation alone will do. If your agents rarely ask questions because you run them with broad permissions, the notification half matters much less. And none of it makes an agent's output correct; it makes the gap between the output existing and you knowing about it close to zero.
The case it is built for is the one where you have deliberately parallelised: several agents, several problems, and a suspicion that you are spending more time coordinating them than deciding anything.
Try it
ByteVoice is a macOS app. Setup is a Google sign-in, two permissions and a keyboard test - it checks each one actually works rather than telling you it should. It connects to Claude Code and Codex automatically once their CLIs are installed, and dictation works everywhere regardless, including in editors and terminals with no agent attached.
Download for macOS or see the plans - the free tier includes 5,000 dictated words and 1,500 spoken words a week, which is enough to find out whether the loop above is the one you have been missing.