Voice for Codex CLI: dictate and listen on macOS
Codex runs in a terminal, and that is the whole problem worth solving. Every instruction you give it is typed into the same window you are trying to read its output from, so composing the next prompt means scrolling away from the thing you were reading to decide what to say next.
Voice splits those apart. The output stays on screen where you can read it; the input arrives through a channel that is not competing for the same window - or the same hands.
Installing Codex CLI
The official standalone installer ships a self-contained binary, so there is no Node or npm to have first:
curl -fsSL https://chatgpt.com/codex/install.sh | shThen run codex once - the first run walks you through signing in with your ChatGPT account. Upgrading later is the same command; it replaces the binary in place.
Talking to it
ByteVoice works at the system level rather than as a Codex plugin, which for a terminal tool is the distinction that matters: there is no integration to configure, because dictation goes wherever your cursor is. Put the cursor in the terminal running Codex, hold fn, say what you want, and it is typed there.
Once Codex is detected, a second path opens that does not involve your cursor at all. ByteVoice reads Codex's session state directly, so each session becomes something it can address - you can pin one and speak to it from anywhere, with the terminal in the background.
readyMembership modal
Checkout flow wired to the pricing page - all 42 tests pass.
Press fn to speak
Approvals without switching back
Codex asks before it does things, and in a terminal that means a prompt sitting in a window you are not looking at, blocking a turn that was otherwise done thinking.
ByteVoice surfaces those requests as a notification wherever you are, and "allow" or "deny" spoken aloud resolves them. There is also a middle setting worth knowing about, because the default of asking for everything gets noisy fast: Risky only auto-approves ordinary reads, edits and safe shell commands, and still stops for the ones that are hard to undo - anything matching a destructive pattern like rm -rf, sudo, a force push, or a hard reset.
Hearing what happened
The input half is the obvious win; the output half is the one that changes how long you spend in the terminal. When a Codex turn finishes, its summary can be read aloud - which means you can start a long run and go and look at something else without that being a decision to stop knowing what it is doing.
You can pick the voice, the tone, and the language it speaks in - the last of which is worth setting deliberately if you work in a language other than the one you write code comments in.
Model and effort, without leaving the composer
Codex reads its model and reasoning effort from ~/.codex/config.toml. ByteVoice exposes both as controls next to the composer and writes that file for you, so switching from a cheap fast model to a slow careful one is a click rather than an editor session - and it lists only the efforts each model actually accepts, because Codex rejects an unsupported one at startup rather than clamping it.
One thing to know: a session already running keeps the model it started with. The config change applies to the next session, not the one in front of you.
Both agents at once
Most people running Codex are also running something else - Claude Code on a different problem, or Cursor in an editor. ByteVoice does not treat them as separate integrations: they are rows in the same list, with the same states and the same pin. Whether the pinned session is Codex or Claude Code changes nothing about how you talk to it.
That case has its own write-up: running three coding agents at once, by voice.
Getting started
ByteVoice needs macOS 15 or later, is signed and notarized by Apple, and sets up in about two minutes: sign in, grant Microphone and Accessibility, test the shortcut. Codex is detected automatically once its CLI is installed and signed in.