Codex provides Voice.
OpenHands does the work.
The current bridge forwards an explicit Voice handoff directly to the saved OpenHands Cat. It does not wait for a second reasoning model to choose an integration tool. Codex handles the experimental authenticated audio connection; the saved Cat owns tools, approvals, history, execution, and results.
Current path: one bound Cat, direct handoff
Agent Server creates an ephemeral Codex thread for each call and binds the call to one existing OpenHands Insider conversation. The browser negotiates WebRTC through the server, keeps its audio connection across Canvas navigation, and polls call status. A fresh call starts with bounded saved Cat context. The selected worker in the Projects board is not included in that Codex session.
Codex Voice: emit thread/realtime/itemAdded with item.type = "handoff_request", handoff_id, and explicit input_transcript.
Agent Server: validate and admit the handoff once within that call, then append the request to the bound Cat with send_message(run=True).
OpenHands: execute that saved Cat using its configured tools and policy. The bridge waits for the entire run, including stop hooks, to settle.
Agent Server → Voice: select the new saved answer after the pre-request event watermark and send it through thread/realtime/appendSpeech. A proposed finish alone is not sufficient.
The Codex thread uses dynamicTools: [] alongside restrictions on other tool sources. The pinned app-server can still create a background turn: clientManagedHandoffs suppresses automatic outbound answers, not incoming background work. That turn has no tools and neither dispatches requests nor supplies the saved OpenHands answer. This restriction is essential; forwarding a notification while leaving an independently capable Codex agent active would have two execution owners.
The browser does not resubmit Codex transcripts or tool events. The separate public OpenAI Realtime API path still uses a browser send_to_insider tool bridge. Its events and credentials are different. Public GPT-Live delegation metadata is also a different protocol: do not substitute it for the pinned Codex input_transcript contract.
There is still a Voice-model boundary. Live must emit a handoff. Direct speech from initial history is not proof that a request was saved in OpenHands. The new bridge removes the additional tool-choice failure after a handoff arrives; it does not guarantee that every spoken utterance will produce a handoff.
Published implementation: Codex Voice broker and handoff and lifecycle tests. Equivalent tested worktree change: 957c1c584.
The Cat uses normal OpenHands tools and skills
The real worker is the saved OpenHands Cat. A September 20 inspection of the Cat that could not count conversations found an earlier bare fixture with an empty configured tool list and no client tools. Its reasoning and finish tools remained, but it lacked the normal terminal tool. Voice had reached the correct agent; its saved configuration needed repair. The absence of a dedicated inventory tool was not the blocker.
With normal workspace tools, enabled OpenHands skills, and the correct backend URL and authentication context, the Cat can call Agent Server APIs from its terminal. Prefer GET /api/conversations/count for an authoritative count; paginate list/search results for detail. Skills supply instructions rather than permissions or credentials. Resuming an old Cat preserves its saved configuration rather than silently replacing it with the active profile. Focused list/read/wait or create/send tools could improve convenience, but are not required for this architecture.
ask_agent is a reasonable description of the desired experience, not a current Codex tool. The existing SDK endpoint of that name is a stateless one-off question helper, not a saved agent run with tools. Replacing the current conversation append-and-run path with that endpoint would lose the intended execution ownership.
Identity, safety, and call lifetime
A handoff ID is admitted at most once within a call; this is not a claim of global exactly-once execution across crashes or new calls. Repeated or conflicting handoffs and overlapping work are bounded. An uncertain outcome is kept distinct from a request rejected before submission.
request_not_sent identifies a rejection before admission; relay_failed tells the user to inspect the saved outcome before retrying; connection_failed covers transport and protocol failures. Accepted append-and-run work is shielded from call cleanup. End closes Voice and its temporary relay, not the accepted OpenHands task.
Opening the same Cat’s regular Canvas conversation preserves the connected peer. New Cat creates a new saved owner on the next typed send. Switching Cat or backend ends the old call. Condensation keeps the Cat ID but ends the old call so the next session loads current context. Long answers use a labelled spoken excerpt while the complete answer stays saved.
Direct-path verification, September 19–20
A real WebRTC call with generated speech produced a saved user request and OpenHands answer, then returned that answer as spoken audio. A fresh call on the same Cat recalled the saved test word, saved the new question and answer, and spoke the correct word. Both used the existing Codex sign-in and a Luna/high OpenHands Cat. These are two observed end-to-end requests across two calls, not a promise of universal delegation.
The earlier Luna/high relay had received a handoff but completed without choosing send_to_insider; no request reached OpenHands. That failure motivated direct dispatch. It does not show that Luna cannot be the saved Cat’s model: the subsequent direct-path checks used it successfully.
A further September 20 test repaired the earlier saved Cat by adding the standard terminal, editor, task-tracker, and browser tools, preserving its ID, model, approval policy, and all existing events. A generated spoken calculation then caused a real terminal action and observation, saved the answer, and returned the matching result as audio. This verified terminal execution by the receiving OpenHands agent; the separate API check below then verified skill-guided backend access. A transient waiting acknowledgment was misleading before the correct final response, so progress speech remains a usability follow-up.
Skill-guided API verification, September 20: after restoring the default OpenHands skills and accurate backend context, a natural spoken request reached the same saved Cat. It invoked openhands-api, read the backend’s OpenAPI to discover the count operation, called GET /api/conversations/count through the terminal, and saved its answer. An independent API check matched the result; the verified answer returned as matching speech with nonzero received audio. The spoken request supplied neither an endpoint nor a shell command. The existing history and Cat identity were preserved. This demonstrates the normal tool-and-skill path without a dedicated inventory tool.
A separate live launch check created a temporary Cat through the active profile and verified terminal, editor, task tracker, browser tools, eleven default skills, project-skill discovery, and the advertised backend/authentication references in its saved configuration. That check started no agent turn and removed only its own temporary conversation. It verifies new-conversation configuration separately from the saved Cat’s real Voice execution.
Physical iPad microphone, speaker, interruptions, approvals, and repeated-use behavior remain device checks. The public API transport needs its own live provider test. A transcript or a plausible spoken answer alone is not enough evidence; verify the new request and result in the saved conversation.
Configuration and development record
The prototype selects Codex with OH_VOICE_PROVIDER=codex, uses a dedicated Codex home (optionally OH_CODEX_VOICE_HOME), and gates the tested codex-cli 0.154.0 executable and restricted effective configuration. Sign-in and account access must work on the server. See the App README for setup; this notebook contains no account credentials or local runtime paths.
Codex background-thread model settings, the Live audio model, and the OpenHands Cat’s model are three separate choices. A background-thread setting cannot add tools to the Cat. The current bridge supplies no Codex dispatch tool; the single send_to_insider fixture below belongs to the retired relay design.
Development records: Insider App #2, Canvas #17558, and Agent Server #5180. They are not a release or account-eligibility guarantee.
Historical source investigation and model-driven relay trials · September 15–16
Historical design below: these sections explain the protocol investigation and earlier tool-based relay. References to a Codex model choosing send_to_insider, or completing without that call, are superseded by the direct server handoff. The observed trials and their limits remain useful evidence.
Earlier development record on September 19
The implementation is recorded in Insider App draft PR #1, Canvas draft PR #17558, and Agent Server draft PR #5180. The tested SDK worktree is also preserved in commit 7c806a9; its Voice and condensation changes match that upstream draft. These are development records, not released features. The live speech evidence below remains dated September 16.
What was inspected
The Codex source was refreshed from openai/codex and reviewed at c51cb968e4. The implementation and upstream mock-test assertions below were read. The September 15 source-only investigation did not open a microphone or create a provider session. The separately dated September 16 live checks below used the authorized existing sign-in and generated speech.
Correction to the earlier conclusion: separately billed public API access does not imply that every Codex Voice integration requires an API key. Codex has a separate authenticated application path. Calling the public Realtime API with a ChatGPT token and asking Codex app-server to manage its own Voice session are different integrations.
Authentication depends on transport
| Path | Authentication in inspected source | Implication |
|---|---|---|
| Insider → public Realtime API | Agent Server uses a standard OpenAI API key. | This transport still needs API access. The optional Codex relay below uses a separate authentication path. |
| Codex standalone realtime WebSocket | Calls realtime_api_key(), including an environment fallback for ChatGPT sessions. | Signing into ChatGPT alone does not satisfy this transport’s current key requirement. |
| Codex WebRTC call | Uses ModelClient and its current authentication. ChatGPT sessions use the ChatGPT bearer and account identity; API-key sessions use their API key. | This is the candidate for ChatGPT-authenticated Voice through app-server. |
| Codex attachment to an existing call | Uses the current model client’s authentication for the sideband connection. | It attaches to an already negotiated call; it is not an independent way to obtain one. |
Evidence: transport preparation, API-key helper, and WebRTC authentication and sideband identity.
The upstream test conversation_webrtc_frameless_chatgpt_sends_codex_headers_to_backend creates dummy ChatGPT authentication, starts WebRTC with protocol v3, and asserts a call to /backend-api/codex/realtime/calls with its access token. It also asserts client delegation and the selected voice model. This is explicit mock evidence of the intended path, not a live account test. Read the test.
The app-server surface exists
The protocol exposes thread/realtime/start, appendAudio, appendText, appendSpeech, stop, and listVoices. Clients opt into experimental APIs during initialization. Starting a call requires a Codex thread; the response acknowledges submission, while notifications report SDP, session start, transcripts, errors, and closure.
Offline executable check: the installed codex-cli 0.154.0 successfully generated JSON Schema and TypeScript bindings using a separate temporary configuration directory. Its experimental schema contains the three transports, v1/v2/v3, and all 19 start-parameter names in the reviewed source. The stable schema excludes realtime request methods. This proves that the executable exposes the protocol; schema generation does not negotiate a voice call.
Browser: after an explicit Start Voice, create the WebRTC offer.
Insider’s trusted backend: send thread/realtime/start with the mapped Codex thread ID, outputModality: "audio", and transport: { type: "webrtc", sdp }.
Codex app-server: establish the call using its own authentication and return the answer through thread/realtime/sdp.
Browser: apply the answer and keep audio attached to that conversation across Canvas views.
Use an explicit protocol version. WebRTC and existing-call attachment accept v1 or v3 and reject v2 in this revision; omitting the version selects v1. Protocol v3 defaults to gpt-live-1-codex; v1/v2 default to gpt-realtime-1.5. A model override field exists, but its presence does not prove permission to use any public API model through ChatGPT authentication.
For v3, initialItems can supply bounded, role-bearing text history: at most 128 items and 8,192 estimated text tokens in total. includeStartupContext: false prevents automatic injection of Codex’s startup history. Existing-call attachment cannot configure a new model, voice, prompt, or initial history.
That initial history seeds Voice, not the backing Codex thread. Subsequent v3 appendText calls discard the supplied role, so they cannot serve as a role-preserving history import. Codex also persists its own realtime timeline; synchronization with the saved OpenHands conversation remains integration work. Text adapter; history persistence.
Evidence: RPC methods, start parameters and transports, and version selection and validation.
The difficult part is conversation ownership
Codex app-server Voice is integrated with the Codex agent. Incoming voice handoffs are routed into the owning Codex thread before the public notification is emitted. Merely listening to those notifications and submitting the same request to OpenHands would risk running the work twice.
clientManagedHandoffs does not turn off that incoming routing. It suppresses automatic delivery of Codex responses back to Voice so the client can append output explicitly. It is not an external-agent delegation switch.
Evidence: incoming handoff routing, outgoing handoff guard, and the parameter’s definition in the protocol.
Two designs deserve separate evaluation:
- A Codex-backed Insider controller: keep one Codex thread as the companion’s owner and adapt Canvas history, new/resume, condensation, approvals, and agent actions to that owner. This is a larger backend integration.
- A Codex relay to the existing OpenHands Cat: the Codex turn invokes a constrained integration tool that addresses one bound OpenHands conversation. OpenHands remains the authority for task outcomes. App-server supports dynamic tools, but the Codex model must choose to call the tool during its turn. This adds a Codex reasoning turn and requires explicit transcript synchronization and duplicate-request prevention. This is the design implemented in the prototype below. Dynamic tool request/response test.
For either design, persist the owner mapping on the server. Navigation must reuse it; New must create a distinct owner; resuming must reload the correct history; condensation must end the old voice session and seed the next from the condensed context. Each request must have one execution owner, with approvals and outcomes visible in the same conversation the user is reading.
Selected prototype: a relay to the saved Cat
The selected design keeps OpenHands as the durable conversation owner. Each Voice call creates a fresh, ephemeral Codex relay thread, bound on the server to that one Insider controller. This refines the mapping proposal above: the saved OpenHands ID survives navigation and resumption; the Codex thread lasts only for the call. Restarting Voice seeds a bounded snapshot of the current saved Cat context instead of accumulating a second permanent history.
Voice → Codex relay: a delegated spoken request starts a Codex turn whose constrained tool addresses only the bound Cat.
Relay → OpenHands: save the request once, retain the Cat’s approval policy, and wait for the actual run to settle.
OpenHands → Voice: return the saved result or an explicit pending, approval, or failure status. Accepted work is not a completed result.
The browser must not also submit the transcript as a second OpenHands request. A backend-selected provider distinguishes this server-owned relay from the existing browser-owned public Realtime tool bridge. The original API-key path remains available.
Stopping Voice closes its connection and relay; it does not undo work already saved in OpenHands. Changing Cat, starting a new Cat, or condensing the current Cat ends the old call. Opening the same Cat in another Canvas view keeps its identity and active call.
The relay keeps its own configuration and can share an existing sign-in with explicit user authorization. The prototype selects Codex with OH_VOICE_PROVIDER=codex and gives it a dedicated home, optionally set with OH_CODEX_VOICE_HOME. The tested executable is codex-cli 0.154.0; an untested version or enabled MCP configuration fails the availability check.
Setup update, 16 September: for an existing file-store sign-in, link only the dedicated home’s auth.json to the existing Codex auth.json, and set cli_auth_credentials_store = "file" in the dedicated configuration. This shares the session without copying tokens or loading the ordinary home’s configuration, plugins, or MCP servers. The pinned version’s file-store save follows the symlink, so token refresh updates the shared file; this arrangement relies on that behavior and does not apply to keyring storage. File-store save implementation.
The successful live trial set model = "gpt-5.5" and model_reasoning_effort = "low" in that dedicated configuration. These select the Codex reasoning relay, separately from the Live voice model. The saved OpenHands Cat also used GPT-5.5 in the trial.
This separation follows an executable test: an empty mcp_servers override still inherited a configured local MCP server, even with empty environments. The restricted fixture exposed only send_to_insider to the model. A dedicated configuration, feature restrictions, and a version check are required for this prototype; prompt instructions alone do not establish that boundary.
If no suitable existing session is available, sign in normally in the dedicated home. On the server machine, run CODEX_HOME=/path/to/insider-codex-home codex login --device-auth, then use that same directory for OH_CODEX_VOICE_HOME. The device flow lets the user finish sign-in on an iPad through OpenAI’s device sign-in page. Ordinary browser login uses a localhost callback, which normally points at the wrong machine when opened on a different device. Keep the generated one-time code private.
Short saved replies can be spoken in full. Long replies use an excerpt from the beginning and explicitly say that the full answer is saved in the conversation. The durable answer is never shortened to fit speech. Live captions are transient; they are not a second saved transcript.
Codex mode offers microphone mute and End call. Its current app-server protocol has no verified speech-only interruption operation, so the App hides the separate Stop speaking button for this provider. Spoken stop, spoken hangup, real microphone handling, and provider voice quality still require their own live validation.
Local verification on September 16
In the earlier fixture checks, the actual installed app-server ran behind the actual Agent Server and Insider App. Only the provider endpoints, model replies, and microphone input were substituted. Browser peer statistics showed audio packets in both directions and an actively playing silent audio stream. Those fixture checks used no physical microphone, real provider session, or real account credential.
Two synthetic Voice requests produced exactly two saved requests and two saved final answers in one Cat. The first call survived App → regular conversation → App navigation; the second request was made from the regular view. Each captured Codex model request exposed only the bound send_to_insider tool. Mute disabled the outgoing audio track, and End closed the browser peer and provider sideband.
Automated checks cover provider selection, duplicate tool calls, stale events, rapid restart, saved-result settlement, long reply excerpts, approval and busy states, scoped call access, and cleanup. A disconnected relay cannot cancel work that OpenHands has already accepted. These checks establish the local integration, not real speech recognition, voice quality, account eligibility, or live-model compliance with the delegation instructions.
Recorded focused checks: 64 App tests, 42 Agent Server Voice tests, and one HTTP regression passed; bundle parity, SDK hooks, and documentation links also passed. That backend revision was exercised with a resumed browser call and simulated provider hangup. Several earlier repeated fixture calls stalled during local ICE setup; a later call connected after all old fixture peers were released. Their exact transport cause was not established, so repeated real-provider calls remain an explicit validation item.
Live delegation and recall on September 16
A real provider call then connected using the existing Codex ChatGPT sign-in, shared with explicit authorization through the file-store setup above. It needed no new login or OpenAI API key. Generated spoken input asked the Cat to save the test word “lantern.” Voice handed the request to the Codex relay, which called send_to_insider; the OpenHands Cat saved the request and its confirming answer. The verified result returned through appendSpeech, and the native WebRTC receiver reported nonzero audio energy. Opening the regular conversation view preserved the same connected peer.
A second spoken request from the regular view was saved once and answered with “lantern,” while the same peer remained connected. After End and a fresh real provider call on the same Cat, another recall request was recognized, saved as plain user text, and answered correctly by OpenHands and Voice. In total, three spoken requests and their answers were saved across two successful calls. These checks used real Voice and reasoning models with generated speech instead of a physical microphone.
The fresh-call result verifies that restarting Voice can recover the updated saved Cat context. Physical iPad microphone input, audible playback on that device, and broader repeated-call and approval behavior still need validation.
Delegation is still model-dependent. An earlier live trial answered a recall question directly from Voice’s initial history without saving a new OpenHands request. The later successful trial used stronger Voice instructions, an explicit relay-mode instruction override, and the GPT-5.5/low relay settings above. That success does not establish that every utterance will delegate, and the protocol exposes no switch that forces it. A correct spoken answer alone is insufficient evidence: check the saved request and result. If a Codex relay turn completes without calling the Cat tool, the server now ends that call with an explicit unsent-request message and a text fallback. Direct Voice speech that never creates a relay turn remains outside that check.
Current focused checks: 67 App tests and 44 Agent Server Voice tests passed, along with bundle parity, SDK hooks, and documentation validation. The temporary browser microphone substitutes were removed; all test peers, tracks, and audio contexts were closed.
Remaining live checks
- Completed offline: verify the installed app-server’s generated protocol against the selected source revision. Updating a source checkout does not update the installed executable.
- Completed with fixtures: initialization, thread creation, WebRTC start/SDP/stop, two saved requests and results, navigation, mute, and cleanup through the running browser integration.
- Completed on this account: three saved request-and-answer round trips across two real Voice calls, including navigation during a call and recall after End and restart, using an explicitly authorized shared file-store session. Other accounts still need their own connection check.
- On iPad, validate audible answers, recognition, live-model delegation, repeated calls, reload/resume, New, condensation, interruption, and approval pauses. Local fixture coverage does not replace this device and provider test.
Recommendation: continue with the tested Codex configuration and verify durable delegation over multiple turns before treating Voice as the Cat’s dependable input path. This account’s existing sign-in carried the live test without an OpenAI API key; that observation does not establish billing terms or access for other accounts. Keep the API-key transport available as a separate option.