Agent Canvas · Apps · behavior and architecture

Insider Cat

A sibling of SmolPaws inside Agent Canvas: one saved OpenHands agent, with a Cat-shaped place to talk, type, and find your work.

Updated . Development implementation · review pending

Voice is a way to reach the Cat. OpenHands does the work. The selected Cat has a durable conversation, agent configuration, tools, and approval policy on Agent Server. The Voice connection carries requests and speaks saved results; it does not replace that agent or grant it access to every conversation in Canvas.

A sibling with its own place

The standalone Insider App contributes the Projects page and, on a compatible Canvas host, a travelling companion. The Cat has its own identity and history. The compatibility tag smolpaws: insider does not give it SmolPaws’ credentials, memories, channels, or permissions.

Canvas owns navigation and the visible controls. Agent Server owns the saved conversation and execution. The chosen OpenHands agent profile supplies the model and actual tools. Selecting an LLM profile or switching models alone does not add tools. The App adds its Insider instructions when creating a Cat; instructions explain the role but cannot manufacture missing capabilities.

The Cat comes first

The September 19 interface puts a warm cream SVG Cat, its status, and Talk beside the conversation. The saved-Cat picker and New Cat sit with that conversation; Your work holds the conversation browser. On narrower screens the work list moves below the Cat. The page has one set of Voice controls.

Resting
Listening
Speaking
Working
Actual Cat artwork from the App, shown here as static state illustrations, not a browser screenshot. Adjacent text and controls carry the status; color and animation are supplementary.

The Cat sleeps when idle, perks up while connecting, listens when the microphone is active, and stays awake while speaking or working. Pending agent work takes precedence over a listening pose. Mute removes the active-listening cue; it does not imply the agent stopped. Errors, unknown state, paused work, and approval requests stay visible instead of looking like a peaceful idle Cat. Reduced-motion preferences disable the decorative animation.

Before a Cat is selected, Talk is visible but disabled, with an explanation. Choose a saved Cat, or select New Cat and send the first typed message. Away from Projects, a compact companion keeps the same call and Cat available in the Canvas layout. It occupies space below the main content instead of covering the chat composer. Returning to Projects restores the main view without creating another connection.

UI and artwork: shared Cat view and pose mapping and SVG.

One saved Cat, across views and calls

ActionWhat happens
Choose a saved CatReload that controller by its full ID on its owning backend. The App discovers controllers across list pages; multiple Cats require a deliberate choice unless a remembered choice is verified.
Open the full conversationContinue the same ID in ordinary Canvas chat, including its tools and approval controls. Its Insider badge and return link distinguish it from a worker. Navigation preserves the active call.
New Cat or /newPrepare a fresh Cat. Its first typed send creates a different controller, using the then-active OpenHands profile and chosen workspace. This does not clear the old Cat or clone its whole history.
/condenseCondense the selected Cat without changing its ID or starting another agent turn. End its Voice session first; a new call loads the updated context. The saved event record remains inspectable.
Select a workerAttach the full worker ID and current metadata to the next typed Cat request, while preserving the draft. Selection alone does not send the worker anything.
End callRelease the microphone and close Voice. Work already accepted by OpenHands continues; its result remains in the saved conversation.

Cat selection is scoped to the backend and organization, and revalidated after loading. Changing Cat or backend, disabling the App, or requesting condensation ends the old call. A regular conversation you open elsewhere in Canvas does not silently become the Voice target. Draft text survives navigation within the active App, but is not persisted across reloads or backend changes.

The App shows text from recent saved events and a link to the full conversation. Live captions are transient. Saved delegated requests and OpenHands answers are the task record; exact audio and every word spoken by the Voice model are not a second complete transcript.

Use the normal OpenHands tools and skills

The Cat can inspect and coordinate conversations through the existing Agent Server API using its normal terminal tool. The OpenHands API skill provides instructions; the runtime must identify the actual backend URL, authentication source, and relevant state or workspace. A dedicated conversation-listing tool is an optional convenience, not a prerequisite.

The Projects page’s JavaScript methods are not automatically model tools, but that does not prevent a normally configured OpenHands agent from making the same supported API requests through its terminal. Skills explain how to use existing capabilities; they do not independently grant network access, credentials, or permissions. Selecting an LLM profile also does not change the saved agent’s tool configuration.

A September 20 inspection found that the reported Cat was an earlier bare fixture: its configured tool list was empty and it had no client tools. The remaining reasoning and finish tools did not include terminal access. The reported parallel entry was not a registered native SDK tool in that configuration. This was an incomplete agent configuration, not evidence that OpenHands needs a new inventory-tool implementation.

The same Cat was repaired by adding the standard workspace tools without rewriting its history or changing its ID, model, or approvals. A real Voice test executed a terminal calculation and spoke its saved result. A subsequent check restored the default OpenHands skills and supplied accurate backend context: the same Cat loaded openhands-api, discovered the count endpoint from OpenAPI, and called it through the terminal. Its saved count matched an independent API check and returned through Voice. No bespoke inventory tool was added.

LayerResponsibilityBoundary
Projects UIList, filter, page through and open conversations; select a target; send to the Cat.A loaded page is not the complete server inventory.
Saved OpenHands CatUse its configured tools, including terminal, under its approval and security policy.Older saved Cats retain their actual configuration; a new model does not repair missing tools or context.
OpenHands skills and runtime contextExplain the APIs and identify the correct backend, authentication source, workspace, and state.Instructions are not permissions. Use available authorized access without exposing credentials or guessing another backend.
Voice bridgeBind a spoken request to that Cat and return its verified result.The OpenHands agent remains the execution owner, whether it uses a terminal API request or another tool.
Optional focused toolsWrap list/count/read/wait or create/send requests with typed inputs and convenient results.They may improve ergonomics and validation; they are not required for API access through the terminal.

For a conversation count, prefer GET /api/conversations/count on the owning Agent Server. It is the authoritative API count for that backend and query. For titles, statuses, or per-conversation detail, follow the relevant list/search pagination and state any filters or incomplete coverage. Count conversations, not messages, events, or files.

A filesystem check is useful only when the runtime confirms the current local backend’s persistence directory and the agent has access to it. Count distinct conversation records there, not every JSON file or an old checkout’s state. For a remote backend, local files are not its inventory. Use the live API when possible and explain any narrower fallback scope.

The OpenHands API skill already documents terminal-based local and Cloud workflows. It is default-enabled for new workspaces in the registry; existing saved skill selections still need to be checked. Confirm the actual URL and authentication source rather than copying example defaults. Keep credentials out of tool output, chat, and notes. For an uncertain write, inspect the saved conversation before retrying.

A name such as ask_agent can describe asking the real agent to investigate. It is not an implemented Codex tool in this path. The SDK’s existing /ask_agent endpoint answers a stateless, one-off question; it does not run the saved Cat with its tools. The current send_message(run=True) handoff preserves that saved execution owner.

Two Voice transports, the same execution owner

1. The person speaks after explicitly choosing Talk and granting microphone access.

2. The selected Voice transport hands off a request to one bound Cat conversation.

3. OpenHands saves and runs it with that Cat’s own model, tools, workspace, and approval policy.

4. The bridge waits for the whole run to settle, including stop hooks, and retrieves the new saved answer.

5. Voice speaks that answer or a labelled excerpt. The full result stays in the same conversation.

Codex: direct server handoff

In the current experimental Codex path, Agent Server consumes thread/realtime/itemAdded items of type handoff_request. It validates the explicit input_transcript, admits each handoff_id at most once within that call, and sends it directly to the bound OpenHands conversation. The saved result returns through thread/realtime/appendSpeech.

The Codex background thread exposes dynamicTools: [] with its other tool access restricted. The pinned app-server may still start a background turn, but that turn neither dispatches the request nor supplies the saved Cat’s answer. The browser does not also forward Codex transcripts. This replaces the older model-chosen send_to_insider relay, which could receive a handoff and finish without calling the Cat.

Live still has to emit a handoff. A sentence spoken directly by the Voice model is not evidence that OpenHands received or saved a request. A new Codex call starts from a bounded snapshot of the saved Cat; the Projects page’s selected worker is not included in that session. The Codex implementation note records protocol, configuration, history, and live evidence.

OpenAI API: browser tool bridge

The separate public Realtime API path uses gpt-realtime-2.1 and a server-held OpenAI API key. Its Voice model calls send_to_insider; the App forwards that request through the Cat’s normal conversation API and waits for confirmed run completion. This tool belongs to the Voice bridge, not the Cat’s own toolbox. This transport retains fixture coverage; it needs its own live provider validation.

The server selects the transport. Codex uses its dedicated, tested sign-in/configuration path; an OpenAI API key belongs to the public API path. Neither substitutes the other’s credential automatically. No provider credential is sent to the browser.

Controls and failures

Both transports offer Mute and End call. The API transport also exposes Stop speaking; the Codex App hides that action because its inspected app-server protocol has no verified speech-only interruption operation. Ending speech is separate from stopping agent work.

request_not_sent means the bridge rejected the request before admission. relay_failed means the saved outcome needs inspection, and must not encourage a blind retry. connection_failed reports a connection or protocol problem. A busy Cat can reject a new spoken request while an earlier accepted request continues.

Install the matching components

The package is an ordinary Canvas App with a prepared, self-contained browser bundle. It replaces the historical Secretary skin server, iframe bridge, fixed browser ports, and page-owned credentials. The enhanced companion and Voice paths require matching Canvas and Agent Server changes; choosing an App ref alone cannot add missing host or backend capabilities.

  1. Use a local Agent Server connection with Apps, agent profiles and launch additions. The inspected Apps host does not support the cloud backend.
  2. Open Customize → Apps. Use Source github:enyst/insider and a reviewed Ref; leave Repository path empty. For a local source, use the package directory on the server machine and leave both other fields empty.
  3. Configure the active OpenHands agent profile in Canvas settings, enable the installed App, and open Projects. Choose a real workspace for a new Cat. Check that it has the normal tools, enabled OpenHands skills, and accurate backend connection context.
  4. Configure the selected Voice transport on Agent Server. The availability check runs before requesting the microphone. Typed chat remains usable without Voice.

The server copies extension.js; it does not build source. Rebuild and reinstall after source, translation, or skill changes. The App README contains setup details; the Codex note distinguishes the current transport from earlier experiments.

An Insider skill explains the role

The standalone package owns skills/insider-cat/SKILL.md. The build embeds its body and appends it through agent_launch_additions.system_message_suffix_append when creating a controller. This preserves the profile’s existing instructions and policies and works independently of the Cat’s working directory. It complements the normal OpenHands skills. It does not install tools or rewrite an already saved Cat’s configuration.

The skill distinguishes the Cat from selected workers, requires actual evidence for actions, and separates acceptance from completion. It grants no shared SmolPaws memory, scheduler, or automatic monitoring. The Odie-local skill and old skin remain part of the development history.

New Cat launches also read the owning backend’s advertised runtime_services and include validated agent-side URLs and authentication environment-variable names. Credential values are not included. The host must actually provide those references to the agent. The September 20 direct-Vite fix restores this metadata on that development path; missing metadata is omitted rather than replaced with a guessed backend.

Make capability and continuity easy to see

  1. Keep the Cat, status, and Talk together. An idle Cat rests; connecting, work, approvals, and errors remain legible in words. Avoid duplicate controls when the page and companion share a call.
  2. Show what this Cat can actually do. Check the agent’s normal tools, enabled skills, and actual backend context. Terminal API access is a valid way to inspect conversations; do not mistake the absence of a dedicated wrapper tool for inability.
  3. Keep one visible target. The controller, selected worker, and merely viewed conversation are separate identities. Confirm ambiguous names and use full IDs for actions.
  4. Make coverage explicit. Your work filters the loaded pages, and Load more expands them. Agent answers about all conversations require complete server data or an honest partial count.
  5. Keep receipts with results. Voice should speak saved outcomes; the full conversation shows tools and approvals. An acknowledgment is not completion, and silence is not evidence that work stopped.

Configure and verify the existing capability first

Use the normal OpenHands toolset and enabled skills with accurate runtime context. Verify a read-only count or list on the intended backend, then read and follow up on an explicitly chosen worker under the existing permissions and approval policy. Focused conversation tools can be added later for convenience. Scheduled “keep me updated” work needs a durable Automation service and its normal API or skills; an open Voice connection is not a scheduler.

Design references

The earlier exploration drew on Amp’s Puck for the distinction between a companion conversation and worker threads, and Talk to Puck for Voice handing work to a real agent. ElevenLabs remains an alternative to evaluate later. Neither is a dependency of the current App. Provider comparisons in the voice landscape are dated research, not current installation guarantees.

Evidence, history, and limits

5 October: integration maintenance

The feature is split across three implementation PRs: Agent Server owns Voice and saved work, Canvas owns the companion surface, and the Insider App owns its controls and audio. The integration issue records acceptance criteria and current verification; the Canvas issue tracks host behavior. The App and notebook repositories have issues disabled, so their PRs link back to those issues.

The failed Agent Server macOS job stopped while downloading a dependency from PyPI, before building or running tests. It was a connection timeout, not a demonstrated Voice failure. A fresh rerun and cross-repository checks are recorded on the linked issues and PRs; the September provider observations below remain historical evidence.

September 19–20 direct-handoff checks: generated speech reached the real saved OpenHands Cat and returned its saved answer as audio. A fresh call then recalled the saved test word, saving its new question and answer in that same Cat. These checks used the real provider and existing Codex sign-in, with generated input rather than a physical microphone. They establish those observed requests, not universal delegation or account eligibility. On September 20, a further natural spoken request also verified skill-guided API use: the Cat independently discovered and queried the conversation-count endpoint, saved the correct result, and returned matching speech. The request did not supply an endpoint or command.

Physical iPad microphone, speaker, interruption, approval, and repeated-use behavior still needs device verification. The public API transport’s provider speech remains unverified. No fresh page screenshot is claimed here; the artwork is copied from the App source.

The Secretary proposal documents the earlier skin. Its authentication, proxy, memory, and dispatch assumptions should not be read as the current App contract. The September 15 fixture and September 16 model-driven relay trials remain historical evidence in the Codex note.

Reviewed against App 7efa337, published SDK 02a51ae46 (the equivalent tested worktree change is 957c1c584), and the Canvas draft. These pages describe development revisions, not a release announcement or live service-status monitor. Beads owns follow-up status.