control-openhands: agents using Agent Canvas like a user
The command-line lever inside the verify-openhands skill from OpenHands PR #17961. What it does, what its commands are, why an agent is better off with it than running the app directly, and what happened when 36 fix agents used it in parallel on the same day.
1What it is
control-openhands is one Node command, about 5,500 lines, that lets an agent use Agent Canvas the way a person would, without fighting the environment, and leaves a record of what it saw. It builds the checkout when needed, starts it as an isolated stack, drives it in a real Chromium through the repo's own Playwright, and writes evidence. Every command prints one JSON object and exits 0 (ok), 1 (action failed), 2 (usage) or 3 (environment). A failure carries an error, a hint and a screenshot.
It comes with a feature map: 27 family files and 715 stable IDs that cover every user-facing behavior, each with recipes written as exact commands from the user's point of view. The CLI is what makes those recipes rerunnable by the next agent. This is the “feature map + control-app CLI for the product surfaces” that What we take from pstack chose to adopt, built for Agent Canvas.
2The commands
| Group | Commands | What they're for |
|---|---|---|
| Lifecycle | launch doctor status restart service stop runs env | Build the checkout only when its build inputs change. Start a stack with its own private HOME, ports, keys and state, plus a browser. doctor checks the stack is healthy and serving this checkout's build. restart keeps state; --rotate-key makes the stored key stale; service stop takes one backend down to reach backend-down screens. stop ends the whole process group. |
| Arrange | api llm fixture | Fast preconditions: API calls with the run's key, LLM profiles from a key file, throwaway git repos, folders and files. Never used as proof. |
| Essential paths | login onboard workspace open conversation start/send/wait/events | Multi-step flows through the real UI that nearly every check needs. |
| Drive and observe | browser plus about 50 verbs | One browser that persists across calls. Clicks, fills, key presses, drag, paste, file choosers and viewports. It also reads what a screenshot can't show: errors, network, toasts --history, storage, downloads, media, clipboard, tooltip and the accessibility tree (snapshot). |
| Evidence and map | evidence add/report map check/coverage/ids/routes | A pass/fail/blocked/not-run record per feature ID, and checks that keep the map in line with the code: valid IDs and commands, every route and feature directory either covered or listed as excluded. |
A recipe is just a sequence of these, for example adding a secret:
control-openhands launch --new --build auto --print-run # isolated stack + browser; prints the run dir
export OH_VERIFY_RUN=<run dir>
control-openhands doctor # is the served build this checkout?
control-openhands onboard --skip
control-openhands browser goto /settings/secrets
control-openhands browser click 'testid=add-secret-button'
control-openhands browser fill 'testid=add-secret-form >> testid=name-input' QA_TMP
control-openhands browser screenshot --feature F14.create --name form
control-openhands evidence add --feature F14.create --result pass ...
control-openhands stop
3Why not just run the app
An agent can run npm run dev and click around. For a one-off, that's fine; the pstack note says the same. The CLI earns its place when verification recurs and other agents must be able to rerun it:
- Isolation, so several agents can work at once. Each run has its own HOME, ports, keys and state. Run directly, the app wants fixed ports and a shared
~/.openhands, so two agents trample each other's settings and conversations, and every agent improvises its own launch. - It knows which code it's testing.
launchrebuilds only when the build inputs change, anddoctorrefuses a stack whose served bundle doesn't match the checkout. The most damaging agent failure is “verified the fix” against a stale bundle, and nothing in the output warns of it. - A browser that stays open, and listeners that never detach. An agent works in separate shell calls. A script per step either relaunches Chromium and loses the page, or re-attaches to a running Chromium over the DevTools protocol: that keeps the page, but misses whatever happened between steps, such as a toast shown for under a second, a request fired during page load, or a console error. The CLI's browser daemon holds one Chromium per run and keeps listening the whole time, recording errors, console messages, requests, dialogs and downloads for later commands to read back. Each command is a quick call to it, so a check is a plain list of commands. That is what lets the map be written as instructions someone can rerun.
- It sees what pixels don't show. Many of the bugs found while mapping were invisible in screenshots: telemetry requests despite Do Not Track (
network), double error toasts that disappear within a second (toasts --history), hydration errors (errors), switches that Tab skips and icon buttons with no name (snapshot). - Setting up stays separate from proving. An agent may create a secret through the API to save time, but the check passes only on what the UI shows. Without that rule, agents quietly verify the API and call it a UI pass.
- Secrets and resources are handled for it. Keys are typed from files and never echoed, storage and network output are redacted, a launch is refused when memory is low, and stopping kills the process group, so nothing is left running.
- Hard-won fixes live in one place. A fix to the harness protects every later agent instead of being rediscovered by each one. Section 5 lists the ones from a single day.
- It reuses the app's own handles, and catches when they break. Recipes select elements by the
data-testidhandles the app already ships for its tests (set in about 1,350 places in its source onmain), so they survive copy and layout changes;testid=add-secret-formis just the CLI's short form for one of them. Because the map relies on them, a broken handle surfaces as a bug instead of a check that quietly finds nothing: the Stop and Delete confirmations rendered without their test IDs because callers passed the wrong prop name (#17915).
4CLI was used in 36 of 38 fixes
On 5 October Engel asked for every ready-for-dev issue from the mapping run to be fixed, but only where nobody else had offered. 38 qualified: 14 other issues already had a contributor's pull request, and two had volunteers in their comments. Each issue got its own git worktree off main, with a copy of the skill inside so that the CLI serves that worktree's code, and a pipeline: fix (reproduce live on main, failing test first, fix, prove live) → independent review (checks the tests fail on the old code, re-checks the fix live, looks at the evidence; up to two revision rounds) → pull request. Four agents ran at a time on a 4-CPU, 16 GB machine; the CLI's memory guard refuses a launch below about 1.5 GB free.
Final numbers across all 120 fleet agents: the fixers, the independent reviewers who re-checked each fix live, and the revisers and publishers. The remaining two of the 38 fixes (SDK #5495 and automation #551) work on Python services without a UI and used pytest plus rendered terminal captures instead.
The result: all 38 fixes passed independent review, three of them after one revision round. 35 became pull requests: 33 on Agent Canvas, and one each on the SDK and the automation service. Two were held back because contributors claimed those issues while the fleet was running, and one is waiting on a push.
Almost every fixer launched twice: once on unchanged main to reproduce the bug and capture a “before”, once on its own fixed build for the “after”. The reviewer launches again. No agent wrote harness code of its own. They picked the verb that fits the bug:
| Bug | Verbs that carried it | What the CLI made observable |
|---|---|---|
| First-run “No backend is configured” toast (#17901) | launch --public, browser toasts ×19 | A toast over the Add-a-backend step and the API-key screen, twice after a reload. |
| A failed “Run now” shows two error toasts (#17906) | api ×14, browser wait-text ×27, toasts ×17 | Arranged through the API (delete the automation elsewhere), proved in the UI. |
| Unknown URLs: bare 404 and a hydration error (#17904) | browser errors ×17 | Page and console errors counted before and after, not just a screenshot of the page. |
| Icon-only buttons without accessible names (#17935) | browser snapshot ×28 | The accessibility tree, where a nameless button is visible as such. |
| Header menu off-screen at phone width, ignores Escape (#17913) | viewport phone|narrow, bbox, press Escape | Evidence such as before-phone-390-menu-overflows-right-edge.png and after-narrow-320-menu-inside-viewport.png. |
| Automation “View logs” says “(no output)” (#17947) | browser network ×9, api ×8 | The request the log view makes and what the server returns. |
| Command menu shortcut hint and missing pages (#17908) | browser press ×13, text ×12 | Control+k and Meta+k driven for real, with the hint text read back. |
Across the whole fleet, the most used verbs were click (455), screenshot (261), goto (245), wait (218), doctor (142), toasts (116), snapshot (79), errors (77) and network (51). doctor running 142 times is the point of rule 2: agents checked that they were looking at their own code before believing what they saw.
5Fixed once, for every agent
The same day turned up several harness faults, from agents running it and from two reviewers of the PR. Each fix now protects every later run:
restartreplayed a stale proxy address saved at launch, so model calls failed after the machine's proxy changed. Machine-scoped settings now come from the current shell.launch --run-idcould reuse a live run's directory, overwriting its keys. A named run must now be new, and the directory is created exclusively.- The build identity ignored
config/, which the app imports, so a config change kept serving the old bundle anddoctoraccepted it. A test now checks that every file the app imports from outsidesrc/is a build input. browser storage --valuesprinted an API key stored as a plain string. Values under secret-looking keys are now redacted whole.- About 85 recipe steps assumed a page an earlier step had left, and 17 depended on state a later step creates. A sweep of all 27 families fixed them, and every changed sequence was rerun as written.
6Compared with the e2e tests
Agent Canvas already has Playwright e2e suites on main: 26 spec files, 77 tests, about 9,200 lines. 23 specs run against a scripted mock model, one against a real Agent Server and model, one checks the launcher's network binding, and one measures markdown rendering speed.
The e2e tests assert about 10% of the behaviors in the feature map (68 of 715), and check some request-level details the map leaves out, such as an uploaded image reaching the model as base64 data.
| Families | |
|---|---|
| No e2e assertion | F02 App shell, sidebar and command menu · F09 Settings shell · F12 Model router · F14 Secrets · F15 Condenser, agent context and verification · F16 Application settings · F19 Plugins · F24 Automation Git Sync |
| Best covered | F26 Launcher modes 8/27 · F01 First run 7/21 · F10 LLM profiles 7/31 · F20 Canvas apps 7/23 |
| Large and nearly untouched | F04 Conversation list 1/40 · F21 Automations dashboard 2/34 · F03 Home 4/39 |
The other direction: 61 of the 77 tests assert at least one behavior in the map. The other 16 test plumbing the map does not describe: how the launcher injects the session key, Docker mode, port conflicts. Across all tests there are 89 assertions with no map entry. Most are outside the map's scope by design: exact request bodies, event counts, settings persisted through the API, “no error banner” guards, and a 500 ms composer budget. About 20 are user-visible behaviors the map is missing and should gain, among them: the active backend is chosen per tab; a Ctrl/⌘-click on a conversation opens it on the backend that owns it; a failed message's Retry survives a reload and sends exactly once; attaching only an image enables Send; the “Fetching older messages” indicator; the empty states for no apps and no MCP servers; and the Cloud provider list loading past its first page.
7Trade-offs and next steps
map check, map coverage and 20 unit tests in CI guard it, but it needs an owner: in practice an agent, under review.login fills Host Name with Local; a user who misses that field still meets a disabled Connect. The map drives screens by hand when the screen itself is under test.agent-canvas --isolated flag), and consider offering the CLI as an MCP server so agents outside the repo can use it.