Unit tests
Check handlers and models in isolation, with no live server behind them.
OpenHands / Agent Server verification 01 / 10
Drive the real server. Know what you should see. Tell a stale instruction from a real bug.
verify-agent-server lets an agent drive the OpenHands Agent Server, the API that agents and programs use, feature by feature, through the same routes and clients they use. It is built by agents, for agents, and works at their speed.
The verify-agent-server story · Engel Nyst
Software Agent SDK PR #5621 · results from the full replay on 9 October 2026
The problem 02 / 10
The Agent Server is what agents and programs drive: about 190 REST and WebSocket routes on the default app, 203 with Docker runtime mode, plus webhooks, telemetry, and the Python SDK and TypeScript clients that consume them.
Check handlers and models in isolation, with no live server behind them.
Catches changes to the published contract, not changes in behavior.
Exercise a real server, for a selection of paths.
Each covered part of it. None gave an agent a single place to drive the real server end to end, know what it should observe, and tell a stale instruction from a real product bug.
The design 03 / 10
A contributor skill under .agents/skills/verify-agent-server/, adapted from Lauren Tan’s pstack verification skills, which OpenHands already used for its Agent Canvas UI.
How do I drive this checkout? Isolated servers, any REST route or WebSocket with assertions, fixtures for everything around the API, and evidence with secrets redacted.
control-agent-serverWhat should happen? 34 families, 930 stable sub-feature IDs, and every agent-facing route owned by exactly one family. Each bullet carries a shell recipe that proves it on a fresh server.
references/feature-map/What is broken today? A known bug’s failing assertion is marked # bug. The replay expects it to fail, and flags the day it stops failing.
# bug → xfail / xpassHow does it stay true? Route coverage, recipe parse checks, route-table diffs between commits, owner lookup for changed files, and a cross-repo test guarding the structure.
map coverage · check · diff · ownersThe lever 04 / 10
control-agent-server launches this checkout as an isolated server: its own HOME, persistence directory, keys and port.
It reaches every route through api and ws, with --expect for status codes and --check for what comes back. It arranges what the API needs around it: git repositories, skills, plugins, MCP servers, webhook sinks, and a scripted LLM stub.
Model-backed recipes run on DeepSeek, an open-weights model, through the server’s own pre-flight. Every command prints one JSON object, and saved exchanges have their secrets redacted.
export PATH="$PWD/.agents/skills/verify-agent-server/scripts:$PATH"
export AGENT_SERVER_VERIFY_RUN=$(control-agent-server launch --new --print-run)
control-agent-server doctor
control-agent-server api PUT /api/settings/secrets \
--json '{"name": "QA_F18_TOKEN", "value": "**********"}' \
--expect 200,422
control-agent-server api GET /api/settings/secrets/QA_F18_TOKEN \
--expect 200 --check . eq "$QA_F18_SECRET" # bug
No untracked one-off scripts. If a path cannot be driven with the CLI, that is a harness gap: extend the CLI, prove it live, then write the recipe, so the next agent can rerun it.
Adapted from SKILL.md and the F18.secret-empty-value recipe at the PR head.
An executable inventory 05 / 10
Server status, auth and sessions, deferred init (warm pool). F01–F03.
Lifecycle; run, pause and interrupt; events; the session socket; confirmation, security and secrets; model switching and plugins; fork, navigate, condense and ask; goals. F04–F11.
Runtime and credentials, bash, files, git, workspace serving and VS Code, the workspaces registry. F12–F17.
Settings and secrets, MCP, LLM profiles, agent profiles, meta profiles, the LLM catalog and connections, OpenAI subscription. F18–F24.
Skills, plugins, hooks, subagents and tools, Canvas extensions, the app backend bridge. F25–F29.
The OpenAI-compatible gateway, persistence and limits, webhooks and telemetry. F30–F32.
The Python SDK remote client and the TypeScript client. F33–F34.
- `F18.secret-empty-value`: a PUT whose value
is empty or the redaction placeholder must
not leave a listed but unreadable secret or
wipe an existing value.
# The recipe, run on a fresh server:
GET …/secrets/QA_F18_TOKEN --check . eq … # control
PUT /api/settings/secrets value "**********"
GET …/secrets/QA_F18_TOKEN --check . eq … # bug
# Today: 200 on the PUT, then 404
# "Secret not found". Filed as #5623.
Every route has one owner. map coverage checks that each agent-facing route belongs to exactly one family and that a recipe drives it: 193 of 203 do, and 10 are excluded on purpose, each with its reason.
Counts from the map index at the PR head; the groups above are presentation groupings and add up to the 930 IDs.
A living contract 06 / 10
The replay reproduces it and reports an expected failure. The run stays green, and the bug stays visible.
The assertion passes and the replay reports an unexpected pass: known bug no longer reproduces: update the map. The run fails until someone drops the marker and the bullet starts guarding the fix.
If a known-bug bullet fails before its # bug line, the arrange steps broke and the bug was not reproduced. The replay says so.
It already happened once, while the map was being written.
Upstream PR #4952 fixed RemoteWorkspace.get_llm() on provider-linked profiles (#5514). The next replay flagged F20.sdk-get-llm-linked as an unexpected pass, and the map was updated: the bullet now checks that the linked-profile conversation finishes on the connection’s key.
A living contract, not a snapshot. Fixes cannot slip by unnoticed, and the map cannot quietly keep describing a bug that is gone.
How it was built 07 / 10
Agents mapped each family in parallel and proved every recipe live on the real server.
Independent agents re-drove each family on fresh servers, without the authors’ state.
A completeness critic looked for routes, consumers and behaviors no family covered yet.
A pass added explicit # bug assertions, then a full fresh replay of all 34 families.
Each family replays on its own fresh server with the launch flags its file declares: control-agent-server map run --all --fresh --keep-going --record.
The evidence 08 / 10
No failures and no unexpected passes in any of the 34 families. About 135 product bugs are recorded as known-bug bullets: 1 high, 37 medium, 89 low, 8 trivial. The six blocked rows need Docker, a GitHub token, the VS Code binary, or a human ChatGPT login.
A pause during a model call isn’t persisted; a /goal loop still reads “running” after a restart.
After the cipher key changes, writes answer 200 and permanently null stored secrets instead of refusing with 409.
Blank or masked API keys accepted over real ones; an empty MCP server config accepted, so new conversations lose all MCP tools.
One failed run blocks the event loop for about 5 s; a locked profile store can freeze the server for up to 30 s.
Deleted conversations leave sockets silent; Canvas app backends outlive a killed server; the bridge drops the WebSocket subprotocol.
Event search matches only full class paths; per-request usage reports conversation totals; the TypeScript client mangles binary downloads.
Repositories below a path are ignored; filenames with spaces or non-ASCII characters are mangled; large command output is truncated.
31 new issues, 6 existing ones reproduced, one umbrella. Issues #5622–#5652 were filed, six earlier issues got a comment with the map’s reproduction, and all of them are gathered in #5653. The bug tracker lists each one with its severity, status and map IDs.
Counts from the replay report, run on commit f688c8d. Security-relevant findings were reported privately. Four small hardening PRs came out of them: #5657 keeps provider credentials out of completion logs, #5659 ignores empty session API keys, #5660 treats an empty OH_SECRET_KEY as unset, and #5661 keeps sockets and /v1 closed until deferred init completes. A review of community PR #4813 by BSmick6 found five follow-ups, offered as fixes on a branch.
Credit where it belongs 09 / 10
Lauren Tan’s (@poteto) pstack treats verification as infrastructure that other agent workflows can build on.
OpenHands first applied it to the Agent Canvas UI with verify-openhands: a browser-driven map of what a user can do. verify-agent-server applies the same pattern one layer down, to the API that agents and programs drive.
The UI map proves what a person sees after a click. The API map proves what an agent, the SDK or the TypeScript client sees after a request.
Known bugs as executable expected failures, single ownership of every route, and route-table diffs that show which families a change touches.
Open source, reproducible 10 / 10
make build
export PATH="$PWD/.agents/skills/verify-agent-server/scripts:$PATH"
control-agent-server map run --all --fresh
This page describes PR #5621 as checked on 11 October 2026, while it is open. Source links are pinned to its head commit; PR and issue links carry their current status.
← Engel’s Code Design Notebook · Static HTML, real evidence.