EN / field notes

OpenHands / Agent Server verification 01 / 10

A map that keeps
itself honest.

Drive the real server. Know what you should see. Tell a stale instruction from a real bug.

verify-agent-server lets an agent drive the OpenHands Agent Server, the API that agents and programs use, feature by feature, through the same routes and clients they use. It is built by agents, for agents, and works at their speed.

The verify-agent-server story · Engel Nyst
Software Agent SDK PR #5621 · results from the full replay on 9 October 2026

Read it as a story
A 2½-minute film, animated with John Lasseter’s principles of animation, with a generated score. Its storyboard and source are open.
  1. 01 Launch a fresh server
  2. 02 Drive the real API
  3. 03 Check what it answers
  4. 04 Keep the evidence

The problem 02 / 10

203 routes.
How do you know it all works?

The Agent Server is what agents and programs drive: about 190 REST and WebSocket routes on the default app, 203 with Docker runtime mode, plus webhooks, telemetry, and the Python SDK and TypeScript clients that consume them.

Unit tests

Check handlers and models in isolation, with no live server behind them.

OpenAPI breakage check

Catches changes to the published contract, not changes in behavior.

Live-server tests

Exercise a real server, for a selection of paths.

Each covered part of it. None gave an agent a single place to drive the real server end to end, know what it should observe, and tell a stale instruction from a real product bug.

The design 03 / 10

A lever, a map,
and a way to keep it true.

A contributor skill under .agents/skills/verify-agent-server/, adapted from Lauren Tan’s pstack verification skills, which OpenHands already used for its Agent Canvas UI.

01

The control CLI

How do I drive this checkout? Isolated servers, any REST route or WebSocket with assertions, fixtures for everything around the API, and evidence with secrets redacted.

control-agent-server
02

The executable feature map

What should happen? 34 families, 930 stable sub-feature IDs, and every agent-facing route owned by exactly one family. Each bullet carries a shell recipe that proves it on a fresh server.

references/feature-map/
03

Known bugs as expected failures

What is broken today? A known bug’s failing assertion is marked # bug. The replay expects it to fail, and flags the day it stops failing.

# bug → xfail / xpass
04

Maintenance levers

How does it stay true? Route coverage, recipe parse checks, route-table diffs between commits, owner lookup for changed files, and a cross-repo test guarding the structure.

map coverage · check · diff · owners

The lever 04 / 10

Every step is a command
the next agent can rerun.

control-agent-server launches this checkout as an isolated server: its own HOME, persistence directory, keys and port.

It reaches every route through api and ws, with --expect for status codes and --check for what comes back. It arranges what the API needs around it: git repositories, skills, plugins, MCP servers, webhook sinks, and a scripted LLM stub.

Model-backed recipes run on DeepSeek, an open-weights model, through the server’s own pre-flight. Every command prints one JSON object, and saved exchanges have their secrets redacted.

From a Software Agent SDK checkout
export PATH="$PWD/.agents/skills/verify-agent-server/scripts:$PATH"
export AGENT_SERVER_VERIFY_RUN=$(control-agent-server launch --new --print-run)
control-agent-server doctor

control-agent-server api PUT /api/settings/secrets \
  --json '{"name": "QA_F18_TOKEN", "value": "**********"}' \
  --expect 200,422
control-agent-server api GET /api/settings/secrets/QA_F18_TOKEN \
  --expect 200 --check . eq "$QA_F18_SECRET"   # bug

No untracked one-off scripts. If a path cannot be driven with the CLI, that is a harness gap: extend the CLI, prove it live, then write the recipe, so the next agent can rerun it.

Adapted from SKILL.md and the F18.secret-empty-value recipe at the PR head.

An executable inventory 05 / 10

The unit is a behavior
an API consumer can observe.

feature families
34
sub-feature IDs
930
routes owned and driven
193 / 203
Server and sessions 52

Server status, auth and sessions, deferred init (warm pool). F01–F03.

Conversations 216

Lifecycle; run, pause and interrupt; events; the session socket; confirmation, security and secrets; model switching and plugins; fork, navigate, condense and ask; goals. F04–F11.

Workspace and runtime 167

Runtime and credentials, bash, files, git, workspace serving and VS Code, the workspaces registry. F12–F17.

Settings, models and profiles 194

Settings and secrets, MCP, LLM profiles, agent profiles, meta profiles, the LLM catalog and connections, OpenAI subscription. F18–F24.

Extensions 147

Skills, plugins, hooks, subagents and tools, Canvas extensions, the app backend bridge. F25–F29.

Gateway, persistence and delivery 79

The OpenAI-compatible gateway, persistence and limits, webhooks and telemetry. F30–F32.

Clients 75

The Python SDK remote client and the TypeScript client. F33–F34.

F18-settings-and-secrets.md · one bullet
- `F18.secret-empty-value`: a PUT whose value
  is empty or the redaction placeholder must
  not leave a listed but unreadable secret or
  wipe an existing value.

# The recipe, run on a fresh server:
GET  …/secrets/QA_F18_TOKEN  --check . eq …  # control
PUT  /api/settings/secrets  value "**********"
GET  …/secrets/QA_F18_TOKEN  --check . eq …  # bug

# Today: 200 on the PUT, then 404
# "Secret not found". Filed as #5623.

Every route has one owner. map coverage checks that each agent-facing route belongs to exactly one family and that a recipe drives it: 193 of 203 do, and 10 are excluded on purpose, each with its reason.

Counts from the map index at the PR head; the groups above are presentation groupings and add up to the 930 IDs.

A living contract 06 / 10

A bug the map knows about
fails on purpose, until it doesn’t.

While the bug lasts: xfail

The replay reproduces it and reports an expected failure. The run stays green, and the bug stays visible.

The day it is fixed: xpass

The assertion passes and the replay reports an unexpected pass: known bug no longer reproduces: update the map. The run fails until someone drops the marker and the bullet starts guarding the fix.

Outside the marker: a real failure

If a known-bug bullet fails before its # bug line, the arrange steps broke and the bug was not reproduced. The replay says so.

It already happened once, while the map was being written.

Upstream PR #4952 fixed RemoteWorkspace.get_llm() on provider-linked profiles (#5514). The next replay flagged F20.sdk-get-llm-linked as an unexpected pass, and the map was updated: the bullet now checks that the linked-profile conversation finishes on the connection’s key.

xfail · #5514→ xpasspass · map updated

A living contract, not a snapshot. Fixes cannot slip by unnoticed, and the map cannot quietly keep describing a bug that is gone.

How it was built 07 / 10

A full agentic run,
checked by agents.

  1. 01

    Authors

    Agents mapped each family in parallel and proved every recipe live on the real server.

  2. 02

    Reviewers

    Independent agents re-drove each family on fresh servers, without the authors’ state.

  3. 03

    A critic

    A completeness critic looked for routes, consumers and behaviors no family covered yet.

  4. 04

    Hardening

    A pass added explicit # bug assertions, then a full fresh replay of all 34 families.

Each family replays on its own fresh server with the launch flags its file declares: control-agent-server map run --all --fresh --keep-going --record.

The evidence 08 / 10

771 pass. 156 known bugs reproduced.
Zero surprises.

pass
771
expected failures: known bugs reproduced
156
blocked
6

No failures and no unexpected passes in any of the 34 families. About 135 product bugs are recorded as known-bug bullets: 1 high, 37 medium, 89 low, 8 trivial. The six blocked rows need Docker, a GitHub token, the VS Code binary, or a human ChatGPT login.

State that isn’t saved, or comes back stale

A pause during a model call isn’t persisted; a /goal loop still reads “running” after a restart.

Data loss on key change (the one high)

After the cipher key changes, writes answer 200 and permanently null stored secrets instead of refusing with 409.

Validation gaps

Blank or masked API keys accepted over real ones; an empty MCP server config accepted, so new conversations lose all MCP tools.

Responsiveness

One failed run blocks the event loop for about 5 s; a locked profile store can freeze the server for up to 30 s.

WebSocket and process lifecycle

Deleted conversations leave sockets silent; Canvas app backends outlive a killed server; the bridge drops the WebSocket subprotocol.

Server and client contracts

Event search matches only full class paths; per-request usage reports conversation totals; the TypeScript client mangles binary downloads.

Git and file edge cases

Repositories below a path are ignored; filenames with spaces or non-ASCII characters are mangled; large command output is truncated.

31 new issues, 6 existing ones reproduced, one umbrella. Issues #5622–#5652 were filed, six earlier issues got a comment with the map’s reproduction, and all of them are gathered in #5653. The bug tracker lists each one with its severity, status and map IDs.

Counts from the replay report, run on commit f688c8d. Security-relevant findings were reported privately. Four small hardening PRs came out of them: #5657 keeps provider credentials out of completion logs, #5659 ignores empty session API keys, #5660 treats an empty OH_SECRET_KEY as unset, and #5661 keeps sockets and /v1 closed until deferred init completes. A review of community PR #4813 by BSmick6 found five follow-ups, offered as fixes on a branch.

Credit where it belongs 09 / 10

The idea came from pstack.

Lauren Tan’s (@poteto) pstack treats verification as infrastructure that other agent workflows can build on.

OpenHands first applied it to the Agent Canvas UI with verify-openhands: a browser-driven map of what a user can do. verify-agent-server applies the same pattern one layer down, to the API that agents and programs drive.

Same pattern, different consumer

The UI map proves what a person sees after a click. The API map proves what an agent, the SDK or the TypeScript client sees after a request.

What this map adds

Known bugs as executable expected failures, single ownership of every route, and route-table diffs that show which families a change touches.

Open source, reproducible 10 / 10

Every feature and every known bug,
replayed on a fresh server.

From a Software Agent SDK checkout with the PR
make build
export PATH="$PWD/.agents/skills/verify-agent-server/scripts:$PATH"
control-agent-server map run --all --fresh

This page describes PR #5621 as checked on 11 October 2026, while it is open. Source links are pinned to its head commit; PR and issue links carry their current status.

← Engel’s Code Design Notebook · Static HTML, real evidence.