The feature map
What can a user do? Stable behavior IDs, entry points, preconditions, exact recipes, expected observations, and known traps.
references/feature-map/OpenHands / verification as infrastructure 01 / 12
A map of the product. A way to run it. A standard of proof.
We taught agents how to use OpenHands itself: open the app, take the paths a user takes, check what actually happened, and leave evidence someone else can replay.
The verify-openhands story · Engel Nyst
Updated 9 October 2026 · implementation and evidence snapshot
The design 02 / 12
The repository now carries the knowledge an agent needs to verify its own work.
What can a user do? Stable behavior IDs, entry points, preconditions, exact recipes, expected observations, and known traps.
references/feature-map/How do I run and use this checkout? Isolated stacks, a persistent browser, reusable actions, diagnostics, and evidence capture.
control-openhandsWhat counts as proof? Exercise real entry points. Re-read state after a mutation. Record each outcome. Make the recipe replayable.
SKILL.md + mapping.mdHow does this stay true? Compare fixed revisions, read the intent of changes, drive affected behavior, and separate map drift from product bugs.
maintenance.md + report.mdMerged in #17961; maintained in #18085. The existing CLI deep dive covers the command surface in more detail.
A readable inventory 03 / 12
“A secret is still there after a reload” gives an agent something concrete to prove.
First run, app shell, Home, conversations and folders. F01–F04.
Composer, agent activity, header controls, files, diffs, terminal, browser and planner. F05–F08, F27.
Providers, LLM profiles, router, agent profiles, secrets, agent behavior and application preferences. F09–F16.
MCP, skills, plugins and Canvas apps. F17–F20.
Dashboard, creation, detail, runs, logs and Git Sync. F21–F24.
Backend selection, Cloud, sharing, launcher, Docker, desktop and library use. F25–F26.
# F14 — Secrets
Source: src/routes/secrets-settings.tsx, …
## Sub-features
F14.create: adding a secret persists it;
it is listed after a reload.
## How to get to it (user POV)
Settings → Secrets · direct URL · command menu
## Driving it with control-openhands
Preconditions → commands → expected observation
## Gotchas
Real traps and linked known failures
Mapped does not mean passed. The map keeps failures, blocked prerequisites, and gaps visible. Cloud accounts, external integrations and platform-specific behavior need their own available environment.
Counts at merged maintenance revision 8793c11, checked 9 October. The original 716 behaviors have grown to 743 across the same 27 families. The groups above are presentation groupings; the linked index owns the inventory.
The lever 04 / 12
control-openhands wraps the production launcher and the repo’s own Playwright.
Each run gets its own HOME, state, keys and ports. The browser stays alive across shell calls, continuously collecting errors, requests, toasts and downloads.
doctor checks that the services are healthy and the browser is looking at this checkout’s build. A stale bundle cannot quietly become evidence for a new fix.
export PATH="$PWD/.agents/skills/verify-openhands/scripts:$PATH"
export OH_VERIFY_RUN=$(control-openhands launch --new --print-run)
control-openhands doctor
control-openhands onboard --skip
control-openhands browser goto /settings/secrets
control-openhands browser snapshot
control-openhands browser errors
# After verification:
control-openhands stop
API calls and fixtures create the preconditions: a repository, a profile, an automation, a backend that goes down.
Click, type, navigate, press keys, resize, upload and download through real product entry points.
Read the UI and accessibility tree, inspect network and toast history, check persisted state, and save artifacts.
The CLI returns structured results with actionable failures. Read the launch, doctor, drive and cleanup contract.
A recipe you can replay 05 / 12
F14.create verifies persistence through the UI. Opening the form is only the beginning.
control-openhands browser goto /settings/secrets
control-openhands browser click 'testid=add-secret-button'
control-openhands browser fill 'testid=add-secret-form >> testid=name-input' QA_TMP_SECRET
control-openhands browser fill 'testid=add-secret-form >> testid=value-input' dummy-value-123
control-openhands browser click 'testid=add-secret-form >> testid=submit-button'A dummy value in a disposable run. These steps exercise the actual form and Save action.
control-openhands browser reload
control-openhands browser count 'testid=secret-item >> has-text=QA_TMP_SECRET'
# Expected observation: count 1 after reload.
# If that is not what happened, do not record a pass.
control-openhands browser screenshot --feature F14.create --name after-reloadReloading asks the product for its saved state again. A success toast alone cannot establish that.
control-openhands evidence add --feature F14.create --result pass \
--entry "Settings > Secrets > Add" \
--expected "row survives reload" --actual "count 1 after reload" \
--artifact evidence/F14.create/after-reload.png
control-openhands evidence report
control-openhands stopThe evidence survives cleanup. Use the actual returned artifact path if the screenshot name received a suffix.
API writes may arrange a test. A UI claim still needs UI evidence.
The ledger records an ID and entry point: pass, fail, blocked or not-run.
Inspect saved state, real files or tool observations. Model wording cannot stand in for those checks.
Adapted from the actual F14 Secrets recipe and evidence contract. Run in order on a fresh isolated stack.
Evidence / 01 06 / 12
A perfectly ordinary user action exposed a real layout failure.
Open Settings on a 390px viewport. Open the Agent Canvas update dialog. The 520px dialog spills past both edges; its title is cut off and its Close button is outside the screen.
The agent reproduced it, captured the before state, applied the shared viewport cap, and checked the fixed build at phone, narrow and desktop widths.
Merged · 5 Oct PR #18008 · fixes #17903
control-openhands browser viewport phone
control-openhands browser bbox \
'testid=agent-canvas-update-modal'
control-openhands browser click \
'testid=close-agent-canvas-update-modal'
control-openhands browser count \
'testid=agent-canvas-update-modal'
# Expected: 0 after Close.The PR reports x = −65 / width 520 before, x = 19.5 / width 351 after. The screenshot is paired with geometry and dismissal checks.
Original, unmodified captures from #18008. The initial verification effort and its 6 October refresh filed 68 issues. This was one of dozens of fixes that landed while an agent kept updating the still-open feature-map PR.
Evidence / 02 07 / 12
One failed “Run now” action showed two error toasts.
The recipe creates an inert automation, loads its card, then deletes it through the API to make the UI stale. Clicking Run now takes the real error path.
Both the caller and a global error handler displayed a toast. browser toasts --history preserved the evidence. The fix leaves one useful error message.
#18043 · merged 6 Oct · fixes #17906
#18011 fixed Delete automation: Escape did nothing, focus stayed outside, and assistive technology had no named dialog. The verification used key presses and the accessibility tree.
#18032 fixed a stale Files preview after an agent edited a file. Even Refresh and a page reload had continued to show old contents.
Original captures from separate before/after stacks; fixture lists differ. The comparison is the number and content of error toasts. Click the image to inspect it.
These PRs explicitly connect their fixes to the full agentic run. Screenshots support the claim; the reproduction and observation steps make it checkable.
Why the structure matters 08 / 12
The index points to the relevant family. A bounded task can load its recipes and gotchas without reconstructing the whole application.
Launch, health checks and build identity are shared infrastructure. Fixing the harness once helps every later run.
The persistent browser retains the page and observations between short commands. Transient errors remain available to inspect.
Each live worker gets its own stack, ports, keys and evidence. Source readers can work in parallel without sharing one browser session.
Stable IDs, expected observations and a second view make “verified” specific. Failures and missing prerequisites have explicit outcomes.
The next agent can run the same recipe against the actual fix. A summary becomes an inspectable record of behavior.
It was used beyond the first sweep. The 5 October fix-fleet report records 36 Canvas fixers using the CLI; across the wider fix/review fleet, 170 isolated launches and 261 screenshots. These are usage counts, not a measured speedup or a correctness guarantee.
The OpenHands adaptation explains the persistent browser and isolation. Unit tests and CI remain part of verification; this adds the live user experience.
The part that keeps it useful 09 / 12
Maintenance asks three separate questions: is the feature present, does it work and look right, and was the change intended?
Pin BASE and TARGET. Comparable runs need known code and separate state.
Map changed paths to behavior IDs. Read PR and issue intent. Shared code widens the scope.
Replay recipes and known failures. Compare BASE and TARGET before attributing a regression.
Keep evidence, fix the right layer, and propose a reviewed baseline only after a completed pass.
A renamed control or intended interaction change belongs in the map, backed by live evidence.
Prove the new interaction before other recipes depend on it.
A broken behavior is not a reason to redefine success. Route it to Canvas, SDK or automation as appropriate.
Name the account, platform or service needed. Partial and blocked passes do not advance the baseline.
Intent and runtime are independent. “Documented change” does not mean “works.” A merged fix still needs to be exercised; known failures get replayed.
The merged maintenance rules define the comparison, intent ledger, live pass and baseline policy.
A living system 10 / 12
The initial verification effort and its 6 October refresh file 68 issues across Canvas, the SDK and automation. The map’s update history names 32 bug-fix PRs incorporated before the feature-map PR itself merges.
An agent keeps the still-open PR current as fixes land: merge in main, commit updated recipes and expected results, rerun each changed recipe on a fresh stack, and have a second agent check it.
Four update commits record the loop: 3 fixes → 17 more → 2 more → 10 more.
#18162 replaces recipe workarounds with shared browser and fixture commands. #18191 drives model-dependent cases with real DeepSeek calls: 29 of 31 checks pass; two expose known failures. Four new behaviors bring the map to 740.
Daily automation now produces maintenance PRs. Its first merged report records a memory constraint that blocked the live run. Recording that limit is part of the result.
#18211 adds browser video recording. #18209 repairs model-profile setup. #18207 replays the merged Hooks fix and finds a remaining mismatch.
The next merged maintenance report updates recipes and adds three behaviors: 743 now mapped. The pass is partial; the accepted baseline remains 6 October.
#18213 reports a run against an earlier 736-behavior revision: 534 pass, 40 fail, 162 blocked. It files four new issues and reproduces older ones. These are the results for that pinned run; they do not establish a clean pass of the current 743-behavior map.
Read the per-behavior evidence · All findings, current PR status and visual walkthroughs
Credit where it belongs 11 / 12
Lauren Tan’s pstack treats verification as infrastructure that other agent workflows can build on.
Its create-verification-skill and maintain-verification-skill turn a repository into an executable guide to its own user-facing behavior.
We applied that pattern to OpenHands: the real launcher, isolated stacks, a persistent browser, a detailed feature map, and rules for evidence and change intent.
The reusable idea is to put product knowledge and a means of checking it inside the repository, where every agent can find them.
“A generated skill that was never executed is a draft, not a deliverable.”
Lauren’s example feature map shows the structure using a fictional application. The maintenance skill explains how to keep it useful.
@poteto’s verification-skills announcement
Matt Pocock’s conversation with Poteto ↗
The OpenHands attribution and adaptation notes record what was borrowed and what was built for this product. See also our earlier pstack decision note.
Take it with you 12 / 12
That is what we meant to build: agents that can use the product, discover what is wrong, prove a repair, and improve the shared instructions.
This presentation describes the public work as checked on 9 October 2026. Source links are pinned where practical; PR links carry current status. The embedded captures are historical test artifacts, not fresh verification of today’s build.
← Engel’s Code Design Notebook · Static HTML, real evidence.