EN / field notes

OpenHands / verification as infrastructure 01 / 12

Give agents a way
to prove it.

A map of the product. A way to run it. A standard of proof.

We taught agents how to use OpenHands itself: open the app, take the paths a user takes, check what actually happened, and leave evidence someone else can replay.

The verify-openhands story · Engel Nyst
Updated 9 October 2026 · implementation and evidence snapshot

See how the pieces fit
Agent Canvas · real verification run
Actual F16 Application-settings evidence from the feature-map PR. Screenshots throughout this page come from the original runs.
  1. 01 Find the behavior
  2. 02 Drive the product
  3. 03 Observe the result
  4. 04 Keep the proof

The design 02 / 12

“Test it like a user”
needs an operating manual.

The repository now carries the knowledge an agent needs to verify its own work.

01

The feature map

What can a user do? Stable behavior IDs, entry points, preconditions, exact recipes, expected observations, and known traps.

references/feature-map/
02

The control CLI

How do I run and use this checkout? Isolated stacks, a persistent browser, reusable actions, diagnostics, and evidence capture.

control-openhands
03

The verification rules

What counts as proof? Exercise real entry points. Re-read state after a mutation. Record each outcome. Make the recipe replayable.

SKILL.md + mapping.md
04

The maintenance rules

How does this stay true? Compare fixed revisions, read the intent of changes, drive affected behavior, and separate map drift from product bugs.

maintenance.md + report.md

Merged in #17961; maintained in #18085. The existing CLI deep dive covers the command surface in more detail.

A readable inventory 03 / 12

The unit is a behavior.

“A secret is still there after a reload” gives an agent something concrete to prove.

feature families
27
mapped behaviors
743
routes inventoried
30
Starting work 127

First run, app shell, Home, conversations and folders. F01–F04.

Conversation & workspace 148

Composer, agent activity, header controls, files, diffs, terminal, browser and planner. F05–F08, F27.

Settings & configuration 178

Providers, LLM profiles, router, agent profiles, secrets, agent behavior and application preferences. F09–F16.

Customization 101

MCP, skills, plugins and Canvas apps. F17–F20.

Automations 128

Dashboard, creation, detail, runs, logs and Git Sync. F21–F24.

Backends & runtimes 61

Backend selection, Cloud, sharing, launcher, Docker, desktop and library use. F25–F26.

F14-secrets.md · condensed anatomy
# F14 — Secrets
Source: src/routes/secrets-settings.tsx, …

## Sub-features
F14.create: adding a secret persists it;
it is listed after a reload.

## How to get to it (user POV)
Settings → Secrets · direct URL · command menu

## Driving it with control-openhands
Preconditions → commands → expected observation

## Gotchas
Real traps and linked known failures

Mapped does not mean passed. The map keeps failures, blocked prerequisites, and gaps visible. Cloud accounts, external integrations and platform-specific behavior need their own available environment.

Counts at merged maintenance revision 8793c11, checked 9 October. The original 716 behaviors have grown to 743 across the same 27 families. The groups above are presentation groupings; the linked index owns the inventory.

The lever 04 / 12

One checkout. One isolated run.
One browser that remembers.

control-openhands wraps the production launcher and the repo’s own Playwright.

Each run gets its own HOME, state, keys and ports. The browser stays alive across shell calls, continuously collecting errors, requests, toasts and downloads.

doctor checks that the services are healthy and the browser is looking at this checkout’s build. A stale bundle cannot quietly become evidence for a new fix.

From an installed OpenHands checkout
export PATH="$PWD/.agents/skills/verify-openhands/scripts:$PATH"
export OH_VERIFY_RUN=$(control-openhands launch --new --print-run)

control-openhands doctor
control-openhands onboard --skip
control-openhands browser goto /settings/secrets
control-openhands browser snapshot
control-openhands browser errors

# After verification:
control-openhands stop

Arrange

API calls and fixtures create the preconditions: a repository, a profile, an automation, a backend that goes down.

Drive

Click, type, navigate, press keys, resize, upload and download through real product entry points.

Observe

Read the UI and accessibility tree, inspect network and toast history, check persisted state, and save artifacts.

The CLI returns structured results with actionable failures. Read the launch, doctor, drive and cleanup contract.

A recipe you can replay 05 / 12

Example: save it, reload it,
find it again.

F14.create verifies persistence through the UI. Opening the form is only the beginning.

After launch, doctor and onboarding
control-openhands browser goto /settings/secrets
control-openhands browser click 'testid=add-secret-button'
control-openhands browser fill 'testid=add-secret-form >> testid=name-input' QA_TMP_SECRET
control-openhands browser fill 'testid=add-secret-form >> testid=value-input' dummy-value-123
control-openhands browser click 'testid=add-secret-form >> testid=submit-button'

A dummy value in a disposable run. These steps exercise the actual form and Save action.

The second view proves persistence
control-openhands browser reload
control-openhands browser count 'testid=secret-item >> has-text=QA_TMP_SECRET'

# Expected observation: count 1 after reload.
# If that is not what happened, do not record a pass.

control-openhands browser screenshot --feature F14.create --name after-reload

Reloading asks the product for its saved state again. A success toast alone cannot establish that.

Only after observing the expected result
control-openhands evidence add --feature F14.create --result pass \
  --entry "Settings > Secrets > Add" \
  --expected "row survives reload" --actual "count 1 after reload" \
  --artifact evidence/F14.create/after-reload.png
control-openhands evidence report
control-openhands stop

The evidence survives cleanup. Use the actual returned artifact path if the screenshot name received a suffix.

  1. Setup is not proof.

    API writes may arrange a test. A UI claim still needs UI evidence.

  2. One pass has a precise scope.

    The ledger records an ID and entry point: pass, fail, blocked or not-run.

  3. The result must be observable.

    Inspect saved state, real files or tool observations. Model wording cannot stand in for those checks.

Adapted from the actual F14 Secrets recipe and evidence contract. Run in order on a fresh isolated stack.

Evidence / 01 06 / 12

Example: the Close button
was off the phone.

A perfectly ordinary user action exposed a real layout failure.

Open Settings on a 390px viewport. Open the Agent Canvas update dialog. The 520px dialog spills past both edges; its title is cut off and its Close button is outside the screen.

The agent reproduced it, captured the before state, applied the shared viewport cap, and checked the fixed build at phone, narrow and desktop widths.

Merged · 5 Oct PR #18008 · fixes #17903

With the update dialog open
control-openhands browser viewport phone
control-openhands browser bbox \
  'testid=agent-canvas-update-modal'
control-openhands browser click \
  'testid=close-agent-canvas-update-modal'
control-openhands browser count \
  'testid=agent-canvas-update-modal'
# Expected: 0 after Close.

The PR reports x = −65 / width 520 before, x = 19.5 / width 351 after. The screenshot is paired with geometry and dismissal checks.

Before · 390 × 844
Clipped title.
No reachable Close button.
After · 390 × 844
Full dialog in view.
Close works.

Original, unmodified captures from #18008. The initial verification effort and its 6 October refresh filed 68 issues. This was one of dozens of fixes that landed while an agent kept updating the still-open feature-map PR.

Evidence / 02 07 / 12

Example: some failures disappear
before the next screenshot.

One failed “Run now” action showed two error toasts.

The recipe creates an inert automation, loads its card, then deletes it through the API to make the UI stale. Clicking Run now takes the real error path.

Both the caller and a global error handler displayed a toast. browser toasts --history preserved the evidence. The fix leaves one useful error message.

#18043 · merged 6 Oct · fixes #17906

Example: keyboard behavior needs keyboard proof.

#18011 fixed Delete automation: Escape did nothing, focus stayed outside, and assistive technology had no named dialog. The verification used key presses and the accessibility tree.

Example: visible content needs a fresh read.

#18032 fixed a stale Files preview after an agent edited a file. Even Refresh and a page reload had continued to show old contents.

Before #18043: one failed Run now action displays both “Automation not found” and a generic 404 toast.

Original captures from separate before/after stacks; fixture lists differ. The comparison is the number and content of error toasts. Click the image to inspect it.

These PRs explicitly connect their fixes to the full agentic run. Screenshots support the claim; the reproduction and observation steps make it checkable.

Why the structure matters 08 / 12

Every new agent inherits
what the last one learned.

01

Less product rediscovery

The index points to the relevant family. A bounded task can load its recipes and gotchas without reconstructing the whole application.

02

Less environment guesswork

Launch, health checks and build identity are shared infrastructure. Fixing the harness once helps every later run.

03

Continuity across tool calls

The persistent browser retains the page and observations between short commands. Transient errors remain available to inspect.

04

Parallel work with separate state

Each live worker gets its own stack, ports, keys and evidence. Source readers can work in parallel without sharing one browser session.

05

A smaller gap between claim and proof

Stable IDs, expected observations and a second view make “verified” specific. Failures and missing prerequisites have explicit outcomes.

06

Reviewers can replay the work

The next agent can run the same recipe against the actual fix. A summary becomes an inspectable record of behavior.

It was used beyond the first sweep. The 5 October fix-fleet report records 36 Canvas fixers using the CLI; across the wider fix/review fleet, 170 isolated launches and 261 screenshots. These are usage counts, not a measured speedup or a correctness guarantee.

The OpenHands adaptation explains the persistent browser and isolation. Unit tests and CI remain part of verification; this adds the live user experience.

The part that keeps it useful 09 / 12

The map changes with the product.

Maintenance asks three separate questions: is the feature present, does it work and look right, and was the change intended?

  1. 01

    Freeze revisions

    Pin BASE and TARGET. Comparable runs need known code and separate state.

  2. 02

    Read the changes

    Map changed paths to behavior IDs. Read PR and issue intent. Shared code widens the scope.

  3. 03

    Drive it live

    Replay recipes and known failures. Compare BASE and TARGET before attributing a regression.

  4. 04

    Reconcile

    Keep evidence, fix the right layer, and propose a reviewed baseline only after a completed pass.

Map drift → correct the recipe.

A renamed control or intended interaction change belongs in the map, backed by live evidence.

Harness gap → extend the shared CLI.

Prove the new interaction before other recipes depend on it.

Product defect → report the bug.

A broken behavior is not a reason to redefine success. Route it to Canvas, SDK or automation as appropriate.

Missing prerequisite → record blocked.

Name the account, platform or service needed. Partial and blocked passes do not advance the baseline.

Intent and runtime are independent. “Documented change” does not mean “works.” A merged fix still needs to be exercised; known failures get replayed.

The merged maintenance rules define the comparison, intent ledger, live pass and baseline policy.

A living system 10 / 12

The follow-ups close the loop.

Product fixes

68 issues filed. Dozens of fixes already landing.

The initial verification effort and its 6 October refresh file 68 issues across Canvas, the SDK and automation. The map’s update history names 32 bug-fix PRs incorporated before the feature-map PR itself merges.

An agent keeps the still-open PR current as fixes land: merge in main, commit updated recipes and expected results, rerun each changed recipe on a fresh stack, and have a second agent check it.

Four update commits record the loop: 3 fixes → 17 more → 2 more → 10 more.

Foundation merged

The map, CLI and rules land together.

#17961 arrives with the map already revised repeatedly against fixes it helped uncover. The PR records 33 of the 68 issues closed as fixed by 6 October; that is an issue count, separate from the 32 named fix PRs above.

Maintenance merged

New fixes change the expected behavior.

#18085 updates hooks, install-error recipes, preset metadata and credential-hidden Git Sync. Changed recipes are run on a fresh stack and independently replayed. The pass also finds new issues: #18088, #18089, #18090.

Merged

Daily passes, and 20 more behaviors.

#18093 adds affected-family discovery, baseline and test-ID helpers, richer reports, and smoke checks for untouched families. #18105 turns 20 gaps found in the E2E comparison into live recipes: 716 → 736 behaviors.

Merged

The harness improves through use.

#18162 replaces recipe workarounds with shared browser and fixture commands. #18191 drives model-dependent cases with real DeepSeek calls: 29 of 31 checks pass; two expose known failures. Four new behaviors bring the map to 740.

Daily automation now produces maintenance PRs. Its first merged report records a memory constraint that blocked the live run. Recording that limit is part of the result.

Credit where it belongs 11 / 12

The idea came from pstack.

Lauren Tan’s pstack treats verification as infrastructure that other agent workflows can build on.

Its create-verification-skill and maintain-verification-skill turn a repository into an executable guide to its own user-facing behavior.

We applied that pattern to OpenHands: the real launcher, isolated stacks, a persistent browser, a detailed feature map, and rules for evidence and change intent.

The reusable idea is to put product knowledge and a means of checking it inside the repository, where every agent can find them.

“A generated skill that was never executed is a draft, not a deliverable.”

The OpenHands attribution and adaptation notes record what was borrowed and what was built for this product. See also our earlier pstack decision note.

Take it with you 12 / 12

A fix should leave behind
a better way to verify the next one.

That is what we meant to build: agents that can use the product, discover what is wrong, prove a repair, and improve the shared instructions.

This presentation describes the public work as checked on 9 October 2026. Source links are pinned where practical; PR links carry current status. The embedded captures are historical test artifacts, not fresh verification of today’s build.

← Engel’s Code Design Notebook · Static HTML, real evidence.