What we take from pstack
A decision record, not a survey. After reading Lauren Tan's (@poteto) pstack and her “software factory” write-up, this note pins down what we chose to adopt, what we declined, how to use it, and how it serves the autonomous software factory being discussed for the OpenHands repos.
1The decision
create-verification-skill + maintain-verification-skill, imported into .agents/skills/ (MIT, attributed).2Why verification is the pick
Strip pstack's 44 skills away and one claim carries the rest: the highest-leverage skill is a verification skill — treat it as critical infrastructure, maintain it like infra. The mechanism is the ordinary loop with the human moved off the hot path: observe → act (write code) → re-observe to verify against intended state → loop until done. Because the agent proves its own work, it closes the loop without a person as the bottleneck, and swarms, routines, and auto-repro all compose on that one skill.
We already run a crude version of this: a cabin sibling closes a bounded transpile interval against a projection oracle, gated on CI, reviewed by a human. So this is less a new idea than naming and hardening the thing we improvise. The gap pstack names is that verification should be a first-class, maintained artifact, not re-derived per task.
3What we imported
Two skills, verbatim except our frontmatter convention and a .cursor/skills/ → .agents/skills/ path swap. Both are MIT (Lauren Tan); source is credited in each skill's metadata.
| Skill | What it does |
|---|---|
create-verification-skill | The generator. Interviews a repo (surface, run, drive, observe, isolate) and writes a project-local verify-<app> skill — Launch / Doctor / Drive / Evidence / Cleanup / Helpers — then seeds a feature map and proves itself once before handing over. |
maintain-verification-skill | The upkeep loop. Parallel source-readers per feature, one live pass driving every feature, at most one PR of proven corrections. Keeps the map from rotting as the app changes. |
verify-canvas / verify-agent-server lever is produced later, when we run the generator against a real repo.4How to use it
- Generate. Run
/create-verification-skillagainst a target repo. Out comes.agents/skills/verify-<app>/: a control-app recipe (launch the real app, drive it like a user, capture evidence) plus afeatures/map of the top user-facing capabilities. - Drive with the harness we already have. The lever is a thin, repo-specific recipe; the driving capability is
agent-browser(accessibility-tree / CDP for web, PTY for CLIs, HTTP for services) — already in the toolbox, and the subject of beadsmolpaws-882. - Prove changes with it. An agent (or CI, or you) reruns the same control-app to produce the same before/after evidence. That is the point: it turns “trust me, I clicked around” into “run this.”
- Keep it honest. Run
/maintain-verification-skillon a cadence; it re-reads source per feature, drives every feature live, and ships at most one PR of proven corrections.
npm run dev and check by hand. The lever earns its place when verification recurs and has to be rerunnable by someone else: determinism, reviewability, cold factory agents inheriting one hardened recipe, and evidence landing in a standard place.5Scope: product surfaces vs the SDK
npm run ci. We call that the SDK's verification and don't bolt a control-app onto it.6How it serves the software factory
Graham Neubig proposed making the OpenHands OSS work an autonomous “software factory” and listed four bottlenecks. This is where our pick pays in:
| Factory bottleneck (Graham) | What verification-as-infra gives it |
|---|---|
| 1. Prioritize which issues to work | Little — that's product vision + user/enterprise feedback. Honest gap; pstack doesn't help here. |
| 2. Fully-specify issues (ready-for-dev) | Partial — a control-app + feature map give ready-for-dev a concrete “can this be driven and proven” gate, not just static checks. |
| 3. Auto-review reliability | Direct support — the feature map is reviewer context, and reproducible evidence lets a human or second agent re-check a verdict instead of trusting it. |
| 4. Auto-testing / live before-after evidence | Direct hit — a control-app driving the real app and capturing state is the before/after-evidence structure Graham said is missing. |
7Open questions
Decisions still Engel's to make:
- First target. Generate the first
verify-<app>against which surface — Canvas, the CLI, or the agent-server? - Formalize the coordinator loop. The wake-cabin / launch-sibling / verify / review sequence we run by hand could be written down as one repeatable loop, with
/swarm-style fan-out as a later, small add. - Maintenance cadence. How often
/maintain-verification-skillruns, and whether it hangs off the heartbeat.
#proj-automation; the general read-for-parts analysis this note replaces is preserved in git history.