decision · pstack → SmolPaws / OpenHands

What we take from pstack

A decision record, not a survey. After reading Lauren Tan's (@poteto) pstack and her “software factory” write-up, this note pins down what we chose to adopt, what we declined, how to use it, and how it serves the autonomous software factory being discussed for the OpenHands repos.

1The decision

Take: verification as infrastructureThe keystone idea. A self-verifying loop that takes the human off the critical path, that everything else composes on.
Take: the two generator skillscreate-verification-skill + maintain-verification-skill, imported into .agents/skills/ (MIT, attributed).
Take: feature map + control-app CLIFor the product surfaces (Canvas, CLI, agent-server) — a user-POV capability map, each entry with a rerunnable drive/proof.
Decline: the 44-skill catalogAgainst “stay small.” We take the pattern and a few principles, not the plugin. Dr Eggbot (an xAI bot, not a repo skill) is skipped too.
ⓘ One line: adopt verification-as-infra and pstack's two verification-skill generators; apply the feature map + control-app CLI to product surfaces; leave the SDK's projection oracle as its own lever; port nothing else.

2Why verification is the pick

Strip pstack's 44 skills away and one claim carries the rest: the highest-leverage skill is a verification skill — treat it as critical infrastructure, maintain it like infra. The mechanism is the ordinary loop with the human moved off the hot path: observe → act (write code) → re-observe to verify against intended state → loop until done. Because the agent proves its own work, it closes the loop without a person as the bottleneck, and swarms, routines, and auto-repro all compose on that one skill.

We already run a crude version of this: a cabin sibling closes a bounded transpile interval against a projection oracle, gated on CI, reviewed by a human. So this is less a new idea than naming and hardening the thing we improvise. The gap pstack names is that verification should be a first-class, maintained artifact, not re-derived per task.

3What we imported

Two skills, verbatim except our frontmatter convention and a .cursor/skills/ → .agents/skills/ path swap. Both are MIT (Lauren Tan); source is credited in each skill's metadata.

SkillWhat it does
create-verification-skillThe generator. Interviews a repo (surface, run, drive, observe, isolate) and writes a project-local verify-<app> skill — Launch / Doctor / Drive / Evidence / Cleanup / Helpers — then seeds a feature map and proves itself once before handing over.
maintain-verification-skillThe upkeep loop. Parallel source-readers per feature, one live pass driving every feature, at most one PR of proven corrections. Keeps the map from rotting as the app changes.
★ Key point so we don't over-build: these generate a repo-specific control-app CLI by interviewing the codebase. pstack does not ship, and we did not import, a generic CLI. The actual verify-canvas / verify-agent-server lever is produced later, when we run the generator against a real repo.

4How to use it

  1. Generate. Run /create-verification-skill against a target repo. Out comes .agents/skills/verify-<app>/: a control-app recipe (launch the real app, drive it like a user, capture evidence) plus a features/ map of the top user-facing capabilities.
  2. Drive with the harness we already have. The lever is a thin, repo-specific recipe; the driving capability is agent-browser (accessibility-tree / CDP for web, PTY for CLIs, HTTP for services) — already in the toolbox, and the subject of bead smolpaws-882.
  3. Prove changes with it. An agent (or CI, or you) reruns the same control-app to produce the same before/after evidence. That is the point: it turns “trust me, I clicked around” into “run this.”
  4. Keep it honest. Run /maintain-verification-skill on a cadence; it re-reads source per feature, drives every feature live, and ships at most one PR of proven corrections.
ⓘ When the CLI is not worth it: a genuine one-off. pstack itself says skip the lever when the task is trivial — an agent can just npm run dev and check by hand. The lever earns its place when verification recurs and has to be rerunnable by someone else: determinism, reviewability, cold factory agents inheriting one hardened recipe, and evidence landing in a standard place.

5Scope: product surfaces vs the SDK

Product surfaces → take itCanvas, the CLI, the agent-server runtime. Real users (or user-shaped callers) touch these, so a user-POV feature map + control-app CLI fit directly. This is where the “turn the E2E suite into a feature map” move lands: same coverage, but a readable, maintainable capability index, each entry pointing at its drive and proof.
vs
The SDK → leave its own leverA library's “user” is code, so a control-app CLI is a weak fit and a feature map would just restate the API. The SDK already has its rerunnable lever: the drift / projection oracle + npm run ci. We call that the SDK's verification and don't bolt a control-app onto it.

6How it serves the software factory

Graham Neubig proposed making the OpenHands OSS work an autonomous “software factory” and listed four bottlenecks. This is where our pick pays in:

Factory bottleneck (Graham)What verification-as-infra gives it
1. Prioritize which issues to workLittle — that's product vision + user/enterprise feedback. Honest gap; pstack doesn't help here.
2. Fully-specify issues (ready-for-dev)Partial — a control-app + feature map give ready-for-dev a concrete “can this be driven and proven” gate, not just static checks.
3. Auto-review reliabilityDirect support — the feature map is reviewer context, and reproducible evidence lets a human or second agent re-check a verdict instead of trusting it.
4. Auto-testing / live before-after evidenceDirect hit — a control-app driving the real app and capturing state is the before/after-evidence structure Graham said is missing.
★ It also fits Engel's own “dark factory” framing from that thread: the machines work in the dark rooms; humans patrol the hallways with flashlights to check the walls hold and the room hums right. A maintained verification skill is that flashlight — it proves the user-facing behavior still holds without the human re-reading every diff. His “self-healing automations” idea then composes on top of it.

7Open questions

Decisions still Engel's to make:

  • First target. Generate the first verify-<app> against which surface — Canvas, the CLI, or the agent-server?
  • Formalize the coordinator loop. The wake-cabin / launch-sibling / verify / review sequence we run by hand could be written down as one repeatable loop, with /swarm-style fan-out as a later, small add.
  • Maintenance cadence. How often /maintain-verification-skill runs, and whether it hangs off the heartbeat.
ⓘ Companion material: the source discussion is Graham's Software Factory thread in Slack #proj-automation; the general read-for-parts analysis this note replaces is preserved in git history.