Engel's Code Design Notebook

Code-grounded notes on how OpenHands, its SDK, and SmolPaws actually work — mechanism traces, architecture studies, and design decisions.

Verification

Title frame of the explainer video: Verifying the agent SDK, OpenHands / software-agent-sdk
OpenHands agent-sdk verification
Feature maps written for agents, a CLI that drives the real server, and known bugs kept as expected failures. Start with a two-minute explainer video, then the Agent Server map and its bug tracker, the core SDK map in progress, and the OpenHands app map.
video + table of contents · 11 October 2026

Investigations & dashboards

Last week’s issues: closed by merged PRs
75 issues posted by enyst and smolpaws on October 4–11 are closed by merged PRs: 61 by those accounts and 14 by others. Every issue links its closing PR and author; one partial tracker is labeled.
created during the week · 11 October 2026
Last week’s issues: still open
31 issues posted by enyst and smolpaws on October 4–11 remain open. 24 have an open PR and 7 do not. Each issue shows its open PR count, titles and authors, including drafts.
created during the week · 11 October 2026
OpenHands PRs: Keep / Close without merge
3 Keep candidates and 39 Close candidates in separate sections. Each PR has a link, title, author and reason. The nine merged positives, including #18127 merged on 11 October, and two already closed submissions are listed separately.
review queue · 11 October 2026
Earlier OpenHands issue / PR survey
Snapshot at 04:24 UTC on October 11: open issues of any age without an identified implementation PR by enyst or smolpaws, with other-author proposals. Its weekly tally uses closure dates. Use the separate pages above for issues posted during the week.
earlier survey · 11 October 2026
OpenHands PRs by enyst and smolpaws
26 open PRs in five OpenHands repositories. See what each changes, its latest review, checks, and next step.
work list · 9 October 2026
Can Cloudflare's new CLI replace Wrangler?
What the cf beta changes for Liberty Labs: prebuilt releases, local D1, credentials, and privacy checks. A source-linked assessment and a bounded trial before changing the deployment workflow.
deployment study · 5 October 2026
Jev for OpenHands: decisions worth testing
57 live requests: issue triage, TOCTOU and supply-chain screening, personal relevance, and evidence sufficiency. Three reproducible typed examples.
experiment · 2026-09-21
Agent Canvas: a live UI audit
137 desktop and mobile screenshots from OpenHands PR #17435's Tailwind theme migration and shadcn-lint trial. Real conversations, workspace edits, settings, and automation screens — with page-by-page coverage, observed defects, and explicit testing gaps.
visual QA · 19 Sep 2026
Reviewing the reviewer
Two verified approval contradictions, independent code checks, and a proposed OpenHands automation that scores reviews without confusing bad paperwork with bad code.
investigation · proposed automation
What OpenHands users actually want
All 648 open issues across OpenHands, software-agent-sdk, and automation — clustered by demand and weighted by signal (comments × reactions), then split into two lenses: the whole community, and the wide community (non-maintainers). Routed to Engel's attention or relayed to other maintainers by passion-fit.
demand survey
OpenHands Review Grove
Sortable reviewer activity dashboard — who reviews what across OpenHands repositories.
data
The Mention Gazette
A daily newspaper of where I was pinged on GitHub — review requests and mentions first, the rest below. Regenerated each morning.
daily

How things work

Before and after the secret-selection switch: fixed green becomes theme-aware
Agentic linting refactoring
Shared theme tokens, component ownership, and shadcn rules an agent can act on. A Before/After video, the proposals merged so far, and the next steps, with a direct MP4 download.
video + rollout · 11 October 2026
How SmolPaws resets its context
The agent saves notes and chooses when to call condense. Two sequence diagrams trace that reset and the separate emergency summarizer, with warning thresholds, configuration and the context each path leaves behind.
implementation guide · 5 October 2026
The path to a software factory
Graham Neubig's plan for turning OpenHands OSS work into an autonomous software factory: profile-scoped Docker workers, four independent automations (triage → develop → review → watchdog-merge), and complete setup and operation through Agent Canvas. A mirror of his document — required outcome, existing machinery, the PRs each change needs, review/release order, and validation against the Box and Airbnb clones.
plan
What we take from pstack
A decision record, not a survey: what we chose to adopt from Lauren Tan's (@poteto) pstack — verification-as-infrastructure and her two verification-skill generators (now in .agents/skills/), a feature map + control-app CLI for the product surfaces, and the SDK's projection oracle left as its own lever. Includes how to use it and how it maps onto Graham Neubig's four software-factory bottlenecks.
decision
control-openhands: agents using Agent Canvas like a user
The CLI behind the verify-openhands skill (OpenHands PR #17961): isolated stacks, a browser that persists across commands, verbs for what screenshots miss, and an evidence ledger. Its commands, why an agent is better off with it than running the app directly, and how 36 parallel fix agents and their reviewers used it: 3,454 calls, 170 launches and 261 before/after screenshots.
how it works · 5 October 2026
ask_oracle: a second opinion without switching your model
PR #3673 adds a built-in tool that consults a saved oracle LLM profile for a one-shot answer — stateless, no history, no tools, and the agent's own active model is never changed. Why it's not switch_llm and not a subagent, with a code-grounded mechanism trace.
show-me
docs #688: redrawing the product & repo boundaries
The OpenHands docs PR that untangles the old "Agent Canvas = UI+backend blob" into four honest nouns after the repo migration: Canvas is the client, Agent Server (in the SDK repo) is the runtime, Automation Server schedules, and Sandbox Server is community-only. What moves where, page by page, grounded to the diff — plus the feline review.
show-me
Two files, one truth: meta.json vs base_state.json
The agent-server persists the whole agent config twice. The duplicate source of truth reverted a restored conversation's LLM config and doubled the trajectory secret leak. A source-grounded trace of where the copy is made, and the case for one owner.
study
Three ways to end the meta.json / base_state.json split
The decision note: (a) exclude the agent from the stored record, (b) a minimal "base_state wins" patch, or (c) stop the stored record from being a create-request at all. Compared field by field, with the resume/reattach edge resolved — and why (c) wins.
decision
GPT-5.6: Programmatic Tool Calling
The model writes sandboxed JS that orchestrates tool calls. What it needs, what the OpenHands SDK sends and parses today, and the two-sided gap that silently stops the loop. Combined show-me + teach-me, with a quiz.
study
GPT-5.6: Tool search (defer_loading)
Load tool schemas on demand to cut prompt tokens while preserving the cache. Why it's a clean greenfield add for the SDK's fat toolset. Combined show-me + teach-me, with a quiz.
study
GPT-5.6: WebSocket mode
Persistent socket + incremental inputs for ~40% faster tool-heavy rollouts — and the cheap previous_response_id Phase 0 the SDK can adopt today. Combined show-me + teach-me, with a quiz.
study
GPT-5.6: Multi-agent vs our subagents
The model spawns subagents server-side — but the SDK already ships a portable client-side subagent system. A design decision, not just a gap: who orchestrates. Combined show-me + teach-me, with a quiz.
study

SmolPaws channels

SmolPaws WhatsApp Entry Point
The former root process — Baileys multi-device, a SQLite-backed poll loop, scopes and permissions, media, and voice notes. Kept as architecture history; the standalone relay bridge now serves the primary channel on the shared server.
historical · legacy path
SmolPaws Slack Entry Point
Socket Mode bridge paws — a standalone process on the durable Message Relay and the TypeScript agent-server, with thread tracking, access control, and guest rate limiting. The first channel on the new stack.
relay · historical canary
SmolPaws Discord Entry Point
The legacy Discord adapter on the /turns runner, kept as history. The standalone relay rewrite has merged; a live check remains.
historical
SmolPaws GitHub Entry Point
A Cloudflare Worker GitHub App + @smolpaws mention poller — each issue/PR maps to one durable conversation, so pings in a thread resume the same context.
live

The next SmolPaws

Three tracks: run the primary channel on the new transpiled stack, put a cat inside OpenHands, and make both cats talk in realtime voice.

Agent-triggered condensation: a note to your future self
The original design for agent-owned notes, staged warnings and a fresh context. Kept as design history; the implementation guide above explains the shipped mechanism.
design history · September 19
LLM Profiles: configuration, runtime and maintenance
The current TypeScript contract: shared role/scope configuration, saved profiles and private credentials, native provider APIs, safe switching, usage accounting, tests and the separate GitHub LLM environments. Source-verified September 16.
implementation notebook
WhatsApp: normal operation on the shared server
Main, OpenHands and Hunting moved to the normal shared server on September 17, preserving their conversations. Current storage paths and cutover verification, with the earlier iPad canary evidence kept separately.
production · shared :8790
Bridges, one server, and the last steps
One TypeScript OpenHands agent-server and standalone relay bridges. WhatsApp now uses the normal shared host; other ingress migrations remain separate. Process boundaries, dated deployment evidence and operational ownership.
current map
WhatsApp as a bridge
The design that moved the primary channel out of the repo root into a standalone apps/whatsapp on the durable Message Relay — implemented on the shared relay runtime and promoted to normal operation on September 17. bead smolpaws-kxa · P0
implemented
The insider cat
The current Cat interface, one saved OpenHands agent across text and Voice, and a direct Codex handoff. Normal OpenHands tools, enabled skills, and accurate backend context let the Cat work through the existing API.
worktree prototype · not released
The insider cat — from your side
Choose or create a Cat, talk and type into the same conversation, keep controls while navigating, and inspect saved results. Current behavior and coordination through the normal OpenHands terminal and API skills.
behavior · Apps and next steps
The Secretary
Historical design for a conversation-management board and voice companion. Retained for the design lineage; the current Insider Cat review separates implemented skin behavior from proposed dispatch, tickets, and monitoring.
historical proposal
ADR: Durable message work belongs around agent-server
The architecture decision the whole cutover rests on — keep the agent-server upstream-shaped and put durable channel intake and delivery in a small local coordinator. The Message Relay now shared by the standalone Slack, WhatsApp and Discord implementations.
ADR · decision
Maintaining the OpenHands TypeScript transpiles
The living maintenance map behind the cutover: one canonical pin, bounded drift intervals, tests-first porting, and generated Python OpenAPI oracles keeping the new stack faithful to upstream.
OpenHands · transpilation
SDK swap surface
Historical symbol inventory and August 30 decisions. Current contracts and provider guidance are linked from the maintenance page; current operational gates are in WhatsApp readiness.
migration

Architecture Studies

Checking the Agent SDK from outside
What the external verifier checks, what its first Cloud run found, and what a regular contract check still needs. The Cloud automation is disabled.
Agent SDK · 9 October 2026
Testing architecture when agents write the code
A testing strategy for the OpenHands Agent SDK when humans no longer read PRs: separately governed architectural contracts, independent REST/WS verification, stateful and mutation testing, and a recurring self-healing loop. The complete response, reproduced verbatim.
testing strategy · 8 October 2026
They really are in the same place: RBAC (#17055), keys, and the sandbox
The RFC wants Read/Run/Admin roles for a shared Canvas — genuinely useful against remote callers. But agent-server = agent-sdk = the docker sandbox = the remote (cabin) = the local Mac: one nested runtime. So the root key that mints scoped tokens must arrive inside the sandbox where the agent's own bash runs. What roles fix, what stays (key sprawl across browsers/iPad), the config file the server reads, and the boundary that only an out-of-sandbox control plane could give. Source-grounded to the live agent-sdk.
security · RFC #17055
Canvas Extensions: the architecture and the hard truth
How an Agent Server installs UI bundles and Canvas loads them as same-realm JavaScript: manifest, adapter, registry, lifecycle, and staged-update patterns—plus the security boundary that does not exist, the contracts already drifting, and the work required before this is honestly a platform.
architecture teardown
Agent Canvas whole-codebase audit: ten seams worth fixing
A read-only review across thirteen subsystems found ten actionable ownership and state-model seams: backend-scoped settings and MCP health, order-independent events, durable form edits, source-driven workspace views, honest Home launch state, canonical wire types, and instance-safe embedding.
architecture review · 10 issues
SmolPaws architecture audit: five seams to harden
A whole-repository, evidence-first review: fourteen explicit scopes and twenty provisional opportunities, independently reduced to five lean correctness and security findings. The carrying pattern is identity drift — one real operation gaining multiple keys or owners.
architecture review
Maintaining the OpenHands TypeScript transpiles
The living maintenance map: one canonical pin, bounded drift intervals, tests-first porting, generated Python OpenAPI oracles, SDK wire goldens, and clear policy for deviations and exclusions.
OpenHands · transpilation
OpenHands SDK, June→August 2026 — what changed
An area-by-area synthesis of v1.29.2 → v1.42.1: the Python prompt registry, the conversation tree, LLM provider connections, agent-server telemetry/plugins/gateway, and the public surface that was removed.
OpenHands · change study
ADR: Durable message work belongs around agent-server
The architecture decision record for SmolPaws message reliability: keep the OpenHands agent-server upstream-shaped and put durable channel intake and delivery in a small local coordinator, borrowing Automation's table-as-queue mechanics and Hermes's per-session lane policy. Status: accepted and implemented as the Message Relay.
ADR · decision
OpenHands Gateway Compat
Live test results: which OpenAI-compatible clients work with the agent-server's /v1/chat/completions endpoint. Curl, Python SDK, LibreChat, Open WebUI, Pipecat.
results
OpenHands OpenAI-Compatible Gateway
End-goal architecture for exposing the OpenHands agent-server as an OpenAI-compatible model endpoint.
architecture
OpenHands Deferred Init Runtime Boundary
Current warm-pool deferred-init flow, PR 400-style app-state DI, and a safer runtime-context pattern.
architecture
OpenHands Client-Defined Tools
How JSON tool specs become SDK tools, WebSocket action events, acknowledgement observations, and persisted client-side execution hooks.
architecture
OpenHands credential boundaries
Issue #15298, verified: reference-only persistence, typed brokers, the real stdio MCP boundary, runtime-visible fallbacks, current-code seams, and a migration plan.
architecture
OpenHands Automations: where work runs
The September 15 operating snapshot: three local jobs, six enabled Cloud jobs, two paused jobs, what each can change, recent results, and decisions still to make.
operations · dated snapshot
OpenHands Automations
How automation definitions, cron and webhook triggers, tarball storage, dispatch, sandboxes, callbacks, and watchdog reconciliation fit together.
architecture
Do We Need typescript-client?
Responsibility map: SmolPaws core does not import it; Agent Canvas depends on its browser-safe agent-server clients. What to keep, what overlaps the transpiled SDK, and where the Work Coordinator must not couple.
report
Hermes Agent Gateway
How NousResearch's Hermes structures its multi-platform gateway, 28 platform adapters, and the OpenAI-compatible endpoint that turns an agent into an LLM.
study
Hermes Plugin API
Deep dive into the Hermes plugin system: how platforms, tools, and providers are registered with zero core code changes. Lessons for OpenHands agent-sdk extensibility.
study
Hermes Telegram Platform
Bot API 10.1 rich messages, streaming drafts, capability latching, approval flows, and notification volume control. The patterns worth adopting.
study
Voice for AI Agents
June 2026 voice landscape, retained as historical background. The Insider Cat review has the September OpenAI Realtime/GPT-Live decision and an ElevenLabs comparison.
historical landscape
Executor E2E Recording
How executor's e2e harness records a test: Playwright video, asciicast terminal capture, a focus timeline derived from operations, watchable slow-mo, and an ffmpeg splice into one film.
study
Agent Canvas Tool Visualizers
How per-tool React visualizers plug into the conversation event renderer, plus a follow-up addon path for custom renderers.
architecture
ToolShield LLM Security Analyzer
PR #2911 map: separate guardrail LLM, ToolShield safety experiences, data flow, risk labels, bypass hardening, and review edges.
security
Stuck Detector: nudge before hard-STUCK
PR #4332 before/after: a model repeating the same failing tool call used to hit terminal STUCK at 3 tries — now it gets one corrective nudge and a self-correction window first, and only goes STUCK if it repeats once more.
show-me
smolpaws: The Turn-Monitoring Loop, Twice
Architecture review of the legacy /turns client: monitorTurn and monitorConversationTurn ran the same poll loop twice. Moot once the legacy runner retires; kept as history.
review
teach-me: triggerFree Groups
Interactive walkthrough of the smolpaws change that lets a group talk to the cat without an @mention — background, intuition, code, and a quiz to check you understood.
teach
Agent Canvas: Local vs Cloud Backends
How a conversation call forks on getActiveBackend().kind — ConversationClient direct to the local agent-server vs callCloudProxy (direct app-host or proxied sandbox), and why pause differs from interrupt.
architecture
SDK swap surface: hooking smolpaws to the new package
Historical wire-by-wire migration map from agent-sdk 0.10 to the fresh transpile, kept for the Aug 30 decisions in its last section. Full ingress cutover remains tracked in the current readiness plan.
historical · migration
Secrets: keyring refs vs. encrypted-at-rest
Design comparison — the TS transpilation stores a keyring reference, the Python SDK stores an encrypted value on disk. The core difference, per-OS support, the Docker question, and whether the Fernet cipher can be dropped (feature request #3988).
study
Teach: Agent Canvas Cloud UI Features
Interactive study page: what in the current Canvas UI is cloud-specific, why it exists, where the code lives, and a quiz to check the model stuck.
teach
Agent Canvas + app_server Bridge
Minimal OpenHands app_server slice for Agent Canvas: sandbox orchestration, sandbox-hosted agent-servers, WebSocket streaming, and gateway gaps in the current PRs.
study
Teach: Agent Canvas Cloud Proxy Refactor
Interactive lesson for PR #1561: background, intuition, code walkthrough, and a quiz on app API vs runtime proxy routing.
teach
Agent Canvas Cloud Proxy Refactor
Before/after map of PR #1561: direct cloud app APIs are split from the remaining legacy runtime sandbox proxy calls.
show-me
Minimal OpenHands app_server
Architecture of the extracted FastAPI bridge: session-key auth, app-conversation metadata, sandbox runtime proxying, WebSocket tunnels, and temporary settings compatibility.
architecture
Before Agent Profiles: LLM Profile Behavior
Source-grounded history across local, Docker, remote, and Cloud backends: who owned profile state, why the last selection powered the next conversation, and where teams broke.
historical study
One LLM Picker, Three Layers of State
Current composer UX, four persistence options, and one invariant: an LLM selection applies to the conversation currently shown and is remembered for future conversations.
decision study
Agent Canvas LLM Profiles
Why the split between raw Cloud LLM settings and agent-server profiles is tech debt, and how profiles become the shared user-facing contract.
design note

Elsewhere