OpenHands · TypeScript · transpile maintenance
Keeping two transpiles alive
Policy audit · September 14, 2026 · LLM accounting update September 15
The SDK and agent-server are not one-time ports. They are a continuing relationship with the Python
OpenHands/software-agent-sdk: one pinned upstream commit, two TypeScript targets, and a
repeatable way to discover, classify, port, and prove every bounded change.
The governing rule: hand-write policy, generate discovery, execute evidence, and freeze update history.
1 · The three rings
SDK transpile
smolpaws/openhands-agent covers the upstream SDK, tools, and workspace packages. Observable Python behavior is the default contract; named deviations are exceptional.
Agent-server transpile
packages/openhands-agent-server preserves the REST/WebSocket boundary. Additive behavior requires an explicit EXT-SERVER-* policy.
Coordinator and bridges
SmolPaws owns durable intake, lane ordering, claims, retries, delivery ambiguity, and channel behavior. This layer borrows useful ideas from Automation but is not a transpile.
2 · One pin, bounded intervals
A machine-readable manifest in the SDK package names the exact upstream repository, full commit SHA, source/test/example scopes, and policy hints. The server consumes the vendored copy and verifies its generated Python oracles against the same SHA.
Work is never “sync to HEAD”. A candidate may be discovered from current upstream, but the review unit
is immutable: OLD_PIN..NEW_PIN. The pin moves only when every relevant change has a disposition
and its required evidence is green.
3 · The update conveyor
PORT, NO_TARGET_CHANGE, DEVIATION, EXCLUDED, DEFERRED, or DELEGATED.4 · Why the dispositions differ
| Disposition | Meaning | Maintenance effect |
|---|---|---|
PORT | Target behavior or tests must change. | Tests first, then implementation. |
NO_TARGET_CHANGE | The change was reviewed and needs no TypeScript edit. | A concrete reason is required. |
DEVIATION | The area remains relevant, but target policy intentionally differs. | Upstream changes still require review against the alternative behavior. |
EXCLUDED | The upstream subsystem is outside declared scope. | Changes wholly inside it create no port work unless scope changes. |
DELEGATED | The review unit belongs to a target this repository does not own. | Required for SDK-discovered server units. Their substantive disposition and evidence belong in the server package; delegation is not completed server review. |
DEFERRED | The change is in scope but intentionally postponed. | Tracking, compatibility consequence, and revisit trigger are mandatory. |
5 · Generated oracles
A green target test suite proves implemented behavior, but not that every upstream change was noticed. The maintenance system therefore combines completeness checks with behavioral evidence.
- Drift inventory: deterministic git facts for one bounded interval.
- Python OpenAPI operation oracle: paths, methods, parameters, media types, and status codes generated from the exact pin.
- Semantic OpenAPI schema oracle: dereferenced request and response schemas with presentation-only prose removed.
- SDK wire cases: the same language-neutral fixtures serialized by pinned Python and TypeScript.
- Runtime scenarios: deterministic server and cross-repository behavior, added incrementally where OpenAPI cannot prove semantics.
Live LLM calls remain separate. They prove that a provider currently accepts the request, not that TypeScript matches Python.
6 · Weekly rhythm
The SDK repository owns the drift-discovery watcher. Its generated interval summary is input to a selected review, not proof that a pin can advance. The SDK update runbook advances a bounded interval; the server re-vendor runbook consumes it and reviews delegated server units with the matching Python oracle. These are separate responsibilities. Active external schedules were not verified in this policy audit. A separate September 15 operations snapshot records the registered Cloud jobs, their latest outcomes, and the unresolved fork/base policy.
7 · Provider compatibility is part of the TypeScript SDK
Python delegates part of provider behavior to LiteLLM. Our TypeScript SDK speaks the provider protocols directly, so it owns the equivalent request/response normalization, capability decisions and continuation handling. A provider bug can need a TypeScript fix even when the Python SDK diff and upstream pin have not changed. Equivalent behavior is ordinary compatibility work, not automatically a deviation or extension.
Keep normalization inside the owning provider client, with small pure helpers for reused capability decisions. Preserve the shared LLMClient boundary; keep provider switches out of agents, tools, servers and bridges. Accept harmless extra fields only at the appropriate wire boundary, validate required fields, and preserve metadata needed to continue reasoning and tool calls.
A sanitized provider fixture should reproduce each quirk before the fix. Example: #28 tolerates extra tool-call fields such as DeepSeek’s index, while still rejecting missing or invalid required fields. A live provider smoke complements deterministic regression tests; it does not replace parity evidence.
Policy and placement: transpilation contract · LLM provider implementation guide.
Usage is recorded once, then accumulated
Verified September 15: SDK #35 · a983e4c and server #184 · 28106cd merged. The canonical SDK was re-vendored and both the shared product host and WhatsApp canary were verified on the new server revision, with existing conversations preserved. Deterministic SDK/server regressions and the GitHub LLM-environment DeepSeek test passed. The deployment observation records the separate live accounting check; Bead smolpaws-via tracks this work.
The provider client normalizes received usage and preserves the original usage details. In Agent.step, the SDK records one delta per LLM completion before dispatching messages or tools; several tools from one response still count once. Each record retains its provider/model identity, response ID, independent accounting ID, timing and cost provenance. Conversation stats group records by usage ID, normally profile:{profileId}, and project per-call histories and accumulated totals. Reloading reads those records without adding the spend again. The server exposes the same accounting through conversation stats; bridges do not calculate another set of totals. Standalone and auxiliary LLM calls need explicit host attribution.
Missing is different from zero. A missing counter or unmeasured earlier history makes the affected complete total unknown; known subtotals and coverage remain available. Cache reads are already included in many providers’ input totals, and reasoning tokens can already be included in output. DeepSeek cache misses are uncached input, not cache writes. Normalize each provider’s actual buckets once, without adding aliases or subsets twice.
A reported cost, including zero, takes precedence and keeps its unit: OpenRouter credits are not silently relabeled USD. Direct DeepSeek Flash can use a calculated estimate carrying the dated pricing source and applied rates; an estimate is not an invoice. Unknown cost stays unknown. Old canary responses have no recoverable accounting records, so new measurements cannot establish their earlier token usage or spend.
This ports the pinned Python SDK’s per-call and accumulated accounting concepts. DEV-SDK-007 documents the native EventLog deltas, explicit missingness and currency/provenance, and the remaining Python metric API boundaries. Upstream DeepSeek cache-hit issue #4491 matters to our direct client too; fixes involving private LiteLLM fields need separate native evidence. See the field and pricing guide and pinned-source port evidence and limits.
Subscription authentication crosses the SDK/server boundary
Endpoint detection alone does not implement ChatGPT subscription support. The SDK must own OAuth credentials, expiry and refresh, verified account headers, and the subscription-specific Responses request and streaming behavior. Its private credential store is ~/.openhands/auth/, shared with the Python SDK. The separate Codex CLI credential file is not an input to this stack.
The server owns the HTTP login flow under /api/llm/subscription/openai/: start, poll, status, models and logout. It holds pending device-login state and calls the SDK auth service. A saved profile records the authentication mode and vendor; runtime credentials are restored through the SDK when validating or starting that profile. Tokens never belong in profile snapshots, conversation events or architecture pages.
Review both layers against the same pinned Python source. Prove credential-store and transport behavior in SDK tests, then device-flow and conversation behavior in server tests after reproducible re-vendoring. A successful live subscription call proves the external path works; it does not replace those tests. The original missed scope is tracked by smolpaws-zlo.1 and smolpaws-zlo.2.
Merged September 15: SDK #31 · OAuth and subscription transport, with pinned-source port evidence, followed by server #173 · re-vendor and LLM routes. SDK fixtures cover Python transformations and streaming edge cases; server regressions cover device state, validation, missing login and restart restoration. The real subscription validation/tool/continuation smoke also passed against the packed SDK.
LLM profile implementation reference: the profile notebook documents the current TypeScript configuration/store boundary, native provider APIs, switching, usage, tests and GitHub environments, with source links and explicit limits.
8 · Documentation precedence
Code and tests describe current factual behavior. Transpilation contracts describe intended policy. Architecture pages explain the current design. Release notes and historical migration studies preserve their moment in time. A conflict between code and contract is investigated; it is not automatically resolved by rewriting whichever side is less convenient.