OpenHands · TypeScript · transpile maintenance

Keeping two transpiles alive

Policy audit · September 14, 2026 · LLM accounting update September 15

The SDK and agent-server are not one-time ports. They are a continuing relationship with the Python OpenHands/software-agent-sdk: one pinned upstream commit, two TypeScript targets, and a repeatable way to discover, classify, port, and prove every bounded change.

The governing rule: hand-write policy, generate discovery, execute evidence, and freeze update history.

1 · The three rings

strictest parity

SDK transpile

smolpaws/openhands-agent covers the upstream SDK, tools, and workspace packages. Observable Python behavior is the default contract; named deviations are exceptional.

protocol parity

Agent-server transpile

packages/openhands-agent-server preserves the REST/WebSocket boundary. Additive behavior requires an explicit EXT-SERVER-* policy.

product-owned

Coordinator and bridges

SmolPaws owns durable intake, lane ordering, claims, retries, delivery ambiguity, and channel behavior. This layer borrows useful ideas from Automation but is not a transpile.

2 · One pin, bounded intervals

A machine-readable manifest in the SDK package names the exact upstream repository, full commit SHA, source/test/example scopes, and policy hints. The server consumes the vendored copy and verifies its generated Python oracles against the same SHA.

Work is never “sync to HEAD”. A candidate may be discovered from current upstream, but the review unit is immutable: OLD_PIN..NEW_PIN. The pin moves only when every relevant change has a disposition and its required evidence is green.

3 · The update conveyor

scan
→
Generate commits, changed paths, tests, examples, target buckets, and policy hints.
classify
→
Assign PORT, NO_TARGET_CHANGE, DEVIATION, EXCLUDED, DEFERRED, or DELEGATED.
red
→
Port or add the upstream test before changing implementation.
green
→
Implement the smallest compatible behavior and run surrounding suites.
prove
→
Run generated OpenAPI, wire, integration, build, pack, and smoke evidence.
advance
→
Freeze the reviewed interval and move both transpiles to the same upstream SHA.

4 · Why the dispositions differ

DispositionMeaningMaintenance effect
PORTTarget behavior or tests must change.Tests first, then implementation.
NO_TARGET_CHANGEThe change was reviewed and needs no TypeScript edit.A concrete reason is required.
DEVIATIONThe area remains relevant, but target policy intentionally differs.Upstream changes still require review against the alternative behavior.
EXCLUDEDThe upstream subsystem is outside declared scope.Changes wholly inside it create no port work unless scope changes.
DELEGATEDThe review unit belongs to a target this repository does not own.Required for SDK-discovered server units. Their substantive disposition and evidence belong in the server package; delegation is not completed server review.
DEFERREDThe change is in scope but intentionally postponed.Tracking, compatibility consequence, and revisit trigger are mandatory.

5 · Generated oracles

A green target test suite proves implemented behavior, but not that every upstream change was noticed. The maintenance system therefore combines completeness checks with behavioral evidence.

Live LLM calls remain separate. They prove that a provider currently accepts the request, not that TypeScript matches Python.

6 · Weekly rhythm

The SDK repository owns the drift-discovery watcher. Its generated interval summary is input to a selected review, not proof that a pin can advance. The SDK update runbook advances a bounded interval; the server re-vendor runbook consumes it and reviews delegated server units with the matching Python oracle. These are separate responsibilities. Active external schedules were not verified in this policy audit. A separate September 15 operations snapshot records the registered Cloud jobs, their latest outcomes, and the unresolved fork/base policy.

7 · Provider compatibility is part of the TypeScript SDK

Python delegates part of provider behavior to LiteLLM. Our TypeScript SDK speaks the provider protocols directly, so it owns the equivalent request/response normalization, capability decisions and continuation handling. A provider bug can need a TypeScript fix even when the Python SDK diff and upstream pin have not changed. Equivalent behavior is ordinary compatibility work, not automatically a deviation or extension.

Keep normalization inside the owning provider client, with small pure helpers for reused capability decisions. Preserve the shared LLMClient boundary; keep provider switches out of agents, tools, servers and bridges. Accept harmless extra fields only at the appropriate wire boundary, validate required fields, and preserve metadata needed to continue reasoning and tool calls.

A sanitized provider fixture should reproduce each quirk before the fix. Example: #28 tolerates extra tool-call fields such as DeepSeek’s index, while still rejecting missing or invalid required fields. A live provider smoke complements deterministic regression tests; it does not replace parity evidence.

Policy and placement: transpilation contract · LLM provider implementation guide.

Usage is recorded once, then accumulated

Verified September 15: SDK #35 · a983e4c and server #184 · 28106cd merged. The canonical SDK was re-vendored and both the shared product host and WhatsApp canary were verified on the new server revision, with existing conversations preserved. Deterministic SDK/server regressions and the GitHub LLM-environment DeepSeek test passed. The deployment observation records the separate live accounting check; Bead smolpaws-via tracks this work.

The provider client normalizes received usage and preserves the original usage details. In Agent.step, the SDK records one delta per LLM completion before dispatching messages or tools; several tools from one response still count once. Each record retains its provider/model identity, response ID, independent accounting ID, timing and cost provenance. Conversation stats group records by usage ID, normally profile:{profileId}, and project per-call histories and accumulated totals. Reloading reads those records without adding the spend again. The server exposes the same accounting through conversation stats; bridges do not calculate another set of totals. Standalone and auxiliary LLM calls need explicit host attribution.

Missing is different from zero. A missing counter or unmeasured earlier history makes the affected complete total unknown; known subtotals and coverage remain available. Cache reads are already included in many providers’ input totals, and reasoning tokens can already be included in output. DeepSeek cache misses are uncached input, not cache writes. Normalize each provider’s actual buckets once, without adding aliases or subsets twice.

A reported cost, including zero, takes precedence and keeps its unit: OpenRouter credits are not silently relabeled USD. Direct DeepSeek Flash can use a calculated estimate carrying the dated pricing source and applied rates; an estimate is not an invoice. Unknown cost stays unknown. Old canary responses have no recoverable accounting records, so new measurements cannot establish their earlier token usage or spend.

This ports the pinned Python SDK’s per-call and accumulated accounting concepts. DEV-SDK-007 documents the native EventLog deltas, explicit missingness and currency/provenance, and the remaining Python metric API boundaries. Upstream DeepSeek cache-hit issue #4491 matters to our direct client too; fixes involving private LiteLLM fields need separate native evidence. See the field and pricing guide and pinned-source port evidence and limits.

Subscription authentication crosses the SDK/server boundary

Endpoint detection alone does not implement ChatGPT subscription support. The SDK must own OAuth credentials, expiry and refresh, verified account headers, and the subscription-specific Responses request and streaming behavior. Its private credential store is ~/.openhands/auth/, shared with the Python SDK. The separate Codex CLI credential file is not an input to this stack.

The server owns the HTTP login flow under /api/llm/subscription/openai/: start, poll, status, models and logout. It holds pending device-login state and calls the SDK auth service. A saved profile records the authentication mode and vendor; runtime credentials are restored through the SDK when validating or starting that profile. Tokens never belong in profile snapshots, conversation events or architecture pages.

Review both layers against the same pinned Python source. Prove credential-store and transport behavior in SDK tests, then device-flow and conversation behavior in server tests after reproducible re-vendoring. A successful live subscription call proves the external path works; it does not replace those tests. The original missed scope is tracked by smolpaws-zlo.1 and smolpaws-zlo.2.

Merged September 15: SDK #31 · OAuth and subscription transport, with pinned-source port evidence, followed by server #173 · re-vendor and LLM routes. SDK fixtures cover Python transformations and streaming edge cases; server regressions cover device state, validation, missing login and restart restoration. The real subscription validation/tool/continuation smoke also passed against the packed SDK.

LLM profile implementation reference: the profile notebook documents the current TypeScript configuration/store boundary, native provider APIs, switching, usage, tests and GitHub environments, with source links and explicit limits.

8 · Documentation precedence

Code and tests describe current factual behavior. Transpilation contracts describe intended policy. Architecture pages explain the current design. Release notes and historical migration studies preserve their moment in time. A conflict between code and contract is investigated; it is not automatically resolved by rewriting whichever side is less convenient.

Sources of truth