SmolPaws · TypeScript OpenHands · Living implementation notebook
LLM Profiles:
a name, a record,
a running model.
A profile owns the model and provider settings. Product configuration chooses a profile by name. The server saves the conversation’s selected snapshot; the SDK turns it into provider requests. Those are separate responsibilities, and changing one does not necessarily change the others.
01One path, four owners
- 1 · PRODUCTChoose a nameRole defaults and trusted scope overrides.models.json
productProfileSelection - 2 · SERVERResolve the recordLook up and validate the saved LLM profile.ServerStateService
profileAgentFactory - 3 · CONVERSATIONSave the selectionKeep an effective snapshot and any pending switch.llm_profile_ref
llm_profile_snapshot - 4 · SDKSpeak the APIResolve private authentication and serialize native requests.createClientFromProfile
provider clients
| Concept | What it means | What owns it |
|---|---|---|
| LLM Profile | A reusable profileId, provider, model, endpoint, generation options and authentication selection. | SDK schema; server catalog in this product. |
| Role selection | A profile name for a job such as agent, condenser or oracle. A scope may override the shared name. | SmolPaws product configuration. |
| Conversation selection | The profile reference and validated snapshot used by this conversation, plus a durable pending choice. | Agent-server, activated at an SDK step boundary. |
| Agent Profile | A separate preset concept for agent settings, tools and context. CRUD support alone does not mean launch applies the preset. | Separate SDK/server schemas and routes; see current limits. |
| Credential | Provider key, per-profile override or subscription account. It is not configuration to copy into a profile record. | SDK private credential abstractions and host-selected storage. |
Bridges submit work for registered lanes. They do not each resolve models, handle provider quirks or implement a separate switching mechanism. A scope is supplied by trusted lane registration; incoming message tags cannot claim another channel’s configuration.
Source: product composition · selection resolver · profile schema · client factory.
02Configure roles with saved profile names
The product reads ~/.smolpaws/models.json. SMOLPAWS_HOME_DIR changes the default directory; SMOLPAWS_MODELS_CONFIG selects a different file. The names below are illustrative: each must already exist in the catalog of the server that uses it.
{
"version": 1,
"roles": {
"agent": "general-agent",
"condenser": "small-summarizer",
"oracle": "consultant"
},
"scopes": {
"whatsapp:main": { "agent": "main-agent" },
"whatsapp:openhands": { "agent": "general-agent" },
"whatsapp:hunting": { "agent": "general-agent" }
}
}
Selection order: exact platform:scopeId override → shared roles entry. With no applicable mapping, creation uses its ordinary explicit/server-default settings and an existing conversation keeps its choice. A specialist role never implicitly borrows agent. Scheduled work retains its originating scope, including isolated runs; an explicitly configured isolated helper's profile takes precedence over these role selections.
Only the agent role currently has a consumer. The configuration accepts condenser, oracle and future role names, but saving those names does not implement or enable those features.
Save the file atomically: write complete JSON to a temporary file, then rename it over the destination. The server rereads the applicable selection when conversation work is requested. It does not watch the file or interrupt an ongoing call merely because the file changed. An absent default file preserves ordinary behavior; an explicitly configured missing or malformed file fails instead of silently choosing a different model.
Configuration contract and examples · parser and selection tests.
Scheduled helpers: a small context and a concrete handoff
September 18, 2026 · architecture update. SmolPaws #210 adds an isolated Slack checker within the shared scheduler. This section describes the public implementation contract, not a particular installation's runtime state.
An isolated occurrence starts a fresh conversation, but ordinarily inherits the owner's context and tools. The Slack checker instead gets a short role brief, a saved deepseek-v4-flash profile and five selected tools. The same shared runtime still owns its schedule, durable run records and delivery; this does not introduce another timer service.
scheduled-agents.json selects profile, context_files and tools by scheduler task ID. SMOLPAWS_SCHEDULED_AGENTS_CONFIG can select another configuration file. Only a recorded isolated occurrence can use its task's entry; request tags cannot claim it. The file list replaces the normal context files for new occurrences, while existing snapshots and ordinary channel conversations remain intact. The profile name resolves through the existing catalog and secret store, independently of the owner's model.
| Tool | Purpose |
|---|---|
check_slack | Read mentions and followed-thread replies through the existing Chrome session; return findings, stable source IDs, quiet success or an error. |
recover_slack | Try other Slack tabs, then permitted Chrome window recovery, only after a failed normal check. |
notify_smolpaws | Queue a labelled automatic user message for the full conversation owning the task. |
terminal | Run commands for unexpected diagnosis. Tool selection is a capability choice, not a shell sandbox. |
finish | Finish quietly with an empty message, or report a specific unresolved failure. |
Handoff keeps the helper small. The checker stays on DeepSeek; it does not switch profiles. notify_smolpaws resolves the owner's current relay conversation, so a conversation rotation does not stale the destination. The complete pending Slack batch enters the existing intake queue as untrusted external input. The receiving SmolPaws keeps its own full context and profile, and its normal response reaches the WhatsApp OpenHands group. Acceptance means durable queue ownership, not a completed response. Retried source IDs cannot duplicate the same handoff; the checker advances its watermark only after acceptance.
Normal checks leave browser windows alone. Google Chrome is the agents' browser; Comet belongs to the human. Slack can remain in a background Chrome tab, and multiple windows alone are not a failure. Recovery first tries other matching tabs, then may retain a known oldest window or close all Chrome windows and reopen one, recreating the configured Slack tab when needed. It preserves the browser profile and login. API failures and incomplete pagination remain errors, never quiet success; seven-day followed-thread retention remains an explicit limit. The checker does not post Slack messages or add reactions.
Verification contract: regressions cover configuration isolation, context snapshots, profile selection, exact tool exposure, quiet completion, durable handoff and retries, pagination, state restoration and Chrome-only recovery. The opt-in DeepSeek script exercises the real tool schemas against fixture browser data, a fixture recipient and a fixture outbound transport.
These policies live in SmolPaws's product host and coordinator. They reuse the existing server composition hooks; the SDK and server API contracts are unchanged.
Configuration, state and activation contract · Checker role and house rules · Opt-in live fixture · Implementation and regression tests.
03The catalog, snapshots and credentials live separately
The SDK defines an LLM profile; the host persists it. In the agent-server, ServerStateService owns state.json, containing llmProfiles, agentProfiles, settings and secret metadata. It loads once per service instance, serializes its mutations and replaces JSON atomically. This is not a multi-process profile database or a file-watching configuration service.
| Data | Location / override | Update through |
|---|---|---|
| Product role/scope choices | ~/.smolpaws/models.json; SMOLPAWS_MODELS_CONFIG or typed host options can override it. | Atomic file replacement; read on requested work. |
| Server profile catalog and defaults | <state directory>/state.json. Directory: OPENHANDS_AGENT_SERVER_STATE_PATH, else <conversations path>/server_state. | Profile and settings REST mutations; editing disk does not update the already loaded service. |
| Conversation snapshot and pending selection | <conversations path>/<id>/meta.json, under the stored request. | Server creation and lease-guarded switching. |
| History and usage deltas | <conversations path>/<id>/events/event-*.json. | SDK EventLog; retained across switches. |
| Always-on product context | smolpaws-context.json beside conversation metadata; selected through the separate product context policy. | Snapshot at context preparation, retained on model switch. |
| Provider and profile keys | SDK SecretStore. The default server uses macOS Keychain; tests can inject an in-memory store. | Credential/settings API or host SecretStore. |
| OpenAI subscription account | OH_PERSISTENCE_DIR/auth/openai_oauth.json, default ~/.openhands/auth/openai_oauth.json; private directory/file modes 0700/0600. | SDK OAuth login, refresh and logout, exposed through server subscription routes. |
The generic conversation-root precedence is OPENHANDS_CONVERSATIONS_PATH → PERSISTENCE_DIR → relative workspace/conversations. The product bootstrap normally supplies PERSISTENCE_DIR=$SMOLPAWS_HOME_DIR/conversations. Canary overrides are deployment choices, not SDK defaults. A models.json shared by two hosts still requires the selected names to exist in each applicable host’s own catalog.
Only a missing state file seeds the initial default profile: OpenAI gpt-5-nano, Responses mode, and server conversation limit 500. Corrupt state fails rather than silently resetting. A profile record alone does not supply a credential.
What goes in an LLM profile
{
"profileId": "main-agent",
"providerId": "openai",
"model": "gpt-5-mini",
"openAiApiMode": "responses",
"reasoningEffort": "low",
"cachingPrompt": true,
"useProfileKeyOverride": false
}
This is a schema-valid illustrative record. The strict SDK schema supplies defaults where declared and rejects unknown properties. Optional provider settings such as anthropicCacheTtl stay absent when omitted. The schema has no raw apiKey field and no fallback-model chain.
Profile field reference and defaults
| Fields | Shape / default |
|---|---|
profileId, providerId, model | Required. Profile ID: 1–64 characters, starts alphanumeric, then letters/digits/dot/underscore/hyphen. Provider ID uses letters/digits/dot/underscore/hyphen. Model is a nonempty string. |
authType, subscriptionVendor | api_key by default, or subscription; vendor openai or null. |
baseUrl, openAiApiMode | URL or null; mode chat_completions by default, or responses. Native clients select their default endpoints when applicable. |
temperature, topP, topK | Null by default. Respectively a nonnegative number, number 0–1, or positive integer. Provider support differs. |
maxInputTokens, maxOutputTokens, timeoutSeconds | Null by default; positive integer token counts / positive numeric seconds. Schema acceptance is not enforcement: see current limits. |
reasoningEffort, reasoningSummary | Null by default. Effort accepts only low, medium, high; summary auto, concise, detailed. |
cachingPrompt | True by default; controls supported explicit cache preparation. |
anthropicCacheTtl | Optional 5m or 1h, for Anthropic caching only. Omitted stays unset; Anthropic's wire default is five minutes. Configuration and verified one-hour behavior. |
promptCacheRetention, promptCacheKey | Null by default. Retention accepts 24h/disabled; key is a nonempty string. Emission is capability-gated. |
headers, useProfileKeyOverride | String map defaults to {}; key override flag defaults false. Do not place credentials in arbitrary headers or URLs. |
Key lookup is explicit
A SecretRef is a service/account pair; the default service is openhands. When useProfileKeyOverride is true and llm-profile:<profileId>:api-key exists, use it. Otherwise look up llm-provider:<providerId>. Missing both fails client construction. Transport classification does not rename the provider ID for credential lookup.
PATCH /api/settings can store llm_api_key through the active profile’s secret reference. General settings secrets and conversation secrets have separate account namespaces and redacted API representations. The session API key (OPENHANDS_SESSION_API_KEY, then SESSION_API_KEY) authenticates server API requests; it is not an LLM credential.
Subscription authentication is an explicit different path. It uses SDK-owned device/browser login, private OAuth persistence and refresh; it neither reads ~/.codex/auth.json nor silently falls back to an API key. The server exposes status, supported models, device start/poll and logout under /api/llm/subscription/openai/. Even a status request may refresh an expired account; it is not guaranteed to be an offline read.
Profile and settings REST operations
| Request | Meaning |
|---|---|
GET /api/profilesGET /api/profiles/:name | List saved profiles and active ID, or read one profile. |
POST /api/profilesPOST /api/profiles/:name | Schema-validated upsert. The named route’s path overrides the body ID. This is a complete record, not a PATCH merge. |
POST /api/profiles/:name/activate | Set the server’s active/default profile for future creation. Existing conversations retain their snapshots. |
POST /api/profiles/:name/renameDELETE /api/profiles/:name | Rename via {"new_name":"…"}, or delete. References/credentials have limitations described below; these are not a universal migration of every consumer. |
POST /api/profiles/:name/validate | Makes a real model completion. Validates body {"llm": <complete profile>} with a ping. It does not look up the path name or save the record and can consume usage. |
GET /api/settingsPATCH /api/settings | Server defaults and settings. An explicit agent_settings.llm_profile_ref can differ from the active-profile marker; read the actual reference. |
/api/llm/providers, /api/llm/models and /api/llm/models/verified expose the curated SDK catalog. They are not a live provider inventory or an allowlist; a saved profile can name a model outside that catalog. The generated OpenAPI document is the route/schema reference.
Server state and mutations · server path defaults · conversation metadata · LLM profile schema · secret-store contract · profile routes · generated OpenAPI · OAuth persistence.
04Create, switch and restore the same conversation
- Create from a saved reference.The request’s agent settings win over server defaults. The selected profile record is validated and snapshotted. REST creation still requires its initial requested/default reference to exist, even when product configuration will select another profile before execution.
- Observe a requested configuration change.A run request causes the product selector to reread its applicable name. A busy conversation applies that request after the current complete model/tool step settles, before another model call.
- Let the agent choose through
switch_llm.The tool acceptsprofile_nameand a reason. The server validates the target, prepares its client and durably queues it. The tool returns acceptance; it does not wait for the run that contains itself. - Activate at the safe boundary.After all results from the tool batch are recorded, the SDK can replace the agent’s LLM binding. Conversation identity, messages, tools, context, step budget and accumulated usage remain. A switch alongside
finishstill persists for the next turn. - Restore from durable state.The effective reference/snapshot and pending choice survive restart. The server does not reconstruct the choice by blindly applying today’s default to old history.
| Change | Effect on an existing conversation |
|---|---|
Change its applicable name in models.json | Observed on a work request; applied at a safe boundary in the same conversation. |
Call switch_llm | Applies before the next model call after the current tool batch; durable even if the run finishes. |
| Leave the configured name unchanged | Preserves a tool-selected profile across turns and restart. |
| Edit another role or scope | Does not reset this conversation’s agent. |
| Edit a catalog record without changing its name | Existing snapshots remain. Explicit switch_llm reselection can load the updated record. |
| Remove a scope override | A shared role selection becomes applicable if present. Removing the final mapping retains the conversation’s current/pending choice. |
| Activate a profile in the server catalog | Changes the server’s creation default; it does not switch existing conversations. |
Changing the applicable configured name again takes precedence over an earlier tool choice. A missing target, malformed selection or authentication failure is surfaced rather than silently falling back. An unknown profile or failed client construction returns a tool error without replacing the active profile. After a persistence or cleanup error, saved metadata is authoritative because a write may already have committed. An activation-boundary failure stops that run; the next turn rebuilds the persisted binding.
Durable profile runtime · requested work and boundary wiring · upstream switch evidence and deliberate differences.
05One thin interface; native provider implementations
The shared interface is LLMClient.complete(messages, tools?). Dispatch uses the profile’s provider and endpoint, then its API mode. It does not choose credentials by guessing from the model name. Known Anthropic and Gemini providers use their native clients; other compatible providers select Chat Completions or Responses. An unknown provider ID may be classified by its base URL, but credential lookup still uses the original provider ID.
| Transport | Implementation | Protocol responsibilities |
|---|---|---|
| OpenAI-compatible Chat Completions | openai.ts | /chat/completions; native tools and text/image parts; compatible reasoning metadata; max_completion_tokens, sampling and reasoning controls where supported. OpenRouter, LiteLLM gateways and direct DeepSeek can use this transport with the right endpoint. |
| OpenAI Responses | openai.ts | /responses; input/output items, native function calls, system instructions and store:false. Preserve response/function item IDs and compatible encrypted reasoning through replay. |
| Anthropic Messages | anthropic.ts | /v1/messages; native system blocks, tool_use/tool_result, signed and redacted thinking. Default output allowance is 4,096; native top_k is supported. Thinking configuration constrains budget and temperature. |
| Gemini Interactions | gemini.ts | /v1beta/interactions; native function calls/results and thought signatures, store:false, output/thinking controls. This implementation explicitly rejects temperature, top-p and top-k settings. |
| OpenAI subscription | auth/openai.ts | Explicit subscription profiles use the Codex Responses transport, refreshable SDK-owned OAuth and supported-model checks. Transport consumes streaming events internally; the common client still returns a completed response. |
A gateway serving Claude still uses the gateway’s wire protocol. Model capabilities can affect caching or reasoning preparation; they do not silently change the chosen transport or credential owner. The TypeScript runtime does not depend on Python LiteLLM, even when its configured endpoint is a LiteLLM proxy.
Prompt caching has provider-specific semantics
Supported Claude families receive ephemeral cache controls on static system content and the latest user/tool content. Dynamic system context stays separate. The SDK modifies outgoing copies, preserves tool-result cache controls, and rejects requests containing more than four cache breakpoints. This works through native Anthropic Messages and supported compatible gateways. cachingPrompt:false disables these explicit markers; it does not disable every provider’s implicit caching.
OpenAI retention is a separate capability, currently restricted to eligible direct-OpenAI GPT-5.6 requests. A null retention setting defaults to 24h there; disabled omits the retention parameter. Subscription and unrelated gateways do not inherit this behavior. A provider cache hit must be demonstrated through usage, not assumed from successful serialization.
Anthropic cache lifetime belongs to the profile
Deployed September 17 at 03:47 Europe/Stockholm. SDK #42 merged as 8c25a50; server #202 re-vendored it in 61e5fed. Main's existing saved profile now requests one-hour caching. The local and GitHub provider proofs below used isolated Haiku conversations.
Optional-setting correction · September 17. SDK #44, merged as 4212592, removes the schema-inserted TTL default. Server #205 re-vendors it in 6e53dc6, with omission preserved through CRUD, snapshots, restore and generated OpenAPI. Explicit one-hour request serialization is unchanged; the earlier provider proof below remains tied to the version tested.
anthropicCacheTtl is an optional Anthropic setting accepting 5m or 1h. Omitting it leaves it absent from parsed profiles, saved records and API output; unrelated providers do not acquire it. When Anthropic cache markers are enabled, omission or explicit 5m keeps {"type":"ephemeral"}, whose provider default is five minutes. Explicit 1h produces {"type":"ephemeral","ttl":"1h"}. This distinction preserves the explicit one-hour setting on the local Fable profile.
One configured value applies to every generated marker, including tool-result markers, through native Anthropic Messages and compatible Chat Completions gateways serving supported Claude models. Existing capability checks and cachingPrompt:false control whether markers appear; setting a TTL does not enable caching on subscription or unrelated providers. OpenAI's promptCacheRetention remains a separate option.
Using one TTL throughout a request avoids mixed-lifetime ordering mistakes: Anthropic requires longer-lived cache entries before shorter-lived entries. One-hour writes cost twice base input; the lifetime runs from the start of a request that writes or reads the cache. A server restart or cache expiry does not erase the saved conversation. Anthropic's lifetime and billing contract.
Existing conversations keep their snapshots. Editing the catalog record alone does not refresh a same-name selection. The 03:47 rollout used an authorized narrow maintenance migration while idle and with the server stopped: it changed only the TTL in the Fable catalog entry and five saved Anthropic profile snapshots, including Main and a Haiku test conversation. Other profile options, selections, conversation IDs and accumulated metrics stayed unchanged. This did not create a new conversation, refresh memory files or change the general catalog-edit policy.
03:47 production preservation proof: the shared :8790 server restarted on 61e5fed with all 45 conversations retained. All 2,666 event files and 22 context snapshots kept their exact bytes. A read-only API check showed Main idle with anthropicCacheTtl:"1h". The WhatsApp process kept the same PID; the paused bridge processes resumed, heartbeat was restored and relay queues were clear. Deployment made zero provider calls and sent Main no test prompt. Main's next natural request will request one-hour cache writes; deployment itself did not pre-warm its cache.
Optional-setting rollout · September 17 at 04:17:28 Europe/Stockholm: after the scheduled work finished normally, the shared server updated to 6e53dc6. Cleanup removed three automatically inserted five-minute fields from DeepSeek conversation snapshots. It retained all 47 conversations, with 2,713 event files and 24 context snapshots byte-for-byte unchanged; metrics and every other profile option were preserved. Read-only verification found all five saved Anthropic profiles still set to 1h, all 42 other saved profiles without the TTL field, and no Anthropic TTL fields in the non-Anthropic catalog records. Main was idle with 1h. The WhatsApp process kept its PID, heartbeat was loaded and relay queues were clear. This rollout made zero provider calls.
Local live proof · September 17: the isolated Haiku smoke through the eval proxy wrote 9,719 tokens to the one-hour cache, then read 9,719 cached tokens on the warm request and 9,829 after restoring the conversation. Provider duration counters confirmed that all reported writes used one hour; outgoing markers and saved usage passed the smoke's checks. The proxy exposes the duration under usage.prompt_tokens_details.cache_creation_token_details.ephemeral_1h_input_tokens; native Anthropic uses usage.cache_creation.ephemeral_1h_input_tokens. This was a short three-request check, not a test after waiting an hour.
GitHub proof · September 17: ANTHROPIC_CACHE_TTL=1h is verified in smolpaws/openhands-agent's LLM environment. Live LLM run 35171163557 passed both the one-hour Haiku cache smoke and DeepSeek check on SDK 8c25a50. The workflow passes the variable to Haiku and defaults to 1h if it is absent. This checks the GitHub provider path independently of production; the September 16 evidence below remains historical.
Profile schema · cache preparation · Haiku smoke · Live LLM workflow · saved profile selection.
Reasoning and tool history must remain valid after switching
Provider clients serialize native thinking/signatures and reasoning fields. GPT-5-family normalization omits unsupported temperature. Subscription transport removes unsupported options and consumes the streamed response shape. DeepSeek reasoning requests with tools need a reasoning_content field even on historical assistant tool messages from another model; when absent, its serializer supplies an empty string while preserving real recorded reasoning.
Completed tool results are placed adjacent to their assistant request in outgoing request views, including when new user input arrived during tool execution. This projection neither rewrites saved event order nor invents a missing tool result. Incompatible opaque provider history is handled separately by origin, as described under usage and history.
A message received while the agent is answering follows that reply in the next request
September 17, 2026 · SDK #41 · DEV-SDK-009. If B arrives while the agent is answering A, the next model request receives A → reply(A) → B. The saved log keeps the actual arrival order: A → B → reply(A). The server still runs the follow-up for B.
For serialized agent steps, the SDK records the request’s input boundary and its response event IDs. These saved markers make the ordering reproducible after restart and across profile changes. The projection uses only events retained after condensation and does not move input across summaries. Old history without these markers keeps its recorded order; the SDK does not guess which messages an earlier request had seen.
Dispatch factory · capability helpers · Anthropic cache preparation · tool ordering · provider implementation guide.
06Record the provider’s usage; preserve uncertainty
Each model response returned to an Agent.step gets one durable llm_usage record, including available usage on invalid or incomplete responses. A response containing several tool calls is still one model call. A user message requiring many model steps produces many records. Statistics are derived from those records, deduplicated by record ID and restored without charging the same call twice.
- Keep prompt, output, total, cache-read, cache-write, cache-miss and reasoning counters distinct. Cache reads are part of input; reasoning is generally part of output. Do not add either again to a provider’s total.
- DeepSeek cache misses are ordinary uncached input, not evidence of a separately charged cache write. Missing fields stay unknown.
- Keep the profile, requested model and actually served model per call. A switch does not relabel earlier calls or reset accumulation.
- Prefer a provider-reported cost with its currency. A calculated estimate must retain its source, dated rates and assumptions. The current built-in pricing calculator is deliberately limited to recognized direct DeepSeek Flash requests.
- Expose measured subtotals alongside missing-usage, missing-cost and unmeasured-history coverage. Never turn absent billing information into zero dollars.
Proxy billing and conversation accounting are different views. A proxy key’s weekly spend includes every consumer of that key. To attribute it to one conversation, match its response IDs against proxy receipts and compare token counts. Missing local Fable cost was resolved this way for an operator report; it was not fabricated from a model alias or backfilled into historical events.
Opaque continuation also has provenance. Usage records can carry history_origin; the initial binding is anchored before a switch. Signed thinking or encrypted reasoning is projected out of incompatible outgoing requests while the saved events remain intact. Visible text, images and valid tool exchanges remain available across profiles.
Metrics contract · records and projections · bounded cost estimates · provider-history compatibility.
07Tests: contract, integration and real providers
Three claims require different evidence: deterministic behavior is correct; it matches the pinned Python contract; a real provider accepted and completed the request. Keep those claims separate.
| Layer | What to cover | Start here |
|---|---|---|
| SDK profile/auth | Strict schemas, settings defaults, provider-vs-profile key precedence, OAuth expiry/refresh/logout races and private storage. | settings tests · subscription auth tests |
| Provider wire behavior | Native requests, response metadata, cache controls, reasoning and tool continuations; sanitized fixtures and pinned Python projections. | LLM fixtures · cross-profile history |
| Switch and usage | Tool errors, safe parallel-tool boundary, switch plus finish, restore, concurrent run requests, exact per-call/aggregate metrics and missing coverage. | switch tool · step boundaries · metrics |
| Server | Profile snapshots, pending changes, activation failures, restart, stored selection precedence, HTTP boundary and subscription wiring. | profile switching · subscription profiles |
| Product | Role/scope precedence, untrusted tags, missing/invalid config, native/registered lane identity and custom factory compatibility. | model config · product workflow |
# In smolpaws/openhands-agent: deterministic checks, no live provider request
npm test
npm run test:drift
npm run typecheck
npm run typecheck:drift
npm run lint
npm run build
npm run typecheck:live
# In smolpaws/smolpaws
npm run ci --prefix packages/openhands-agent-server
npm run relay-server:test
npm run relay-server:typecheck
npm run typecheck
The server CI also checks provenance/review evidence, generated OpenAPI parity, a local endpoint smoke, coverage, lint, build and packed-consumer use. The SDK projection gate compares against the pinned Python oracle. Ordinary CI typechecks live scripts; that does not execute them against a provider.
The GitHub environments are repository-scoped
Names and protection settings verified September 16, 2026. The table lists secret names, never values. An environment called LLM in one repository is independent of the same name in another.
| Repository / environment | Credentials and variables | Trigger and restrictions |
|---|---|---|
| smolpaws/openhands-agent LLM | Secrets: DEEPSEEK_API_KEY, LITELLM_PROXY_API_KEY.Variable: DEEPSEEK_MODEL. Optional ANTHROPIC_CACHE_MODEL is referenced but absent, so the workflow’s Haiku default applies. | Manual Live LLM workflow. Environment permits main only; jobs also require canonical repository/main. No reviewer or wait-timer rule. Separate workflow runs serialize; the two jobs may run together. Keys enter only their live-test steps. |
| smolpaws/smolpaws LLM | Secrets: OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY. No variables.The current server workflow consumes only OpenAI. | Manual agent-server Examples workflow on macOS; 25-minute timeout. No environment branch/reviewer/wait rule or workflow canonical/main guard. OpenAI key is job-level. Stored keys do not establish test coverage. |
| smolpaws/openhands-agent examples | OpenAI, Anthropic and Gemini API-key secret names, separate from the SDK LLM environment. | Older Examples path: manual dispatch or a PR receiving test-examples. No canonical/main-only guard or explicit environment protection. Do not assume it has the bounded Live LLM workflow’s safeguards. |
The bounded SDK live checks
| Check | Required evidence | Bounds |
|---|---|---|
DeepSeek v4 Flashnpm run live:deepseek-flash | Exact text, real tool call, overlapping user input, restored continuation, one tool side effect and exact usage accumulation. Zero cache hits are allowed; missing credentials fail. | Default deepseek-v4-flash; ≤12 requests, ≤4,096 output tokens/request, 45-second request and 3-minute test deadlines; synthetic echo/finish tools. |
Haiku prompt cachingnpm run live:anthropic-cache-smoke | Ordinary agent generates markers; positive cold cache write, warm cache read and read after restore; provider counters match the event ledger. Zero cache hits or missing credentials fail. | Default anthropic/claude-haiku-4-5-20251001 through eval proxy in the Live LLM workflow. Three completions expected; ≤6 requests, ≤192 output tokens/request, 45-second request and 3-minute test deadlines; finish only. |
These are request/token/time bounds, not a fixed-dollar guarantee. The hardened checks use synthetic prompts, in-memory secrets and sanitized accounting summaries. They do not operate on production conversations. Each GitHub job has a six-minute timeout.
# Intentionally invokes real providers using the SDK repository's LLM environment.
# Choose this separately from ordinary PR/CI checks.
gh workflow run llm.yml --repo smolpaws/openhands-agent --ref main
Dated evidence: Run 35102961257 passed both DeepSeek and Haiku live steps on September 16 at SDK cf17c3e. It predates the SDK revision documented here; it is not a claim that today’s entire head has run through that workflow. The subsequent profile-switch work has separate local live evidence: DeepSeek → Haiku → actual switch_llm → DeepSeek, then restart with the tool-selected profile preserved.
Server examples, subscription smoke and other diagnostics
The current server manual:llm example uses real OpenAI profiles, REST settings/secret lifecycle, separate workspaces and conversations, tool-driven file work, events/forks and plaintext-key checks. It defaults to gpt-5-nano/gpt-5-mini with iteration limits 12/6; it does not share the SDK smoke’s exact request/output budget. A local Gemini example exists, but the GitHub job does not run it. The latest successful server Examples run found during this audit was an older July 14 direct-local smoke, not proof of a current full profile-workflow run.
npm run manual:subscription --prefix packages/openhands-agent-server uses an already connected SDK-owned OpenAI account. It is a local temporary-state validation/two-turn smoke, not an API-key test or GitHub subscription job. A September 15 local pass is recorded separately. It consumes subscription usage.
The older SDK examples workflow, keyring smoke, and Responses reasoning diagnostic have different triggers, logging and failure behavior. The reasoning diagnostic writes private request/response artifacts and can skip successfully when its key is absent. A green run there is not interchangeable with the strict cache test. Read each script before using it.
SDK CI · SDK Live LLM · live-check runbook · SDK Examples · server Examples workflow · server profile scenario · subscription smoke.
08Rules for changing this part of the system
- Give each decision one owner. Product policy chooses role/scope names. The server owns its catalog and durable conversation selection. The SDK owns schema, credential abstractions, native APIs and safe execution boundaries. Bridges should not acquire provider-specific branches.
- Select profiles at the product boundary. REST/product callers pass saved references. Low-level provider clients remain advanced/testing surfaces. Do not copy provider settings into every channel configuration.
- Keep authentication private. Persist credential references, never raw keys or OAuth material, in profile records, settings, snapshots, events or notebooks. Do not use provider failures as a reason to spill raw request/response objects into logs.
- Honor provider behavior directly. TypeScript owns compatibility that Python may obtain through LiteLLM. Put a quirk in its provider serializer/parser or a small shared capability helper; add a fixture at that boundary. A provider fix can be required even when Python has no corresponding change.
- Switch only at complete step boundaries. Finish tool observations before replacing the model. Persist intent and activation; preserve history and metrics. Test both a continuing run and a switch in a final tool batch, including restart and failure paths.
- Preserve Python’s observable contract. Read the pinned source and port its relevant tests first. Classify a bounded
OLD_PIN..NEW_PINinterval; do not treat “latest upstream” as an unbounded implementation target. - Document deliberate differences and missing work honestly. Register named
DEV-*/EXT-*/EXC-*policy in the canonical manifests/contracts. Use trackedDEFERREDfor required behavior that is absent. Correct current evidence without rewriting frozen historical reviews. - Change the SDK at its source. Merge
smolpaws/openhands-agent, then reproducibly re-vendor it throughscripts/vendor-openhands-agent.sh. Do not hand-edit the built copy under the server’svendor/directory. - Prove each claim at the right layer. Provider wire fixtures prove normalization; pinned differential tests prove parity; real-server tests prove persistence; credential-gated live tests prove provider viability. A successful network call alone is not parity evidence.
- Finish the maintenance chain. Run relevant tests and typechecks, review feedback, merge upstream, synchronize the corresponding
enyst/fork and clean local branches to exact commit IDs, and update Beads. Deploy only after conversations and queued work settle, preserving durable state.
| Policy | Meaning to preserve |
|---|---|
DEV-SDK-003 | Secret references replace persisted raw provider keys. |
DEV-SDK-004 | Profile-first configuration and host mapping for the optional switch tool. |
DEV-SDK-005 | ACP execution is excluded; compatible data schemas do not enable its runtime. |
DEV-SDK-007 | Native durable usage accounting with explicit coverage. |
DEV-SDK-008 | Project incompatible opaque history out of outgoing requests across profile switches. |
DEV-SDK-009 | Use recorded request boundaries to place concurrent input after its preceding reply in model requests; preserve saved arrival order. |
EXT-SDK-003 | Host preparation at complete model/tool-step boundaries. |
DEV-SERVER-009 | Durable named-profile switching and host selection at the shared boundary. |
SDK compatibility contract · canonical pin and policy registry · server rules · maintenance notebook.
09Current limits to keep visible
Role names and AgentProfile schemas can exist before their consumers
Only the product’s agent role is connected. Condenser/oracle entries are accepted configuration, not working consumers. The default server’s ask_agent and condense routes remain unimplemented. Agent settings may describe an enabled summarizing condenser without constructing one. Related role/consumer work is tracked in smolpaws-w2d.7/smolpaws-w2d.8; the broader profile review remains smolpaws-jz7.
AgentProfile CRUD, activation and materialization routes exist, but activation currently stores active_agent_profile_id. The default launch path consumes agent_settings and its LLM reference; it does not resolve that active AgentProfile into a working full preset. Native switching also does not implement ACP or delegated-agent rebinding. It does not add LLM-generated titles; the server currently derives titles from the first user message.
Catalog edits are not reference migrations
Rename updates the saved ID and matching server-default reference. It does not rewrite models.json, existing snapshots or move profile-key SecretStore entries. Deletion removes the saved record and its profile-key override, clears the active marker, and can leave agent_settings.llm_profile_ref pointing to the deleted profile. A subsequent creation then needs a valid reference/default. A saved conversation snapshot still requires usable credentials when its client is reconstructed.
Some schema options and capability flags are narrower than they look
maxInputTokens is currently metadata, not an enforced input limit. timeoutSeconds is serialized only as Chat Completions body timeout; it does not create a transport abort deadline. Host/test wrappers can impose their own timeouts. Provider option support differs; for example, native Anthropic supports top-k, OpenAI clients ignore it, and Gemini Interactions rejects it.
llm_api_key_set primarily reflects the active profile key’s presence, not universal provider fallback or subscription connectivity. The REST compatibility field supports_runtime_model_switch still reports false; use the implemented native tool/config path, not that field, to understand current switching. No new HTTP model-switch endpoint was added.
Scope keys come from registration, not a naming guess
Selection literally joins platform + ":" + scopeId. WhatsApp supplies its configured folder, producing whatsapp:main. A native conversation currently registers scope ID agent-server:<uuid>, yielding selector key agent-server:agent-server:<uuid>. Generic relay fallback scope IDs may also already contain a platform prefix. Inspect registered lane identity when configuring another bridge; there is no wildcard matching.
Snapshot and accounting guarantees have explicit boundaries
A persistence cleanup error can happen after pending intent or the effective snapshot has been saved. Restore follows that durable state; do not infer rollback solely from an error observation. The origin digest deliberately excludes keys, headers and URL credentials/query values. Credential/header/query-only changes under the same profile identity therefore cannot distinguish a new opaque-history origin. The schema is not a general secret scanner; keep secrets out of arbitrary configuration strings.
Standalone or auxiliary client calls are not automatically attributed to conversation metrics. Earlier transcript text cannot reconstruct absent usage, and importing Python base_state.stats is not implemented. A server fork resets metrics by default through an explicit ledger marker; reset_metrics:false preserves them. Ordinary model switching does not reset metrics.
Actual launch path · rename/delete behavior · runtime projections · unimplemented routes · accounting limits · option serialization.
Source map and related pages
The repository contracts are authoritative. This page explains their relationship at the revisions in the header; environment settings are a dated observation, not a permanent guarantee.