SmolPaws · production cutover · September 17, 2026 · canary evidence retained below

WhatsApp is on the shared server.

Main, OpenHands and Hunting now use the normal TypeScript product host on :8790, with their existing conversations and memory snapshots. The temporary WhatsApp host on :8791 is retired. The September 15–16 text, image, voice and scheduled-reply observations remain dated evidence below; they are not a fresh visible-message test of this cutover.

Bridge implementation: SmolPaws #172 and SDK #30. Subscription follow-up: SDK #31 · 775869e and server #173 · e634d0b. Beads owns progress: smolpaws-zlo / smolpaws-kxa.

Normal operation, same conversations

The standalone com.smolpaws.bridge.whatsapp LaunchAgent connects to the shared SmolPaws product host on :8790. Main, OpenHands and Hunting keep their existing conversation IDs, history, configured profiles, context snapshots and scheduled work. The canary server and bridge LaunchAgents are retired; the legacy WhatsApp process remains disabled. Other consumers of the legacy server must be migrated before that server can be removed.

The linked-device credentials and message journal remain in ~/.smolpaws/whatsapp/. Group settings live in ~/.smolpaws/whatsapp/registered_groups.json; OpenHands and Hunting retain mention-free operation. The relay uses ~/.smolpaws/coordinator/whatsapp-relay-v1.db, alongside the shared scheduler. Saved relay ownership, intake and delivery progress move with the database, so the cutover does not replay completed messages.

Cutover verified September 17 at 02:48:56 Europe/Stockholm (00:48:56 UTC): the normal WhatsApp bridge connected to the shared product host on :8790, running SmolPaws dc064a2. All 23 former canary conversations joined the 20 existing shared-host conversations, for 43 total. The migration preserved all 2,580 event files and 20 context snapshots byte for byte; saved profiles, workspaces and accumulated metrics matched.

All 29 scheduler tasks and 15 run records were retained. Main, OpenHands and Hunting kept their conversation IDs and profiles. All 201 WhatsApp relay work rows were complete. Nothing was listening on :8791, and the canary LaunchAgents and legacy WhatsApp process were off. This verifies the operational handoff and preservation; it does not claim a fresh user-confirmed reply after cutover.

Where the conversations are saved

New-stack conversations now live under ~/.smolpaws/conversations/<uuid>/events/, as indexed event JSON files. Each conversation keeps its metadata and private smolpaws-context.json snapshot beside that directory. The former canary storage is retained as a backup, rather than an active server.

Legacy Main transcripts remain separately at ~/.openhands/conversations/main-*/events.jsonl. Promoting the new bridge preserves its current conversations; it does not convert the legacy transcripts into the new event store.

Initial canary proof

The September 15 evening canary began with SmolPaws c4ba8c4, SDK 5f28eb8, and the existing deepseek-v4-flash profile. That implementation passed 61 coordinator tests, 6 product-host tests across the relay platforms and full server-package CI. Real-provider multi-tool execution and continuation after restart also passed. These revisions identify that observation, not subsequent deployments.

Historical canary scope · September 16 at 00:46 CEST: Main, OpenHands and Hunting are enabled on the existing account. The running bridge reloaded all three registrations at 00:46:35.891 Europe/Stockholm (September 15, 22:46:35.891 UTC), without a restart. OpenHands and Hunting each retain their legacy triggerFree: true setting. The active registered-groups JSON, selected by SMOLPAWS_WHATSAPP_REGISTERED_GROUPS, owns these per-group settings. This was a configuration update.

At reload, the two added groups had no saved messages awaiting dispatch and no active legacy schedules. The product host on :8791 was healthy on SmolPaws 28106cd and SDK a983e4c.

OpenHands transport proof · September 16 at 00:48–00:49 CEST: two real user messages each produced one completed intake and one completed delivery, with one processing attempt, one send attempt and no recorded error. WhatsApp accepted the two replies at 00:48:55.483 and 00:49:24.329 Europe/Stockholm (September 15, 22:48:55.483 and 22:49:24.329 UTC), with distinct external message identities. This verifies the real ingress-to-reply path through transport acceptance. Hunting had received no new messages at that observation. Later Hunting messages reached the agent and exposed the run-limit issue; its unfinished work must be assessed separately from the transport repair.

Initial overnight canary · September 15 at 18:49 UTC: the bridge began with only Main allowlisted. Legacy WhatsApp was stopped and disabled to preserve one device owner across reboot; it remains available through the tested rollback procedure. Existing Main tasks were preserved. Other local servers and ingress services were left unchanged at that handoff.

The user confirmed exactly one NIGHT-CANARY READY and one NIGHT-SCHEDULE OK. Each has one completed delivery row, one send attempt and an external message ID. The scheduled test originated in an operator request through the real Main conversation. Earlier iPad text arrived through actual WhatsApp ingress; image and voice were queued while disconnected, then delivered after reconnect and confirmed on the iPad.

These are bounded observations, not ongoing telemetry. smolpaws-957 tracked overnight health, real replies, unsettled work and the existing Main cron due September 16 at 07:00 UTC. The later permanent cutover is recorded above; other ingress migrations remain separate work.

Startup notices · September 16: after #190 · a4304ab, the WhatsApp bridge restarted with startup notices enabled. WhatsApp accepted exactly one smolpaws: 🐾 I'm up. send in each of Main, OpenHands and Hunting. A fresh bridge process announces once per destination; ordinary reconnects do not repeat it. This fixed operational message makes no LLM call, and a failed destination does not block the conversation outbox. At that observation, the agent-server was not restarted and reported 02a8187; all 407 canary event/context files and live projection cursors were preserved, with the 500-step default unchanged. Slack and Discord implement the same behavior for explicitly configured destinations. The Slack process inspected at that time had no startup destination list, and this observation does not establish a live Slack notice or Discord migration. Startup-notice configuration and delivery rules.

Three small fixes from the live trial

The first iPad input reached durable intake but exceeded the upstream launch request’s 32,768-character suffix limit. SmolPaws #175 worked around that limit by referencing oversized local files for the agent to read. The merged context-file update uses the SDK’s existing always-on skill path to supply full contents; the request limit stays unchanged.

The first reply appeared twice because the agent both sent a message and finished with identical text. SmolPaws #176 compares final text with durable messages already queued in that turn. It suppresses the final echo while preserving different final text, explicit repeated sends and later turns.

A multi-tool response also copied reasoning metadata into every ActionEvent, breaking the next history reconstruction. SDK #32 restores the pinned Python rule: only the first action carries that metadata. This is a missing transpilation behavior, separate from provider compatibility. The server re-vendors the canonical SDK; malformed older history needs explicit reconciliation without repeating completed effects.

The 18:19–18:30 UTC trial delivered text, image and voice, then rolled back with scheduled tasks unchanged. The fresh evening canary verified both repairs: a real two-tool response survived continuation and restart, and deliberately identical send_message/finish output produced one visible reply. Offline replay of the captured iPad intake at the relay acceptance boundary preserved its completed row without another server call or send.

Conversation errors are visible; runs use the normal limit

Merged and deployment verified September 16 · SmolPaws #188 · 02a8187 · smolpaws-wws

The initial canary used an explicit 12-step limit per run. That temporary setting remained during ordinary use, so a task could stop after recovering older context without producing a reply. It was not a provider requirement or a context-window limit. That September 16 repair set the :8791 settings and four then-existing WhatsApp conversations to the normal 500-step limit, retained after promotion to :8790. A new user prompt gets a fresh run budget while keeping the same saved conversation.

The shared bridge extractor turns each newly projected ConversationErrorEvent into a concise message. A limit error says, “I stopped because this run reached its 500-step limit. Send another message to continue.” The number comes from the actual limit event. Other conversation errors say, “I encountered a conversation error. Send another message to try continuing.” WhatsApp adds its normal smolpaws: prefix. Raw provider details stay out of channel messages; a conversation error need not mean every part of the run stopped.

The agent-server records an event when a run throws, including during provider calls or agent creation, so those errors reach the same delivery path. The relay selects the conversation-error type, not a list of provider error strings. Recoverable tool errors remain available to the agent and do not produce false conversation-error notices. WhatsApp, Slack and Discord use the same extractor.

A notice uses the error event’s identity in the durable outbox. Restarting or replaying a projection cannot send that event twice, while a later error gets its own notice. Normal follow-up prompts resume saved history; completed tools are not automatically replayed. Deployment does not rewind live cursors, backfill historical error notices or automatically restart unfinished turns.

Deterministic proof: one test runs a real local agent-server with a scripted LLM and a fake WhatsApp socket through five turns: a step-limit error, successful continuation, another limit error after bridge restart, a thrown provider failure, and successful continuation. It verifies distinct notices, durable replay suppression, preserved original context and no replay of completed tools. This tests the full bridge path without sending synthetic messages to the live account.

Deployment proof: the shared product host on :8790 and WhatsApp host on :8791 both reported 02a81875f97b26b19c391351c2544803ddc218cb, and the canary WhatsApp process restarted on the update. The 500-step settings and four stored WhatsApp conversation limits were verified. All 429 pre-existing event/context files and the live projection cursors were preserved; conversation counts remained 20 on the shared host and seven on the canary. Producers resumed with empty queues and legacy WhatsApp stayed unloaded. No user turn was replayed. Slack’s existing process was paused and resumed, so this does not claim its new extractor is live.

Validation: full server-package CI passed with 114 tests, alongside 65 coordinator tests, 22 WhatsApp tests and the merged pull request’s GitHub checks.

Remaining work: normal conversation use and scheduled work still need observation. The user’s follow-up already resumed Hunting with its previous context; this repair does not claim that Hunting’s requested work is complete.

Conversation-error contract · Real-server WhatsApp regression.

Full context files, selected by the product host

Merged and deployment verified September 16 · SmolPaws #186 · d7b351a

The SmolPaws host selects identity and memory files from ~/.smolpaws/context.json, or the path selected by SMOLPAWS_CONTEXT_CONFIG. Version 1 has common files and additional scopes keyed by exact platform:scopeId. Omitted files uses the public identity defaults; [] supplies no common files. Private memory belongs in the intended scope’s list. Selection uses the registered scheduler lane, so request tags cannot select another scope’s context.

Each selected file is an always-on, non-AgentSkills Skill with trigger: null. The SDK renders its full body in REPO_CONTEXT, trimming outer whitespace without truncating the body or requiring a file-reading tool call. The agent_launch_additions.system_message_suffix_append request retains its 32,768-character cap.

A private smolpaws-context.json snapshot beside the conversation records the exact file contents and hashes. Later turns and restarts reuse it before consulting configuration or source files. Existing conversations capture their first snapshot on their next run through the updated host; edits then affect conversations without a snapshot, normally new ones. A missing explicitly selected configuration, unreadable selected files and corrupt snapshots fail the run. A fork gets a fresh snapshot for its new native lane; it does not inherit the source scope’s private file selection.

Deployment proof: both the shared product host on :8790 and WhatsApp host on :8791 reported d7b351a. Main, OpenHands and Hunting retained their existing conversation identities. Snapshot inspection confirmed each had nine selected files, with the complete configured private memory present exactly once, and snapshot permissions 0600. All 329 pre-existing event files kept their hashes. Producers resumed with empty queues; the existing Main cron remained active, next due September 16 at 07:00 UTC. Its due-run result remains part of overnight observation.

Isolated live proof: DeepSeek Flash returned a random marker available only in a 58,827-character configured test file. Its only action was finish; it made no file-reading call. A second turn succeeded after the source file was deleted. This proves the first request received the full context and continuation used the snapshot; restart behavior is covered separately by deterministic integration tests. No synthetic WhatsApp messages were sent. The later Main and Hunting iteration-limit follow-up is recorded separately; context delivery alone did not resolve unfinished work.

Upstream’s separate opt-in load_memory/memory_context feature, with its default 6,000-character combined user/project memory-index budget, remains deferred under smolpaws-45n, together with server preference propagation and full AgentProfile skill discovery. Canonical SDK context and memory notes distinguish the existing skill support from that missing port.

Context configuration, snapshot lifecycle and ownership · SDK documentation correction #37.

One execution core; channel reliability around it

The TypeScript SDK supplies the agent and provider clients. The separate TypeScript agent-server supplies the upstream-shaped API and conversation persistence. Bridges and the relay own platform intake, delivery and retries.

Each platform owns a separate relay work database, while task scheduling is shared. The trial exposed an inherited bridge override that could make the native API worker claim bridge deliveries. The isolation fix gives the native worker its own SMOLPAWS_AGENT_SERVER_RELAY_DB_PATH; SMOLPAWS_RELAY_DB_PATH remains bridge-specific. The trial used separate stores before opening the socket.

WhatsApp message path through the new stackWhatsApp bridge and ledger send durable intake to the Message Relay, which communicates with the TypeScript server and SDK. The server EventLog returns through the relay outbox to WhatsApp. Ledger, relay and server persistence are distinct state boundaries. WhatsApp bridgesocket + chat policyledger + cursorsMessage Relaydurable intake / outboxlane → conversationTS server + SDKagent + provider clientconversation EventLog accepteventsEventLog → outboxdeliverSeparate stores. A fresh relay DB does not imply a fresh server conversation.
  1. WhatsApp bridge: socket, chat policy and ledger.
  2. Message Relay: durable intake, lane binding and outbox.
  3. TypeScript server + SDK: agent, provider and EventLog.
Implementation: RelayRuntime, Message Relay design. WhatsApp/Discord identity derives directly from lanes; do not add a versioned namespace to solve deployment isolation.

The controlled-canary gates are complete

01

History and state handoff — implemented

smolpaws-kxa.6 · deterministic tests passing

The actual legacy data/router_state.json is imported once into a journal of handled message identities. Both updated host generations use it, preserving pending and late arrivals, equal timestamps and duplicates across rollback and re-cutover.

One process lock owns the linked device. The entrypoint exposes relay, allowlist and startup-ping controls; SMOLPAWS_HOME_DIR now relocates the default relay store. A persisted relay-owner tag rejects adoption of an unrelated server EventLog, and startup rejects stored destinations outside the allowlist.

Use the latest production ledger, relay, scheduler and conversation state; never restore a stale canary snapshot over newer progress. Drain and stop the bridge before npm run whatsapp:handoff -- legacy. It verifies settled work and projected events, refuses unfinished scheduled runs, exports task changes and marks the ledger ready for the updated legacy host. The command sends nothing and does not start that host; the operator starts it only after handoff succeeds. Existing pre-change processes have no lock: stop them explicitly before transferring socket ownership.

Full handoff and isolation procedure · History regression tests.

02

Engine and deployed host — verified

smolpaws-kxa.7 · verified September 15

The fixed SDK is now reproducibly vendored. Verify the deployed revision, state paths, active profile, credential availability and per-scope workspace. Start the SmolPaws product host: it composes the transpiled server with real scheduler and media executors.

The new server uses its own profile state; legacy LLM_PROFILE_ID does not select it. The product host adds messaging and task tools automatically. The internal tool call and continuation passed on September 15; recheck the deployed host and its state before opening WhatsApp. The bare transpiled package CLI intentionally has no product scheduler/media executors.

Evidence: agent factory · launcher health check · re-vendor runbook.

03

Initial Main-chat canary — passed

smolpaws-kxa.5 · passed September 15; soak tracked separately

Use an explicit allowlist containing the trusted control chat, a bounded live-send window and a tested rollback handoff. For the same account, stop/drain the legacy socket owner during the test and account for scheduled work. Suppress the automatic startup ping or include it in the agreed test.

Trace a unique input across every boundary below. Then test duplicate inbound replay, restart with queued outbound, and reconnection. Reconcile delivery_unknown; never blindly retry an ambiguous send.

  1. Inbound row exists in the WhatsApp ledger.
  2. Durable intake reaches done.
  3. The expected conversation has the user event and completed run.
  4. Expected mid-turn and/or final delivery rows exist.
  5. Delivery is done, with one send attempt and an external message ID.
  6. One correctly prefixed reply appears in the intended chat.

Acceptance contract: repository readiness plan. A visible reply alone is not a pass; a text pass does not prove full replacement.

Shared capabilities and the remaining rollout

One product host binds the tools, one shared SQLite scheduler owns tasks, and each bridge submits due work through the existing relay. There is no separate channel-specific scheduling engine or new scope hierarchy.

CapabilityImplemented behavior / remaining proofBead
Scheduled workImplemented across WhatsApp, Slack, Discord and native API conversations: real scoped lifecycle tools, cron/interval/once, group/isolated runs, durable intake, restart and rollback. The September 15 live scheduled WhatsApp reply was confirmed on the iPad. Existing schedules were retained through the later production cutover.kxa.4
Outbound media / voiceImplemented: immutable file spool, native image/video/audio/document delivery, WhatsApp OGG/Opus PTT and the existing voice-outbox producer. The iPad image and playable voice were confirmed after queued delivery across reconnect. Other platforms still need their own live upload and permission checks.kxa.2
Scope context / permissionsExisting rules preserved: WhatsApp groups keep their workspace, control main can manage other registered scopes, others manage their own tasks. The merged context-file update selects private files through explicit scope configuration. This retains local workspaces; it adds no sandbox model.kxa.8
RecoveryImplemented: reconnect creation retries, stale-socket fencing, transient network guard and shared-server child supervision. Idle interrupted requests can resume; tools without an observed outcome are parked for reconciliation.kxa.9
Independent chat progressImplemented: ingress commits locally without HTTP; server requests include a body-read deadline. Disconnected sends wait; timed-out external sends remain delivery_unknown.39y
Provider usage and costsMerged and deployed September 15: SDK #35 and server #184. Provider usage is recorded once per call, with cost provenance and restoration without double counting. The isolated live check matched disk and HTTP metrics. Earlier unrecorded canary usage remains unknown. Accounting boundary and proof.via
Cutover, rollback and soakPermanent WhatsApp ownership moved to the normal shared :8790 host on September 17, preserving the three current conversations. The temporary :8791 host is retired. Current storage and verification; dated canary transport evidence. Migrate remaining ingress consumers before deleting the legacy server.b1r.24

All IDs have prefix smolpaws-. Shared scheduler/media design · Recovery and ownership.

One model configuration across channels and roles

LLM Profiles: the complete implementation notebook covers configuration, the profile and secret stores, provider APIs, runtime switching, tests and the GitHub LLM environments.

SmolPaws selects saved LLM profiles through ~/.smolpaws/models.json. Shared roles supply defaults; exact scopes such as whatsapp:main override them. A role contains a profile name, never a copied model configuration or credential. WhatsApp, Slack, scheduled work and direct agent-server conversations use the same product selector.

A changed agent selection applies to the existing conversation before its next model call, after the current call and all tools have settled. The agent’s switch_llm tool uses the same durable activation mechanism. It preserves conversation identity, history, context and accumulated usage across restart. An unchanged channel setting does not undo a tool choice; changing the applicable configured name again does. Editing another role or scope does not reset the agent.

The schema also accepts names for condenser, oracle and future roles. Only the agent role is currently connected to an LLM consumer. The LLM summarizing condenser and oracle remain separate implementation work in smolpaws-w2d.7 and smolpaws-w2d.8; saving those selections does not enable them.

The SDK owns the optional upstream switch tool, the safe step boundary and provider history projection. The server owns validated profile snapshots and persistence. The product host owns the selection file; bridges do not each implement switching. Signed or encrypted reasoning from an incompatible profile is omitted from outgoing requests while the original events remain saved. Each completion keeps its actual serving profile and model in usage records.

Isolated live proof, September 16: DeepSeek Flash → configured Haiku → an actual switch_llm call → DeepSeek Flash preserved the test context. Restart retained the tool-selected profile despite unchanged Haiku configuration, with identical old event hashes and no duplicated usage. The trial exposed DeepSeek’s requirement for a reasoning field on historical assistant tool calls; its serializer now supplies an empty field when a foreign model supplied none, while preserving real reasoning exactly.

Tracked in smolpaws-anc. Configuration and precedence · Server and product integration · SDK implementation and review · Pinned-source evidence and deliberate differences.

Provider quirks have a home in the SDK

Python gets part of its provider compatibility from LiteLLM. Our TypeScript implementation owns that behavior directly. Keep protocol normalization inside each provider client and shared capability decisions in small pure helpers; keep those decisions out of the agent loop and channel bridges.

The merged tool-call fix accepts harmless extra fields such as DeepSeek’s index while validating required fields. Its fixture tests and the new provider guide make this a maintainable compatibility rule, even without a corresponding Python SDK diff.

ChatGPT subscription support also requires the OAuth lifecycle, not just a Codex endpoint URL. The SDK owns its private credential store, refresh and transport; the server exposes device login and restores authentication for profile validation and resumed conversations. See the subscription boundary and verification requirements.

Subscription proof, September 15: after SDK #31 · 775869e merged and was reproducibly vendored, the real gpt-5.5 subscription test passed profile validation, a think/finish tool round trip and a second turn in the same conversation. It used temporary state and the SDK-owned OAuth account, without an API key or bridge sends. The existing active DeepSeek profile was not changed.

LLM provider implementation guide · Transpilation policy and provider responsibility.

Measure new usage without inventing old totals

Deployment verified September 15: the shared product host on :8790 and WhatsApp canary on :8791 both reported SmolPaws 28106cd, using the re-vendored SDK a983e4c. Existing conversations and historical event contents were preserved, and ingress resumed. Each completed LLM call in the agent loop now gets one durable record, including token and cache details, model identity and cost provenance. Per-usage histories and accumulated stats are derived from those records, including after restart. A multi-tool response counts once; a user request that needs several LLM calls records each call.

Missing data remains unknown, while measured subtotals remain useful. Cache hits and cache writes are distinct, and DeepSeek cache misses do not mean cache writes. A provider-reported cost keeps its unit; a DeepSeek estimate retains its dated pricing source and rates. The earlier canary history predates accounting, so its full token usage and cost cannot be reconstructed from conversation text.

Isolated live proof: a two-turn DeepSeek test completed with 3,317 input and 162 output tokens, 3,479 total. Input included 1,536 cache hits and 1,781 misses; the 72 reasoning tokens were a subset of output. Cache writes remained unknown because the provider did not report them. The calculated estimate was USD 0.000368958, with its dated quote retained; provider counts, saved records and HTTP totals agreed. These are test-conversation measurements, not the Main chat’s totals.

Validation included 524 SDK tests, 106 server tests, 61 coordinator tests, 8 product-host tests and the GitHub LLM-environment live test. See the SDK/server responsibility and upstream relationship and the metric fields and coverage rules. Bead smolpaws-via owns this work.