← Engel's code design notebook

Implementation guide · 5 October 2026: How SmolPaws resets its context explains the implemented behavior, configuration and event sequence. This page preserves the September proposal; its pending decisions and task status are historical.

SmolPaws · context and memory

A fresh context.
A note from yourself.

Give the agent advance warning, let it save what matters, and let it call condense() when ready. A new condenser clears the old active history. The agent resumes through its own tool call, a recovery instruction, and an optional message to its future self.

Proposal · not implementedUpdated 2026-09-20smolpaws-te15SDK first → server → pilot

Let the agent prepare its own return

Engel's proposal starts from a simple expectation: the agent already has memory, daily notes, and tools for reading and writing them. Before clearing context, it should use those facilities to record unfinished requests, decisions, locations, and whatever else it will need to continue. The new tool can also carry a short personal handoff: a pointer to notes, the task it was doing, or simply “hi from your past self.”

Selected behavior

An opt-in condenser type, provisionally agent_reset, clears prior active history after an agent-originated request. It makes no summarizer LLM call. There is no retained old prefix, recent-history tail, or generated summary on this voluntary path. A separately configured emergency summarizer can recover after an actual context-window error from the main provider.

What persists

The full event log, conversation identity, files, saved notes, profiles, usage accounting, and normal fixed context remain available. “Clear” describes what enters the next model request. Historical events remain stored and inspectable.

The September 19 proposal replaced the original summary-style tool handoff with a full voluntary reset. On September 20, Engel selected a separate, explicitly configured hard-condensation fallback. That fallback produces a summary; the voluntary tool path still does not.

Four chances to prepare

75%Prepare notes
80%Reminder
85%Less room remains
90%Prepare to reset soon

The new condenser owns a configurable ascending list of warning thresholds. These defaults mean percentage of the resolved effective main-model input budget, counting the projected history, fixed context and tool declarations. They do not mean output tokens or the separate summarizer model's window.

Warnings are visible to the agent and ask it to save notes before calling condense. The agent alone chooses when to call the tool. Crossing a warning threshold, exceeding a local token estimate, or reaching an event-count threshold does not automatically condense or block an otherwise valid model request in this mode. The agent can also call the tool early.

Recommended warning lifecycle

Emit each level once per reset cycle and persist that progress across restarts. If one step jumps from 74% to 87%, emit one 85% warning and mark the lower levels passed. A committed reset rearms the levels for the next cycle. Unknown budgets or unavailable token counts remain explicitly unavailable; do not invent a percentage. Current local counting is an estimate, so leave headroom for notes and recovery.

The existing CondensationRequest is an internal trigger and is excluded from the model's view. A warning therefore needs explicit model-visible rendering. The separately configured provider-error fallback below handles emergency recovery; warning thresholds alone do not invoke it.

Two user-facing options

Keep the existing condenser choice and add a separate optional hard_condenser. The first selects the agent-controlled reset and its warnings. The second supplies the LLM summarizer used only when the main provider rejects a request for exceeding its context window.

OptionWhat the user configures
condenserSelect agent_reset and warning thresholds, defaulting to 75%, 80%, 85%, 90%. This path needs no summarizer LLM profile.
hard_condenserOptionally select llm_summarizing and reference an existing LLM profile for emergency hard condensation. This is an independent setting; it does not inherit the main model or activate merely because a profile exists.
{
  "condenser": {
    "condenser_kind": "agent_reset",
    "warning_thresholds": [0.75, 0.80, 0.85, 0.90]
  },
  "hard_condenser": {
    "condenser_kind": "llm_summarizing",
    "llm_profile_ref": "my-condenser-profile"
  }
}

Illustrative future configuration, not accepted by current settings. Field names are provisional; these example warning values are fractions of the effective input budget. The profile reference is a placeholder for an existing saved profile, not a secret or an inline model definition.

Omitting hard_condenser or setting it to null disables automatic fallback. The agent can still use its reset tool, but a provider context-window error cannot trigger an unconfigured summarizer. If hard condensation is configured, validate and resolve its own profile, preserve that binding across restore, and expose the existing bounded retry controls as advanced settings. keep_first, event-count and proactive token triggers are not controls for this full-view fallback.

The initial change applies these two options to agent-controlled mode. Existing standalone summarizer configurations retain their current behavior. Exact validation and host/API presentation belong in the contract and server tasks; no condenser pipeline is part of the proposal.

The agent calls condense

condense({
  message_to_future_self: "Hi from before the reset. Read the daily note I just saved; the next task is recorded there."
})

The parameter is optional free-form text. Field spelling is provisional; its meaning is the agent's own message, preserved verbatim. The tool does not write daily memory on the agent's behalf, discover missing notes, or ask another model to select what matters. Memory writes should have succeeded before it is called.

The tool records a correlated condensation request for the configured condenser. The new reset implementation handles that request at a safe step boundary and records a durable Condensation. Existing LLMSummarizingCondenser remains a separate supported configuration with its ordinary summarization behavior. Whether this new tool is exposed for that configuration is a contract decision; it must never silently discard a supplied future-self message.

Recommended first-version constraint: call condense as the only tool in its response. A mixed or duplicate tool batch receives an explanatory rejection for the reset request. Retaining one action while dropping its sibling actions would break existing batch integrity. Ordinary sibling execution follows the existing tool policy.

Exactly what the next context starts with

  1. Fixed context

    System prompt, identity, configured memory context, runtime/channel information, skills, and tool definitions, composed as usual.

  2. 1 · user

    The agent triggered context condensation.
    Exact synthetic notice. Its provenance is the environment; it is not a new message from the human.

  3. 2 · assistant

    The agent's genuine condense tool call, including its optional message. Preserve its real call ID and required reasoning/provenance.

  4. 3 · tool

    The matching result: confirmation of the agent-triggered reset, recovery instructions, and the message to its future self.

  5. Then · new input

    Any genuine user input that arrived after the model request which authored the reset. It was not available when the agent prepared the reset and must survive, in order.

Proposed tool-result wording

You triggered context condensation. Your earlier active history has been cleared; your notes and the stored conversation history remain available.

Take a deep breath. Find and read your notes to regain your bearings, then continue the user's work. The user may not know what just happened. Consider whether a brief heads-up or a little joke would help; use your judgment.

Message from your past self:
[the exact optional message you supplied]

The optional message block is omitted when no message was supplied. This wording guides the agent; it does not automatically send a message to the user. There is no CondensationSummaryEvent in the reset view.

Memory detail: SmolPaws currently snapshots fixed context for a conversation. Rendering that context again does not reload freshly edited memory files. This initial proposal keeps the snapshot and lets the agent read its newly saved daily notes through its tools. A refresh of trusted memory snapshots would need a separate explicit policy.

New tool payloads inside existing event kinds

The host already has conversation.condense(); the Python SDK and our pinned port do not expose an agent-callable condensation tool. Existing request events carry no notes or originating tool references, and the current view reduces requests to a flag. A thin wrapper around that API would not implement this proposal.

RecordProposed responsibility
ActionEvent
CondenseAction payload
The genuine condense call, validated by a new tool-specific input schema, including the optional message_to_future_self. Keep its tool-call identity and original arguments.
CondensationRequestVersioned request metadata distinguishes an agent tool request, an emergency fallback and a host command. Only the agent request references a genuine originating tool call; the current untyped request flag is insufficient.
ObservationEvent
CondenseObservation payload
A new tool-specific output schema for the matching result, recovery text, future-self message and explicit failure outcome. Successful recovery text is exposed to the next model call only after the associated reset commits.
CondensationThe committed removal set and operation metadata. Voluntary reset: summary: null and references to the real tool pair. Emergency fallback: generated summary and fallback provenance, without an invented tool pair.
Synthetic MessageEventDeterministic environment-origin, user-role notice at the front of the new view; distinct wording for voluntary reset and emergency fallback.
CondensationSummaryEventOnly the summarizing fallback renders this event, after its environment notice. It is absent on the voluntary reset path.

Yes, the tool needs its own Action and Observation schemas. The proposed names CondenseAction and CondenseObservation are payload types, following existing tools such as SwitchLLMAction and SwitchLLMObservation. In TypeScript these are validated schemas and inferred types. They travel inside the existing ActionEvent/ObservationEvent envelopes; separate top-level CondenseActionEvent/CondenseObservationEvent kinds are not required. The previous table omitted this distinction. Typed, versioned reset metadata is also needed; its exact carrier remains a contract decision. A reset-aware projection must express both the voluntary notice/tool pair and the emergency notice/summary.

Commit the boundary without losing arrivals

Capture the history boundary consumed by the model that authored the tool call. Record intent without recursively running the conversation from inside its executor; that could wait on its own step lock. Settle the step, durably commit the tool result and reset operation with defined crash recovery, and only then make the next model call. Clear only the eligible earlier history. Messages arriving during note-taking, model completion, tool execution or reset must remain pending until a subsequent request actually consumes them.

Replay is a projection, not tool re-execution

The durable log keeps chronological source records. Replay reconstructs the new view from stable operation references and synthesizes one deterministic notice. It must not execute condense again, duplicate the pair, revive cleared history, reset accounting, or declare a failed reset successful. Repeated resets and forks need the same explicit rules. A crash between intent, observation and reset commitment needs defined completion or failure behavior before another model request is allowed.

Provider histories must stay valid

Verify assistant/tool adjacency, real call/result identity, response-batch integrity, and signed or encrypted reasoning handling across native clients. The synthetic user notice must remain distinct from genuine incoming user work. Existing profile-origin and concurrent-request projection rules still apply. The reset must not manufacture a claim that the agent authored a call which never occurred.

TRANSPILE treatment: tool/schema additions may be additive extensions; changed request handling, view replay and prompt ordering affect parity-governed behavior. Register the appropriate stable policy in TRANSPILE_CONTRACT.md and the canonical manifest, with tests and source evidence. Review and merge the SDK first, then vendor that exact commit into the server and update its contract.

If the agent runs out of room first

Selected September 20: only after catching an actual, recognized context-window error returned by the main provider, invoke the explicitly configured emergency LLMSummarizingCondenser.hard_context_reset(). The agent decides when to condense; warnings help it prepare. An estimated hard ceiling, event count, ignored warning or judgment about note readiness does not trigger fallback. Even an estimate at or above 100% is not a substitute for a provider error. Unrelated authentication, rate-limit and transport errors retain their own error handling.

Configure both strategies

The user must configure the agent-reset condenser and a separate emergency summarizer, including its LLM profile and bounded retry settings. Merely having a condenser profile in the catalog does not enable the fallback. A reset-only configuration makes no summary calls; without the fallback, a caught context-window error stops the affected run with an actionable error and retains the source history.

Simple error recovery

The normal path warns the agent and waits for its condense() call. The main-model error handler catches a recognized context-window error and invokes the configured hard-condensation fallback. No pipeline or general-purpose condenser composition is needed. Local token and event thresholds do not activate the fallback.

What “hard” means in the Python SDK

A HARD condensation requirement first tries ordinary condensation, retaining its configured prefix and recent tail. If ordinary condensation raises NoCondensationAvailableException — for example, because no valid cut exists or summary generation fails — RollingCondenser calls hard_context_reset(). That method summarizes all events in its input View, forgets them and inserts the summary at offset zero. It already ignores keep_first. It retries failed summary generation with shorter event representations, subject to a configured bound. This proposal selects that full-view method directly for its emergency fallback; it does not mean “set keep_first to zero and run ordinary condensation.”

The recovery context starts with:

  1. Fixed context

    The normal system, identity, memory snapshot, runtime information, skills and tools.

  2. 1 · user

    The environment triggered emergency context condensation after the model provider returned a context-window error.
    Earlier active history was replaced by the summary below. Your notes may be incomplete. Take a deep breath, read the summary and your notes, and check unfinished user requests before continuing. The user may not know this happened.

  3. 2 · summary

    The fallback’s CondensationSummaryEvent, rendered in its existing user role after the environment warning. It is a generated summary, not a message authored by the agent’s past self.

  4. Then · new input

    Genuine pending user input protected by the request boundary, followed by arrivals during fallback processing. Preserve its original text/media and order.

There is no retained historical keep_first prefix and no fabricated condense Action/Observation pair on this automatic path. The new notice is separate from the summary and must be replayed exactly once with truthful environment provenance.

Failure, concurrency and accounting

Apply the upstream hard-reset algorithm to the eligible active View while protecting unseen user input separately. The added notice, pending-input protection and routing are explicit target policies, not claims of unchanged Python behavior. Token estimates can describe the rebuilt context but must not become another automatic condensation or request-blocking threshold. If configuration, bounded summary attempts or persistence fails, or the provider still rejects context after the bounded recovery, retain the source log and stop with an actionable recovery error; do not loop indefinitely or silently perform an empty reset. Repeated provider rejection caused by fixed context alone cannot be repaired by forgetting more history. Record each actual fallback LLM attempt as condenser usage. Do not repeat a paid completion merely because accounting persistence failed, and rearm warnings only after a committed recovery.

“All events” describes the selected input View, not a promise of lossless full-text transfer: the current summarizer uses event previews, and retry clipping can shorten them further. The existing preview-loss bead remains relevant.

Decisions to settle before implementation

QuestionBoundary for this proposal
Fallback configuration and routingThe two-option shape is selected above; finalize field names, validation and host/API presentation. Specify recognized provider context-window error classification, protected-input boundary and operation metadata. Both strategies must be configured to enable fallback; local token estimates, event pressure and unrelated provider failures must not invoke it. The current agent catch also handles malformed conversation history; this new path must distinguish that error from actual context-window overflow, including wrapped errors.
Host or channel /condenseA human command has no genuine agent-authored call/message. Decide how the new mode handles it. Existing summarizer commands continue their current contract; do not fabricate an agent tool call.
Optional message limitChoose a bounded size and an actionable rejection before any destructive view change. The limit must leave room for fixed context, tool exchange and recovery.
Metadata and warning renderingChoose versioned persistence, event validation, restart behavior and model-visible warning form. Keep synthetic maintenance notices separate from real user input.

The agent owns the content of its notes; the harness cannot prove they are sufficient. If notes are missing, unreadable, or incomplete, the recovery instruction should lead to inspection or an honest question to the user. Reset cannot repair a fixed prompt that already exceeds the budget.

Tracked work, in dependency order

smolpaws-te15 remains the proposal's parent. All seven implementation tasks are open. Publishing this page and filing the beads does not implement or enable the feature.

te15.1 contract → [te15.2 tool/reset + te15.3 warnings] → te15.7 emergency fallback → te15.4 SDK review & merge → te15.5 server → te15.6 verification & controlled rollout
BeadDeliverableBlocked by
smolpaws-te15.1Event, lifecycle and TRANSPILE contract; settle open policy decisions.—
smolpaws-te15.2SDK tool and reset projection, with red/green evidence for source-history and late-input preservation..1
smolpaws-te15.3Configurable staged warnings, persisted crossing state, available/unknown measurements..1
smolpaws-te15.7Explicitly configured fallback after a caught provider context-window error, environment warning before summary, bounded recovery and source-history preservation..2, .3
smolpaws-te15.4Substantive SDK review, checks, documentation and merge..2, .3, .7
smolpaws-te15.5Server vendoring, both condenser settings and tool wiring; resolve the independent summarizer profile only when fallback is explicitly configured..4
smolpaws-te15.6Protected-server/relay tests, bounded live recovery check and separately authorized scoped rollout..5

Required evidence includes the complete warning → note write → tool → reset → note retrieval → continued task path; concurrent text/media; crash/restart at every persistence boundary; repeated reset; missing notes; mixed calls; exact native provider message ordering; and warning rearming. Also cover ignored warnings → actual provider context-window error → emergency fallback → notice/summary → continuation. Prove that estimates at/above 100%, high event counts and unrelated provider errors do not trigger fallback or preempt the agent’s decision. Cover reset-only configuration, missing fallback credentials, exhausted retries, oversized fixed context and user arrivals during fallback. Live cost and actual continuation must be reported separately from deterministic tests.

Related work remains separate: smolpaws-d3fw (existing 1M/400k budget snapshots), smolpaws-29dr (summary preview loss), smolpaws-neo (history invariants), and smolpaws-eh4z (Fast-Jev evaluation). None silently changes this proposal.

Source anchors and scope

This is a dated design proposal based on Engel's September 19 instructions and September 20 fallback clarification, the prior bead, and inspected source. Authorship and note selection are motivations to test; a message-role change alone has not been demonstrated to solve continuity.