OpenHands Automations: where work runs operations note
September 17 update · bounded local review trial
The bounded local OpenHands Automation retry completed and published four requests for changes as enyst. Both local trial definitions are verified disabled. The first attempt failed before publication and posted zero reviews; its evidence remains preserved. The installed service, 1.11.1, requires temporary enablement by PATCH for manual dispatch. Both definitions have literal-false event filters and no recurring schedules; the old Cloud Roasted reviewer remains paused. The uploaded bundle uses a relative entrypoint and pins OpenHands SDK 1.46.0. A separate dependency check prompted a verified editorial correction to one published review: an unsupported finding was withdrawn while the supported request for changes remained. This supervised trial does not validate unattended review quality. All twelve completed PRs are retired from future selection. Persistent host checkpoints and a separate model profile retain trial state; Cloud KV and native-event delivery remain unresolved. Verified publications, quality correction and deployment limits.
September 16 update · review automation status
Engel's decision: recurring reviews stay off. The account's Roasted Code Review definition was already paused; the disabled setting was reapplied and verified. Its last recorded run completed on August 1. Returned histories for all eight Cloud definitions and the local runtime showed no pending or running automation runs at this check.
Astra completed two manually supervised batches: the first published four reviews as smolpaws and withheld one after the PR changed; the second published four as enyst—one approval and three requests for changes. At the September 16 check, no Astra automation was registered in Cloud or locally, and the local deployment had no review definition. The bounded local trial above is a later change; continuous review automation remains off. The upstream all-hands-bot reviewer is separate and is not controlled by this setting. The chosen voice is friendly to humans and snarky about verified bot contradictions. Verified outcomes and remaining deployment requirements.
September 15 update · weekly attention review
Deployed · Cloud input validation passed. The existing attention router now uses the weekly policy below. Cloud configuration was read back, the running source files matched the tested bundle byte for byte, and an input-only Cloud run completed successfully at 21:27 UTC on September 15.
- Schedule: Monday at 09:00 Europe/Stockholm, cron
0 9 * * 1. The local time stays at 09:00 across daylight-saving changes: 07:00 UTC in summer, 08:00 UTC in winter. - Repository scope: OpenHands/software-agent-sdk remains the current default. Expansion to other repositories is undecided.
- Base selection: the 50 newest currently open PRs, ordered by creation time, newest first.
- Mention additions: union that selection with every PR containing an actual
@enystmention from the preceding rolling seven days within the scope. These additions are uncapped and include closed or merged PRs. A PR in both groups is reviewed once. - Reading context: use each selected PR's full description and the full bodies of all linked closing issues returned by GitHub's sidebar relationship,
closingIssuesReferences. This covers manually linked issues and issues linked for automatic closure. - Weekly reassessment: rescore every selected candidate each week, including unchanged PRs. Remove the previous skip based on
updated_at. - Run limit: 30 minutes, the Cloud maximum. The collector does not silently cap mention additions or clip context; incomplete reads or failed scoring fail the run.
- Model and destination: retain
deep-pro, resolving toopenhands/deepseek-v4-prothrough OpenHands. Continue filing or updating the same Review Notebook notes.
Verification: 17 focused tests cover selection, pagination, mention timing, complete linked context and failure behavior. The Cloud check resolved the actual profile and required credentials and collected the full inputs; it made no scoring calls or notebook writes. The normal scoring entrypoint is restored for Monday, September 21 at 09:00 Stockholm time. Versioned source and tests. The historical ten-PR, eight-hour pilot below records the earlier configuration.
September 15 audit · before the attention-router change
9 enabled definitions: 3 local and 6 in OpenHands Cloud. Another 2 Cloud definitions are paused. No automation run was pending or running at the capture. Enabled means eligible for a future trigger; it does not mean the last attempt worked.
The immediate issue is local reliability. All three local jobs failed on their latest attempt. The Canvas entry point was unavailable, while the Agent Server, Automation service, and scheduler remained alive. Closing or losing the UI did not stop scheduled work.
The local backend · 3 enabled
The local service runs work on the host through the Agent Server. The two maintenance jobs clone fresh copies of upstream repositories for each run. Gazette is different: it publishes from the shared documentation checkout, which makes ordinary editing work a publishing dependency.
These schedules use Europe/Amsterdam, currently the same civil time as Stockholm. All dates below are in 2026.
| Job and target | When · budget | Purpose and permitted changes | Latest recorded result |
|---|---|---|---|
| Mention Gazetteenyst/enyst.github.io | Mon–Thu, 09:005 minutes · deterministic Python; no LLM | Groups unread GitHub notifications into a Gazette page. Commits that page and pushes documentation main directly. Refuses unrelated checkout changes. |
Failed · Sep 15Shared checkout had unrelated edits; callback also failed. Last recorded completion: Sep 2. |
| Weekly test sweep + simplify (agent-sdk)OpenHands/software-agent-sdk | Mon, 03:0030 minutes · profile gpt-5-6 |
Establishes green tests, conservatively removes redundant tests, simplifies code where justified, and reruns validation. May push a branch and open one ready-for-review PR; no-change runs produce no PR. | Failed · Sep 14Timed out after cloning and starting the conversation. No completed run in recorded history. |
| Weekly test sweep + simplify (odie)OpenHands/OpenHands | Mon, 05:0030 minutes · profile gpt-5-6 |
The same bounded cleanup procedure for Canvas: tests first, conservative removal, production simplification only when justified, then one PR with evidence. | Failed · Sep 14Remote-conversation environment error. Last recorded completion: Aug 12. |
Gazette's saved model field says deep-flash, but its inspected script makes no LLM call. Its macOS credential lookup and shared checkout currently tie it to this host. The next scheduled Gazette attempt is September 16 at 09:00; both weekly sweeps next fall on September 21.
OpenHands Cloud · 6 enabled
Cloud dispatches these definitions into run sandboxes. The account currently combines event-driven triage with scheduled notebook updates and TypeScript maintenance. There is no Cloud duplicate of the two local test-cleanup jobs in this inventory.
Cloud cron schedules are stored in UTC. Stockholm times shown here use September's CEST, UTC+2; they move one hour earlier locally when Stockholm returns to standard time.
| Job and target | When · budget | Purpose and permitted changes | Latest recorded result |
|---|---|---|---|
| Attention routerCurrent test: software-agent-sdk PRs only | 00:23, 08:23, 16:23 UTC02:23, 10:23, 18:23 Stockholm · 20 minutes · deep-pro |
Scores topic fit for Engel and files or updates Review Notebook notes. Current entrypoint limits the test to at most 10 SDK PRs. Source skips unchanged items. TEST_MODE is not a dry run: the source writes real notebook notes. | Completed · Sep 15Latest scheduled attempt at 16:23 UTC. Completion does not prove every note was correct. |
| Issue Duplicate Checker — detectOpenHands, software-agent-sdk, automation, extensions | GitHub issue opened25 minutes · deep-flash |
Investigates a newly opened issue for an existing duplicate. The source workflow can comment and mark a candidate for the later close sweep; it does not immediately close the issue. | Failed · Sep 13Only recorded attempt: timeout with sandbox unavailable during verification. |
| Issue Duplicate Checker — auto-close sweepDeployed repository scope unverified | Daily, 09:00 UTC11:00 Stockholm · 25 minutes · saved profile deep-pro |
Local source describes revisiting marked duplicates after a waiting period and checking replies before closing eligible issues. This is an intended issue-closing effect; deployed scope and safeguards were not verified. | Completed · Sep 15Latest scheduled attempt at 09:00 UTC. Closures were not independently checked. |
| QA Changes — Auto-Testsoftware-agent-sdk | PR opened or ready for review30 minutes · deep-pro |
Exercises changed behavior and appends an Auto-Test report to the PR description. The event filter excludes drafts and any case-insensitive Auto-Test text in the body. Its saved prompt claims a gate requiring more than three author commits merged to main; the deployed script gate was not verified. Report instructions prohibit new comments or reviews. |
Failed · Sep 13Only recorded attempt: command timed out or was killed. |
| Weekly upstream driftConfigured clone: enyst/openhands-agent | Mon, 06:00 UTC08:00 Stockholm · 30 minutes · deep-pro |
Reviews one bounded Python SDK release interval, ports changes tests-first, and advances the pin only with required evidence. Pushes a branch and opens or resumes one PR; never merges or pushes main. | Completed · Sep 14Reported successful no-op: no new upstream release. Next scheduled: Sep 21. |
| Weekly re-vendorConfigured clone: enyst/smolpaws; SDK source: enyst/openhands-agent | Wed, 06:00 UTC08:00 Stockholm · 30 minutes · deep-pro |
Re-vendors SDK main, reviews delegated server changes, regenerates the pinned Python OpenAPI, and opens or resumes one PR. Never merges or pushes main. | Not run yetFirst scheduled attempt: Wed, Sep 16 at 06:00 UTC / 08:00 Stockholm. |
“Completed” is the runner's result. It does not independently prove the right notebook notes, GitHub comments, labels, closures, or PR edits happened. A failed run may also have made partial changes. Cloud downloads of custom execution bundles returned server errors during this audit, so details drawn from their local source remain unverified against the deployed bytes.
The two weekly transpile jobs are separate stages. SDK drift owns the bounded upstream review; re-vendor consumes the SDK result and handles the server units. Wednesday reads the configured SDK main; it does not merge Monday's PR or wait for that PR to be accepted. Both resume existing work and use a draft PR with a handoff when the time budget expires. The SDK's GitHub Actions discovery watcher is a separate system and is excluded from this OpenHands Automations inventory.
Paused Cloud jobs and retired local work
| Paused job | Saved schedule | Historical behavior and write effects | Last recorded result |
|---|---|---|---|
| Roasted Code ReviewOpenHands/OpenHands and software-agent-sdk | Hourly, at :00 UTC30 minutes · saved profile gpt-5-6 |
Reviews up to five qualifying PRs on Engel's behalf, posts COMMENT reviews, and records them by pushing the public review log to documentation main. Its Opus 5 disclosure conflicts with the saved gpt-5-6 profile; identity wording needs review before reuse. |
PausedLast run completed Aug 1. Saved cadence does not currently trigger work. |
| Daily external PR security screenSaved targets: SDK, legacy agent-canvas name, OpenHands | Daily, 09:00 UTC11:00 Stockholm · 30 minutes | Inspects external contributors' PR diffs without executing their code. Historical policy posts ordinary review comments and sends one email if suspicious findings exist. The proposed lightweight replacement is not deployed. | PausedLast run completed Jul 2. Earlier planning records describe frequency and verbosity concerns. |
The local typescript-client Agent Server release maintainer was disabled and deleted on September 12, with no recorded runs. It is retired history, not an enabled job or an approved future plan.
Decisions ahead · proposals, not applied changes
The inventory records what exists. The following work still needs a decision or implementation; this audit changed no definitions, schedules, or run state.
- Repair the local execution path. Restore Canvas ingress and callback reachability, then investigate the weekly conversation errors. Decide whether the two failing cleanup schedules should pause during that work. Their eventual home—local or Cloud—remains undecided.
- Give Gazette its own checkout. A dedicated checkout would prevent ordinary documentation edits from blocking publication. Keep local execution while its host-specific credential lookup remains, or explicitly redesign that dependency before moving it.
- Make weekly repository ownership explicit. The live prompts use the real
enyst/*forks. The canonical repositories aresmolpaws/openhands-agentandsmolpaws/smolpaws. At this audit the SDK fork matched canonical main, but the server fork had diverged. The next re-vendor run would therefore start from a different baseline. A proposed policy is canonical main for source and PR base, with branches pushed to the enyst forks. That target policy has not been applied. - Decide whether to expand attention routing beyond SDK PRs. The September 15 update above replaces the pilot with weekly complete-context PR scoring. Both candidate lists still use the SDK repository; expansion to other repositories remains undecided. Topic fit alone continues to drive routing.
- Finish the security redesign before re-enabling. Beads
smolpaws-crm.9–.12propose a compact clean-pass stamp, terse request-changes only for real findings, stateful re-checks, and a quieter cadence. The paused definition still contains the older review-and-email workflow. - Recover deployed bundle evidence. Retrieve the custom execution bundles or a trustworthy source digest. Confirm the duplicate sweep's repository scope, waiting period, and reply veto, and verify the QA author gate before relying on those safeguards.
- Keep authored definitions distinct from registrations. The source-controlled “PR → Review Notebook note” event automation exists as code, but was not registered in either inspected deployment. The external reviewer/QA product ideas remain separately scoped work.
Evidence and keeping this useful
This snapshot combines read-only Cloud definitions and run history, local service records and current execution artifacts, repository source, and dated Beads plans. The local scheduler was observed polling even while ingress was down. No pending or running automation rows were present; that does not establish whether unrelated agent conversations were active.
For each later change, record where it is registered, what it targets, what it can write, trigger and timezone, enabled state, latest verified outcome, and the reason for the decision. Keep the observation date beside the claim. Beads owns work status; this page explains the operating arrangement and decisions.
- OpenHands Automations architecture — definition, queue, execution backend, callbacks, and watchdog.
- Moving an automation between deployments — code portability and deployment-specific registration.
- Maintaining the TypeScript transpiles, the SDK drift runbook, and the server re-vendor runbook.
- Review Notebook automation source — broader authored scope and topic-only routing policy.
- OpenHands Automation service and Extensions — runtime and catalog sources.