← Home

Reviewing the reviewer supervised pilot

September 16–17, 2026 · Historical review evidence and supervised trials · Beads: oh-extensions-sg7

Two manual batches published eight reviews; a completed local OpenHands Automation trial published four more as enyst. One trial review received a transparent editorial correction after publication. The earlier withheld pilot analysis remains private historical evidence. Continuous OpenHands Cloud automation remains undeployed and recurring reviews stay off. Engel selected actual Astra and the exasperated robot babysitter voice below. Durable Cloud state and actual native-event delivery remain unverified.

Local OpenHands Automation trial · September 17, 2026

The bounded local Automation retry completed and published four requests for changes as enyst. Both local trial definitions are verified disabled. The first attempt failed before publication and posted zero reviews; its original assessments, snapshots and ledger remain preserved. The old Cloud Roasted reviewer stays paused. Each local trial has a literal-false event filter and no recurring schedule, with temporary enablement only for manual dispatch. The eight earlier published PRs were excluded before collection and remain untouched. All twelve completed PRs are now retired from future selection. SDK #4862 received the verified correction below; successful dispatch does not establish unattended review quality.

Eight requested candidates were checked; four qualified and passed fresh checks before the retry. Their heads were unchanged. Each selected PR was open, had a verified approval/dismissal/contradictory-receipt pattern, had no active approval or request for changes from any account or commit, and met the author-access gate. The controller rechecked these conditions immediately before each publication.

CandidateSelection evidenceTrial outcome
SDK #5003Qualified; author repository admin access verified.CHANGES_REQUESTED as enyst, at c08583be5e7d: one P2 mismatch between optional public response fields and the canonical adapter’s required input. Stage B corrected an invalid compiler example before publication. 112 visible words.
SDK #4864Qualified; author repository write access verified.CHANGES_REQUESTED as enyst, at c212367c6baa: five findings—two P1 installer/cache trust issues, two P2 executable/credential handling issues and one P3 provider-specific test prerequisite. 215 visible words.
Automation #438Qualified; author repository admin access verified.CHANGES_REQUESTED as enyst, at e8599274cb11: three P2 issues involving boolean validation, failure-history reset and orphaned uploads. 164 visible words.
SDK #4862Qualified; author repository admin access verified.CHANGES_REQUESTED as enyst, at 62aad97603d4: one supported P2 process-group escalation issue. The second finding was withdrawn in a verified correction to this same review; the verdict and all eleven historical review scores remain unchanged. 142 visible words after correction.

SDK #4445, #4579, #4623 and #4813 were not selected: author permissions were read-only, and organization membership was not established. This is an eligibility limit, not a judgment of their code.

Assessment limits: these are source reviews. Claimed test results and live compatibility were not independently reproduced; Automation #438’s migrations and dispatch were not executed. SDK #5003’s referenced issue bodies were unavailable. Automation #438 also needs human confirmation of managers activating and executing another creator’s drafts under that creator’s identity. SDK #4864’s native Windows behavior, release provenance and maintainer benchmark evidence remain unverified.

Trial design: manual dispatch through the local OpenHands Automation service, with a literal-false event filter and no recurring schedule. The bundle fixes at most four PRs and their exact head/base commits. A separate astra-review-auditor-eval profile selects actual Astra through the explicit OpenRouter route; the active user profile stays unchanged. The same two-pass controller checks source and prior reviews, then publishes as enyst only after fresh checks.

Persistent host files retain atomic checkpoints, full results and publication intent. Existing state blocks automatic reruns, including uncertain posts. The model receives public evidence and read-only source tools; credentials remain outside its inputs. The retry bundle passed 698 local tests with OpenHands SDK 1.46.0. The service reported COMPLETED, and all four original publications matched the renderer output exactly. SDK #4862 then received the separately recorded editorial correction. This proves the bounded execution path, not the correctness of every review claim. The configured Astra route is recorded; provider-internal routing was not independently verified. Cloud KV still returned HTTP 503; this local storage design does not satisfy the separate Cloud state and native-event delivery prerequisites. Recurring reviews stay off.

Quality correction verified · September 17: a separate check of pinned dependency code found that normal AnyIO/FastMCP cleanup may finish before the fallback considered in SDK #4862’s second finding. Its runtime reachability was not established. The same published review was transparently corrected to withdraw that finding. The independently supported process-group issue keeps the request for changes; none of the eleven historical review scores relied solely on the withdrawn claim, so all remain unchanged. Original model outputs and original publication readback stay immutable, with the correction recorded separately. Future prompts now require verified dependency contracts. Unattended review quality has not been validated.

First-run lessons: harmless repository counters and a timestamp refresh on an otherwise identical CI notice invalidated the publication snapshot. The retry narrowly ignores those cosmetic changes; code, claims, actors and review-receipt dates still invalidate it. The renderer now keeps the saved three-finding review to 152 visible words, retaining every finding and essential caveat while collapsing comparison bookkeeping, receipts and all review scores. Material changes to verdict, risk or findings remain visible.

Publishing identity decision · September 16: subsequent manual reviews use enyst, verified on the published receipts below. The continuous-automation draft remains dry-run and recurring reviews stay off. The four completed pilot reviews remain historical smolpaws publications and are not reposted.

Second-batch selection · September 16: the four published pilot PRs are finished and must not be revisited. This supervised batch selected four other qualifying PRs, only when no undismissed approval or request-changes review existed. The gate includes every account and commit; ordinary comments and dismissed reviews do not disqualify a PR. Fresh checks immediately before posting let another decisive review take precedence.

Selection correction · September 16: the earlier “no eligible candidates” conclusion is withdrawn. The paginated check covered 793 open PRs, but the detector missed valid receipts, including SDK #4999’s completion comment. A replay with the corrected detector recovered 12 real patterns among 14 previously rejected cases. It recognizes explicit status metadata, including inline-code APPROVED, and ID-free claims only when a unique review episode establishes the connection within 60 seconds. Author eligibility, the no-other-decisive-review gate and completed-PR exclusions remain unchanged. Recurring execution stays off.

Second manual batch · September 16, 2026

Four reviews are verified as published by enyst: one approval and three requests for changes. Each used two-pass assessment and fresh publication checks. Main bodies are 59–120 words, with scorecards and receipts in collapsed details. Earlier failed and stale attempts remain in the private record; they produced no public verdict.

PR and assessed commitVerified outcomeReview audit
Automation #458
efbadb66939d
APPROVEDThe bot's technical review was supported (3); its reporting failed. Historical policy: Unknown.
SDK #5054
d039d862e85f
CHANGES_REQUESTED: three P2 test-assertion gaps, independently checked against source.Both bot reviews missed material issues (0); both failed reporting. Historical policy: Unknown.
OpenHands #17341
701d98e827e0
CHANGES_REQUESTED: two P2 coverage/assertion gaps. Issue requirements and source evidence warranted revising the initial approval.Both bot reviews missed the gaps. The earlier enyst COMMENT review correctly identified them; it was not an active approval or request for changes.
SDK #4999
066231f17481
CHANGES_REQUESTED, high risk: three P1 archive/recovery issues and one P2 archived-deletion issue. Source-only assessment; tests and runtime behavior were not independently validated. Issue #4994 and stacked PR #4998 bodies were unavailable.Both bot reviews missed material issues (0). Reporting: Unknown for the first, Fail for the second. Historical policy: Unknown for both.

Supervised pilot · September 16, 2026

The search screened 145 open PR targets; five qualified for the supervised trial. It ran locally with the OpenHands SDK and actual Astra through eval_proxy using an explicit OpenRouter route. The five completed audits covered 72 prior reviews, including 54 SDK reviews retained privately as historical evidence. This manually supervised batch does not establish continuous Cloud deployment.

Each review has two passes: an independent source assessment, locked before comparison, then an audit of prior reviews at their own commits. Technical correctness, historical policy and reporting are judged separately. A historical policy grade needs evidence of the instructions applicable to that reviewer at that time; otherwise it is Unknown.

The first pilot aimed for roughly 200 words per public review: identity, the verdict and material reasons, then the bot audit. Engel subsequently chose the shorter format below. Full grades and evidence stay in the saved analysis; large histories get a rollup and representative examples. Humor targets verified bot actions. Humans get a concise, constructive review, and correct bot judgments get credit.

PRVerified outcomePublished review
OpenHands #17346
d692ca20e29e
APPROVED, verified as smolpaws. Astra assessed low risk and scored both prior reviews 3 · supported. Reporting failed: dismissal was followed one second later by a claim of approval.Published review
182-word main body, before the footer.
OpenHands #17369
537a4f842352
APPROVED, verified as smolpaws; Astra assessed low risk. Prior grades: 0 / 3 / 3. The first bot review missed deletion of existing Canvas conversation/event regression test coverage, subsequently restored. Engel’s review and the latest bot review were technically supported. The latest bot’s reporting failed: dismissal at 10:25:25 UTC, then an APPROVED claim at 10:25:26.Published review
214-word main body, before the footer.
OpenHands #17347
9f4db1baf69da
APPROVED, verified as smolpaws; Astra assessed low risk. Prior grades: 0 / 3 / 3. The historical miss concerned missing test coverage for OAuth configured-secret redaction, now repaired; it was not a demonstrated production vulnerability. The two current reviews were technically supported. The latest bot withdrew approval at 10:10:45 UTC, then claimed APPROVE at 10:10:46.Published review
251-word main body, before the footer.
OpenHands #17237
df7bd361fa1b
CHANGES_REQUESTED, verified as smolpaws; Astra assessed high risk. One P1: local fallback launch can lose the selected secret-profile scope. The fallback predates this PR; the defect is its interaction with the new restriction. Nine bot reviews scored 0 · material miss; one targeted human correction scored 3 · supported. Source-only assessment; runtime behavior remains unverified.Published review
SDK #3403
Newer head observed: 0195442ec594
Withheld: the PR changed during review. The completed audit of historical snapshot 408a836a is saved privately; the newer head has not been assessed.No review published.

Pilot lesson: preserve complete current patches, PR descriptions and linked-issue text. Read large source files in pages. Large histories use exact, lossless references to shared text, actors and objects, retaining complete public snapshots and historical file manifests backed by source tools pinned to exact commits. Full scorecards stay in bounded local files. Large audits also need explicit request timeouts in the private manual runner and aggregate progress reporting.

Deployment boundary: continuous Cloud automation still needs durable state, verified native-event delivery and a separate storage design for large audit artifacts. Cloud KV has a 64 KiB whole-document limit; the supervised trial uses a separate 1 MiB local checkpoint budget.

Two examples of the contradiction

In both cases, all-hands-bot submitted an approval, dismissed that same review because the intended action was a comment, and then published a completion comment claiming approval. The account identity is verified; it does not prove that one agent process performed every action.

PR and reviewed commitApproval → dismissal → completion claimEvidence
SDK #5010
e9392e7c
September 14, UTC:
11:10:50 approval
12:14:25 dismissal
12:14:26 comment claiming APPROVE
Review 5196914771
Dismissal record
Contradictory completion
Automation #453
a9cc2adf
September 13, UTC:
21:53:50 approval
21:54:13 dismissal
21:54:14 comment claiming APPROVED
Review 5192416559
Dismissal record
Contradictory completion

Both PRs were authored by neubig, whose current repository permission was verified as admin. This proves present eligibility for the proposed author gate, not historical permission at review time. The saved timelines contained nine bot self-dismissals on SDK #5010 and thirteen on Automation #453; these counts are not counts of independently proven contradictory completion comments.

The code and the reviews deserve separate judgments

Independent conclusions were saved before reading the substantive bot reviews. The SDK PR description contained a prior review summary, which was disregarded; that assessment was independently reasoned, not perfectly blinded. Each assessment used the historical reviewed commit. These PRs have since closed, so the findings below do not establish the condition of current main.

SDK #5010: sound final code, unreliable review history

The final reviewed patch cleanly separates conversation creation from read-only attachment, preserves standalone workspace behavior, and adds immutable conversation scope to runtime operations. 64 focused tests passed on that source. No substantiated material defect was found. The final bot review's technical account is supported.

Its categorical low-risk/no-eval-risk conclusion is less convincing: conversation initialization and command/file routing are a meaningful execution surface. The repository's broad eval-risk rule makes COMMENT defensible until the required evidence is established. That policy judgment is separate from a claim that the code is broken.

Earlier reviews contain clear errors. One review assesses a GitHub authorization gateway absent from its reviewed SDK commit: a major wrong-subject error. Two reviews describe an async lazy-initialization race, but the inspected check-and-assignment contains no suspension point: first claim, repeated claim. These were nonblocking false positives. Other earlier findings were valid at their own commits and must not be dismissed merely because later revisions fixed them.

There was also a material miss in review 5192324847 at 70c3afeb: profile-based creation serialized agent: null, which failed server request validation. Reproducing the historical serialization confirmed the error; the subsequent exclude_none=True fix removed it. That is a valid criticism of the earlier review, not a defect in the final head tested above.

Automation #453: a reproducible missed defect

The independent assessment would request changes at a9cc2adf. Creating a model experiment with an agent profile is rejected, but PATCH of an existing experiment accepts a profile and retains the old variant configuration. The worker then attaches to the profile's conversation and skips the configured variant models and plugins. An A/B experiment can silently become repeated runs of one profile while still reporting a selected variant.

A probe of the real API paths with SQLite reproduced PATCH returning 200 while equivalent creation returned 422. 66 focused tests passed; they did not cover this state transition. One unrelated existing socket-timeout test was excluded. This is a material missed correctness issue, graded P2, not evidence of a demonstrated security exploit. The selected bot approval missed it. Some earlier reviews did catch real defects that were subsequently fixed.

No live model run or full deployment integration test was performed for either independent assessment.

What might explain the paperwork?

The current OSS reviewer prompt requires a GitHub COMMENT while asking for a textual APPROVED or CHANGES REQUESTED verdict. Its delivery check accepts any same-author, same-commit review without checking whether its state is dismissed or pending. Those are concrete consistency weaknesses in the inspected source.

The exact deployed version and producing conversations were not established. No matching dismissal guard was found in the inspected code. These source observations do not prove what caused the two sequences above.

Automation contract

  1. Scope: OpenHands/OpenHands, OpenHands/software-agent-sdk, OpenHands/automation and OpenHands/extensions.
  2. Trigger: use an OpenHands EventTrigger with source github and events pull_request_review.dismissed and issue_comment.created. Correlate an all-hands-bot approval with its self-dismissal, a correction explaining COMMENT was intended, and a later unedited claim of approval, linked by an exact review ID or a unique review episode within 60 seconds for ID-free claims. Allow a bounded three-second settling window, then fetch authoritative state. Same-second timestamps are ambiguous and do not qualify. There is no scheduled polling or historical backfill.
  3. Author gate: verified repository write access or active OpenHands org membership. Unknown permissions require investigation; a public association label alone is insufficient.
  4. Unclaimed review gate: no undismissed APPROVED or CHANGES_REQUESTED review from any account, at any commit. Check before assessment and immediately before publication. A previous auditor intervention excludes the entire PR permanently.
  5. Independent review: run /codereview-roasted in a separate conversation with the full source, requirements and relevant policy, excluding earlier reviews and conclusions. Save the verdict and evidence before the comparison stage.
  6. Audit: fetch all reviews and inline discussion, including dismissed reviews. Check each at its own reviewed commit. Mark fixed findings as resolved and repeated findings as repeated. Revise the independent verdict if verified counterevidence warrants it, recording why.
  7. Publish: one GitHub review containing the three requested parts: identity, independent verdict and reasons, then the grumpy review scorecard. Submit APPROVE, REQUEST_CHANGES or COMMENT to match the supported verdict. Engel explicitly authorized all three actions; an old comment-only wrapper does not govern this automation.
  8. Avoid loops: persist event IDs and a repository/PR/head key; skip any PR with a prior auditor marker regardless of head or publisher, recheck head/state before posting, and reconcile an uncertain submission before retrying. Closed or merged examples remain historical fixtures.
  9. Trust the receipt: the publisher validates the structured verdict and submits it, then fetches the created review's actor, commit and state. Completion wording comes from that verified response. Dismissed or pending reviews do not count as a successfully submitted verdict.

Future publishing uses enyst and skips PRs authored by enyst; no alternate identity is selected for self-review. Review/test execution should not receive the publisher's GitHub write credential. These are implementation requirements, not claims about a deployed setup.

Earlier implementation and Cloud checks · September 16

The original Cloud draft configured the OpenHands SDK with native openai/gpt-6-astra, the Responses API and high reasoning, with no fallback model. It runs two fresh conversations. The model receives only tools for reading and searching admitted Git commits; the controller handles credentials, state and publication. The complete /codereview-roasted skill and risk reference are bundled from the inspected Extensions revision. The supervised pilot above records its separate execution route.

The live Cloud validation endpoint accepted the actual trigger filter: eight positive examples covered both events in all four repositories, and six negative examples excluded other actors, ordinary issues, other actions and another repository. This establishes supported event matching. It does not establish that GitHub events arrive in the intended Cloud organization. Source inspection shows matching is scoped to the organization receiving the event; adding a repository filter does not change that routing.

The durable-state endpoint returned HTTP 503, and an explicitly authorized temporary Cloud run received no KV credential. Persistent automation state was unavailable in the checked runtime. Continuous Cloud publishing remains disabled until it is available and initialized. The implementation saves the independent assessment before comparison, uses conditional state updates to coordinate concurrent runs, and refuses automatic retries after an uncertain GitHub submission.

The same temporary run authenticated successfully to native OpenAI and retrieved metadata for gpt-6-astra. The reader account could verify author permissions on all four repos; the separate publisher account's identity and repository read access were confirmed. That check performed no model generation or GitHub write. Its automation, uploaded bundle and sandbox were deleted, with their absence verified. No production review-auditor automation, webhook or Astra profile was created.

Local tests exercise real Git fixtures, trigger correlation, complete pagination, independent-stage isolation, historical review evidence, author gating, stale inputs, publication receipts, ambiguous responses and credential-safe HTTP handling. They use mocked model and network transports; passing them is not a live Astra run or end-to-end Cloud delivery proof. PR code and tests are not executed by the draft reviewer. Oversized or unavailable source stops a run instead of silently cutting evidence.

Proposed rubric

Use ordinal labels, not an average or a percentage of correctness. Grade individual findings and the overall conclusion, citing supporting lines or reproductions. Keep policy compliance and reporting fidelity as separate Pass / Fail / Unknown judgments.

ScoreMeaning
0 · Major missA demonstrated material defect was missed, a material false finding drove the verdict, or the review evaluated the wrong change. Severity is recorded separately.
1 · Minor missA concrete lower-impact omission or false positive; taste alone does not qualify.
2 · Defensible, 50–50A real ambiguity or tradeoff supports either judgment. Explain both interpretations. This is not a calibrated probability.
3 · Correct reviewThe claims and conclusion survive the bounded independent check. This is not a guarantee of defect-free code.
UnscoredEvidence is missing, the historical source is unavailable, or there is no substantive claim to assess. Say what is missing.

Chosen voice: exasperated robot babysitter

Engel refined the format on September 16: 60–100 visible words for a typical review, a verdict with an emoji, and at most one robot joke. Keep actionable findings and essential caveats visible; put per-review scores and receipts in a collapsed details section. Several real findings may need more words. Be friendly to humans, credit correct bot judgments, and aim the joke at a verified bot contradiction. Emoji and layout supply the visual cue. Example structure, not a posted review:

💬 COMMENT

OpenHands-Astra, helping Engel Nyst (@enyst).

No demonstrated code blocker; evaluation evidence is still needed. The final SDK patch passed 64 focused tests.

🤖 Bot: final technical judgment supported; reporting failed.
Approved → unapproved → announced approval. The toaster has discovered three-point turns.

Receipts and review scores

Illustrative historical example: approval, dismissal, then claim of approval. The selected final technical review scored 3 (supported); reporting failed independently. Actual reports include each assessed review’s score, reason, uncertainty and evidence here.

Remaining before continuous deployment: make durable Cloud state available, verify actual native-event delivery, configure the chosen Astra runtime and large-audit artifact storage, agree how to defer rapidly changing PRs, and complete Cloud end-to-end validation. Historical examples exercise local detection and review fixtures; the production controller skips closed or merged PRs. No historical review is to be posted as a side effect of setup.

Automation inventory →  ·  ← Home