Current state: The Cloud automation is registered but disabled. We removed its temporary test endpoints and keys. The job has no active schedule or required pull-request check.
Why check from outside?
The testing strategy describes the larger aim. Code-writing agents can change repository tests. A trusted checker must hold the active rules and assign the result outside the candidate.
A candidate is the Agent Server build under test. A contract is the list of required behaviors. The verifier sends requests to the candidate and checks the replies.
The verifier uses public REST and WebSocket interfaces. It does not load candidate tests, plugins, helpers, or expected results. The verifier assigns each result from its own rules.
What exists
| Part | Current state |
|---|---|
| Contract bundle | Three required scenarios cover session replay, restart with saved storage, and a limited legacy event check. The live event-order rule is still a proposal. |
| External verifier | A deterministic program sends REST and WebSocket requests. A separate scripted provider returns a fixed tool call. No paid model call is needed. |
| Cloud automation | OpenHands Cloud can run the verifier bundle. The definition is disabled and has an inert staging schedule. Manual dispatch needs a separately prepared candidate. |
| Evidence | Each required scenario gets pass, fail, or blocked. The saved JSON contains the candidate revision, artifact hash, bundle and configuration hashes, received frames, and REST observations. |
Repository CI checks the verifier's test cases and package. It does not test a live Agent Server candidate.
What the manual Cloud run found
On 9 October, the verifier tested SDK commit 9f47d471. A local controller started that build in a restricted process sandbox. This manual setup did not prove that an unattended service can isolate and identify each candidate.
The first Cloud attempt found a DNS fault in that sandbox. We corrected the test setup and ran the same verifier bundle again. The second run reached the scripted provider and recorded a request witness.
| Scenario | Observed behavior | Result |
|---|---|---|
WS-REPLAY-001 | Saved history and replay matched. The live stream sent sequence 11 when the rule expected 10. | Fail against proposed order rule |
WS-RESTART-001 | The controller killed the server. Saved history and replay matched after restart. The live stream sent 13 when the rule expected 12. | Fail against proposed order rule |
LEGACY-EVENTS-001 | The legacy event endpoint returned all 14 saved events. | Pass for its limited scope |
The verifier also saw a socket error that did not enter durable history. The restored history matched before the post-restart live-order failure.
The durable live-order rule is still a proposal. An independent owner must decide whether it is a public contract. The run records a real difference from that proposed rule.
See the dated evaluation and the exact confirmed-run evidence. The first run remains there as a setup diagnostic.
Why it is disabled
Cloud runs the verifier. It does not build a pinned SDK candidate for this job. The manual candidate, provider route, and restart control used temporary services. Those services are closed.
The current Cloud definition has no active test keys. It cannot repeat this test without a new candidate, new controls, and new configuration.
The personal Cloud account can change its own automation bundle. An agent with that account key could change the test it faces. A separate owner must approve the active bundle and publish accepted results.
A run status is not a contract result. All required scenarios must pass in the evidence. A failed observation is a fail. A missing candidate, key, or trace is blocked.
Path to regular checks
-
Approve the rules
Decide the live-order guarantee first. Approve the exact verifier bundle digest and the effective configuration digest. Keep activation outside the implementation agent's control.
-
Build a known candidate
Choose one exact main or release commit for each run. Build its artifact and record the hash. Independently prove that the test endpoint serves that artifact. Isolate the candidate from verifier and publisher keys.
-
Provide test controls
Host the scripted provider where the candidate can reach it. Protect the provider's request witness. Add a trusted restart control that kills the process and keeps its storage. Issue new limited keys for each run.
-
Save and judge the evidence
Run the approved verifier in Cloud. Export the exact evidence and its hash before the sandbox ends. Compare the executed bundle and configuration with approved digests. Check the independently signed artifact binding, provider witness, and restart record. Treat missing proof as blocked.
-
Publish and repeat
Use a separate identity to publish the result for the exact tested commit. Then start a regular main or release check. The proposed first cadence is daily at 09:00 in Stockholm. Retain failing storage until investigation ends. Then remove test keys and candidate resources.
A required pull-request check can follow after the regular path is reliable. A new PR head or base needs a new test result.
What to test next
The present bundle focuses on event transport. Later rules can cover run ownership, cancellation, confirmation, secret delivery, streaming, storage faults, and client compatibility.
Each new rule needs a clear scope. A known good case must pass. A deliberate wrong case must fail. The approved bundle must change through the separate review path.
Sources
- Testing architecture when agents write the code — the larger strategy.
- Verifier design and Cloud setup — the merged implementation guide.
- Required scenarios — the proposed contract bundle.
- First Cloud evaluation — results, exact evidence, and limits.
- OpenHands Automations — how the service starts and records a job.