Archived QA snapshot · observed 19 September 2026, published here 20 September 2026. This preserves the original findings, not a statement about today's application. PR #17435 · Original QA report

OpenHands / OpenHands · PR #17435 · 19 September 2026

Real browser evidence, not a terminal screenshot

Application head 3e5e1ca; Chromium at 1440×1000 and 390×844. Built frontend, authenticated Agent Server/SDK 1.49.2, local automation 1.13.2, real openhands/claude-haiku-4-5-20251001 responses. No mocked API/LLM responses. 137 reviewed screenshots below.

This is not an all-green certification. Theme colors, typography, fields, borders, cards and most dialogs look consistent in exercised states. The audit also found the issues below. Automation execution and Cloud-only flows are explicitly not passed.

What does not look or work right

  1. Mobile conversation dialogs: Skills, Hooks and Tools are 640px wide at a 390px viewport (x = −125px), clipping content and close controls. Existing fixed sizing; PR changes w-[640px] to equivalent w-160, not the behavior.
  2. Mobile automation edit: a 940px-tall dialog centered in 844px yields y = −48px, clipping the title/close control and footer. Existing missing height/scroll constraint; changed lines only replace token classes.
  3. Mobile dashboard: action buttons squeeze the title/description to ~52px. Existing row layout.
  4. Mobile plugin search: input shrinks to ~40.5px next to category tabs. Existing row layout.
  5. Terminal history: a fresh real command appears correctly in the Terminal tab; reload shows an empty tab while chat retains the command. Relevant terminal/history implementation is not changed by this PR.
  6. Monaco lifecycle: opening an expanded Git diff and switching to Planner reproduces TextModel got disposed before DiffEditorWidget model got reset. Panel still renders afterward; two reproductions. Diff changes in this PR are styling-only; causal baseline execution not performed, so treat this as an observed issue, not a proven new regression.

Page-by-page coverage

All 15 changed route files were mapped and visited directly or through their parent screen; shared components add settings, panels and dialogs. A visit is not a pass for unavailable functionality. The table records tested states, not every permutation of every component.

Page / surfaceResultEvidence and limits
Home / conversations / workspace selectionPASSSecond workspace-backed conversation created via UI; real Python edit/assertion passed; both conversations visible. Workspace picker and manager checked. Screenshot
Files, Markdown, Git commits/diffsPASS + findingReal source and rendered Markdown; uncommitted changes and commit rows. Diff itself renders correctly; switching an expanded diff to Planner triggers a Monaco disposal exception. Screenshot
Task list, Planner, overview, usagePASSTask statuses from a real task_tracker call; populated planner file, overview and actual token/cost values. Cloud balance endpoint unsupported. Screenshot
Terminal / Browser panelsMIXEDFresh terminal output renders. Reload loses Terminal-tab history although commands remain in chat. Browser empty state checked; no live browser-agent run. Screenshot
MCPPASSCatalog, custom/install dialogs; installed Time MCP via UI and successfully tested reachability/tool listing. No external-account OAuth tested. Screenshot
SkillsPASSCatalog and tags, filters layout, detail and add dialogs; enabled-state rendering. Screenshot
Plugins / launchMIXEDCatalog, detail and add dialogs; launch-without-plugin error state. Mobile plugin search shrinks to about 40.5px. No third-party plugin launch executed. Screenshot
Apps / extension pagePASSInstalled, reviewed and enabled bundled dependency-free demo-page fixture; real extension mounted. Disabled/missing extension fallback also checked. Screenshot
Agent / LLM settingsPASSLists, existing/add profile editors, profile menus, provider-connection dialog. Saved real model profile used in actual conversations. Credential fields masked in evidence. Screenshot
Application / Condenser / Agent Context / VerificationPASSEvery page plus available Basic/All views; toggles, fields, help text and disabled/save states inspected. Not every configuration combination was mutated. Screenshot
SecretsPASSEmpty/add/populated/edit/delete-confirm states using an explicitly non-sensitive dummy value; no real service secret published. Screenshot
Backends / command menu / conversation dialogsMIXEDBackend chooser/manager, command menu, transcript export, skills/hooks/tools, cost and delete confirmation inspected. Three conversation dialogs clip on mobile. Screenshot
Automation dashboard / templates / Git SyncMIXEDReal local automation service and disabled QA record. Grid/list, filters/empty state, template setup, import/create-help dialogs, Git Sync disabled form. Dashboard header cramped on mobile. Screenshot
Automation detail / edit / activityPARTIALReal detail/configuration/empty activity log plus edit and delete-confirm dialogs. Edit modal clips vertically on mobile. No manual run was performed during this audit; populated activity/logs were not verified. Screenshot
Device authorizationPARTIALCode-entry UI inspected at both sizes; no genuine Cloud device authorization approved. Screenshot
Shared conversationBLOCKEDLocal Agent Server lacks /api/shared-conversations and /api/shared-events/search (404). Only error state exercised; not a sharing success pass. Screenshot

Remaining gaps

Screenshot gallery

Open an image for full resolution. Credential inputs are deliberately covered (magenta). Earlier captures may show missing emoji glyphs from the minimal headless image; standard emoji fonts were installed for the final PR images. System-message/Tools screenshots are omitted from public evidence; the UI was still inspected.

This report was generated by an AI agent (OpenHands) on behalf of Engel Nyst. Only reviewed screenshots were copied here; no browser storage, keys, transcripts, raw API responses or private state files were published.

Password fields — recaptured without screenshot masks

Addendum · 20 September 2026 · 8 new captures. All 137 original screenshots above are preserved unchanged.

The bright pink rectangles in eight original images were Playwright screenshot masks, not application styling. These new captures show the same four forms at desktop and mobile sizes, using the UI's native type="password" masking instead. The fields contain only the deliberately invalid dummy value visual-qa-not-a-real-credential; no real provider credentials were used.

The original PR application build (3e5e1ca; evidence-only HEAD f18fa0b) was reopened in Chromium against a fresh, real Agent Server 1.49.2. These are new browser captures, not edited versions of the old images. The empty conversation sidebar and demo profile reflect the fresh backend. No application CSS was changed, no screenshot overlay was applied, and no model calls or credential-validation checks were run.

In all eight captures, the credential input remained type="password", with a computed background of #202020 and border color of #404040. The password field is fully visible at both viewport sizes. This addendum verifies that field's appearance and concealment, not the validity of the dummy credentials or the unrelated automation flows.

Recaptured and published by an AI agent (OpenHands) on behalf of Engel Nyst.