Archived QA snapshot · observed 19 September 2026, published here 20 September 2026. This preserves the original findings, not a statement about today's application. PR #17435 · Original QA report
OpenHands / OpenHands · PR #17435 · 19 September 2026
Real browser evidence, not a terminal screenshot
Application head 3e5e1ca; Chromium at 1440×1000 and 390×844. Built frontend, authenticated Agent Server/SDK 1.49.2, local automation 1.13.2, real openhands/claude-haiku-4-5-20251001 responses. No mocked API/LLM responses. 137 reviewed screenshots below.
What does not look or work right
- Mobile conversation dialogs: Skills, Hooks and Tools are 640px wide at a 390px viewport (x = −125px), clipping content and close controls. Existing fixed sizing; PR changes
w-[640px]to equivalentw-160, not the behavior. - Mobile automation edit: a 940px-tall dialog centered in 844px yields y = −48px, clipping the title/close control and footer. Existing missing height/scroll constraint; changed lines only replace token classes.
- Mobile dashboard: action buttons squeeze the title/description to ~52px. Existing row layout.
- Mobile plugin search: input shrinks to ~40.5px next to category tabs. Existing row layout.
- Terminal history: a fresh real command appears correctly in the Terminal tab; reload shows an empty tab while chat retains the command. Relevant terminal/history implementation is not changed by this PR.
- Monaco lifecycle: opening an expanded Git diff and switching to Planner reproduces
TextModel got disposed before DiffEditorWidget model got reset. Panel still renders afterward; two reproductions. Diff changes in this PR are styling-only; causal baseline execution not performed, so treat this as an observed issue, not a proven new regression.
Page-by-page coverage
All 15 changed route files were mapped and visited directly or through their parent screen; shared components add settings, panels and dialogs. A visit is not a pass for unavailable functionality. The table records tested states, not every permutation of every component.
| Page / surface | Result | Evidence and limits |
|---|---|---|
| Home / conversations / workspace selection | PASS | Second workspace-backed conversation created via UI; real Python edit/assertion passed; both conversations visible. Workspace picker and manager checked. Screenshot |
| Files, Markdown, Git commits/diffs | PASS + finding | Real source and rendered Markdown; uncommitted changes and commit rows. Diff itself renders correctly; switching an expanded diff to Planner triggers a Monaco disposal exception. Screenshot |
| Task list, Planner, overview, usage | PASS | Task statuses from a real task_tracker call; populated planner file, overview and actual token/cost values. Cloud balance endpoint unsupported. Screenshot |
| Terminal / Browser panels | MIXED | Fresh terminal output renders. Reload loses Terminal-tab history although commands remain in chat. Browser empty state checked; no live browser-agent run. Screenshot |
| MCP | PASS | Catalog, custom/install dialogs; installed Time MCP via UI and successfully tested reachability/tool listing. No external-account OAuth tested. Screenshot |
| Skills | PASS | Catalog and tags, filters layout, detail and add dialogs; enabled-state rendering. Screenshot |
| Plugins / launch | MIXED | Catalog, detail and add dialogs; launch-without-plugin error state. Mobile plugin search shrinks to about 40.5px. No third-party plugin launch executed. Screenshot |
| Apps / extension page | PASS | Installed, reviewed and enabled bundled dependency-free demo-page fixture; real extension mounted. Disabled/missing extension fallback also checked. Screenshot |
| Agent / LLM settings | PASS | Lists, existing/add profile editors, profile menus, provider-connection dialog. Saved real model profile used in actual conversations. Credential fields masked in evidence. Screenshot |
| Application / Condenser / Agent Context / Verification | PASS | Every page plus available Basic/All views; toggles, fields, help text and disabled/save states inspected. Not every configuration combination was mutated. Screenshot |
| Secrets | PASS | Empty/add/populated/edit/delete-confirm states using an explicitly non-sensitive dummy value; no real service secret published. Screenshot |
| Backends / command menu / conversation dialogs | MIXED | Backend chooser/manager, command menu, transcript export, skills/hooks/tools, cost and delete confirmation inspected. Three conversation dialogs clip on mobile. Screenshot |
| Automation dashboard / templates / Git Sync | MIXED | Real local automation service and disabled QA record. Grid/list, filters/empty state, template setup, import/create-help dialogs, Git Sync disabled form. Dashboard header cramped on mobile. Screenshot |
| Automation detail / edit / activity | PARTIAL | Real detail/configuration/empty activity log plus edit and delete-confirm dialogs. Edit modal clips vertically on mobile. No manual run was performed during this audit; populated activity/logs were not verified. Screenshot |
| Device authorization | PARTIAL | Code-entry UI inspected at both sizes; no genuine Cloud device authorization approved. Screenshot |
| Shared conversation | BLOCKED | Local Agent Server lacks /api/shared-conversations and /api/shared-events/search (404). Only error state exercised; not a sharing success pass. Screenshot |
Remaining gaps
- Manual automation execution, populated run/activity logs, errors/cancellation and output views: not exercised during this audit; the QA record was left disabled.
- Cloud shared-conversation success, device authorization completion, subscriptions/balance, external OAuth and Git Sync to a real remote.
- Browser agent output, every third-party plugin/automation template, every theme/configuration, and native Electron windows. This was Chromium UI QA, not native or cross-browser certification.
Screenshot gallery
Open an image for full resolution. Credential inputs are deliberately covered (magenta). Earlier captures may show missing emoji glyphs from the minimal headless image; standard emoji fonts were installed for the final PR images. System-message/Tools screenshots are omitted from public evidence; the UI was still inspected.









































































































































This report was generated by an AI agent (OpenHands) on behalf of Engel Nyst. Only reviewed screenshots were copied here; no browser storage, keys, transcripts, raw API responses or private state files were published.
Password fields — recaptured without screenshot masks
Addendum · 20 September 2026 · 8 new captures. All 137 original screenshots above are preserved unchanged.
The bright pink rectangles in eight original images were Playwright screenshot masks, not application styling. These new captures show the same four forms at desktop and mobile sizes, using the UI's native type="password" masking instead. The fields contain only the deliberately invalid dummy value visual-qa-not-a-real-credential; no real provider credentials were used.
The original PR application build (3e5e1ca; evidence-only HEAD f18fa0b) was reopened in Chromium against a fresh, real Agent Server 1.49.2. These are new browser captures, not edited versions of the old images. The empty conversation sidebar and demo profile reflect the fresh backend. No application CSS was changed, no screenshot overlay was applied, and no model calls or credential-validation checks were run.
In all eight captures, the credential input remained type="password", with a computed background of #202020 and border color of #404040. The password field is fully visible at both viewport sizes. This addendum verifies that field's appearance and concealment, not the validity of the dummy credentials or the unrelated automation flows.
Original masked capture
Original masked capture
Original masked capture
Original masked capture
Original masked capture
Original masked capture
Original masked capture
Original masked capture
Recaptured and published by an AI agent (OpenHands) on behalf of Engel Nyst.