The original 68 findings, followed by new reports from later feature-map runs. Grouped by repository and theme, with the status of the associated fixing PRs on the right.
The fixes include both community and agent-authored PRs. Existing contributor work was reviewed rather than duplicated.
Report snapshot . GitHub issue and PR metadata refreshed . Screenshots, recordings and reported test results remain the original evidence; no new browser or API verification was performed.
issues in this list
120
associated PRs merged
61
associated PRs open
63
open PRs with failing checks
15
PR totals count distinct PRs, including identified alternatives. An issue can have several PRs, and a merged partial fix can leave an issue open.
68First 68
21Follow-up runs & fixes
31Agent Server API · draft map
The follow-ups include Canvas runs and bugs found while verifying earlier fixes. The 31 Agent Server API reports come from a separate feature map; they are not added to Canvas’s 743 mapped behaviors.
MergedOpen PRDraft PRClosed unmergedChecks failingNo fix PR found
Follow-up to #17914: the text was restored, but Send stayed disabled. The selected draft also covers suggestion and Git-action prefills; alternative #18226 focuses on draft restoration.
The before evidence is a container log; the after screenshot is the original capture from the locally tested PR branch. The current PR state is shown in the ledger.
PR #18219 is the selected fix and includes the original real-server before/after captures; #18216 closed without merging and #18220 is an alternative proposal.
PR #5690 locks metadata updates, but installation still swaps the directory outside that lock; concurrent listing can prune the entry and lose its enabled state.
PR #5694 rejects new empty MCP configurations, but leaves persisted-settings migration and isolation of valid servers from an invalid sibling unresolved.
The merged fix notes remaining gaps for in-flight sync cycles and environment-only encryption-key changes.
No matching issues
Try another term or clear the filters.
The findings, up close
What the issues looked like.
One short walkthrough for each issue. Original before/after pairs are selected from the PRs where available; single issue captures, recordings and terminal evidence are labeled. API reproduction cards condense the linked reports; they are not new test runs. Open and closed-unmerged proposals show what was tested on that branch, not a shipped fix.
Opening an unknown conversation id also shows a misleading "data this UI does not understand" toast
Opening a missing conversation produced two errors, including an incorrect warning about incompatible server data. The merged fix recognizes the missing conversation and shows only the relevant message.
First 68
BeforeA missing conversation triggers two different error messages.AfterOnly the accurate missing-conversation message remains.
Conversation header "⋯" menu overflows the screen at phone width and does not close on Escape
The conversation menu ran off the edge of a phone screen and ignored Escape. The merged fix keeps every item on screen and returns focus when Escape closes the menu.
First 68
BeforeAt 390px, the menu extends past the right edge and clips its labels.AfterAt the same width, every menu item fits within the viewport.
"Branch from here" on a user message drops the message but leaves the new conversation's composer empty
“Branch from here” removed the chosen user message but left the new conversation’s composer empty. The merged fix restores that message as a draft ready to edit.
First 68
BeforeThe new branch opens with an empty composer.AfterThe new branch keeps “Reply with the single word: pong” in the composer.
Stop/Delete conversation confirmation buttons render without their data-testid (BrandButton expects testId)
The Stop and Delete confirmation buttons passed the wrong prop name, so their intended test IDs never reached the DOM. PR #18151 restores those selectors.
First 68
This is a nonvisual DOM-attribute fix. PR #18151 reports five passing component tests; its image attachment returned 404 when the original evidence was collected. Alternative PR #18140 supplied no screenshots. Earlier PR #17983 closed without merging.
"Delete all" conversations deletes only the loaded first page and states a smaller count
“Delete all” only removed the first loaded page: twenty of twenty-three conversations in the reproduction. Its confirmation also understated the total; the issue remains open with no confirmed fix PR.
A rejected (Cancel) action leaves its event group spinning as "k/N actions completed" after the conversation is idle
Cancelling an action could leave its group spinning after the conversation was idle. PR #17995 adds a Rejected marker and aims to count rejected actions as resolved.
First 68
BeforeThe rejected action still appears to be running, with “1/2 actions completed.”After · PR branchThe row gains a Rejected marker; the captured header still says “1/2 actions completed.”
Both images are the PR’s original captures. The after image shows the new Rejected label but still contains the old progress count/spinner, so it does not visually establish the settled-group result claimed in the PR text. The run used a scripted mock LLM with the real app and Agent Server.
Composer "More input actions" menu closes as soon as it opens at narrow widths, so the model cannot be switched
At narrow phone widths, “More input actions” closed immediately after opening, blocking model selection. The merged fix keeps the overflow menu open and usable by pointer or keyboard.
First 68
BeforeAt 320 px wide, the menu disappeared after the click.AfterThe same overflow menu stays open.
Conversation draft typed within 500 ms of navigating away is lost; the older draft comes back
Typing and leaving a conversation within 500 ms could restore an older draft and lose the latest text. PR #17968 flushes the pending draft before navigation while keeping conversations’ drafts separate.
First 68
EvidenceOriginal issue capture: the old “QA draft text” returns instead of the newly typed draft.
The selected PR provides before/after recordings rather than a usable pair of still screenshots. This is the original issue’s reproduction capture.
A failed conversation launch from Home shows duplicated error toasts and clears the typed prompt
A failed Home launch cleared the prompt and showed duplicate error toasts. The merged community fix retains the prompt and, for a refused launch, presents one readable error.
First 68
BeforeRejected launch: two raw error toasts and the typed prompt is gone.AfterOn the PR branch: one readable error and the prompt is preserved.
Conversation drawer at phone width: agent and file-link tab opens do nothing, Planner hides the tab row, and agent-chosen tabs are not persisted
At phone width, clicking a file link or an agent’s request to open a tab could do nothing. These actions now open the correct drawer panel; the fix also keeps Planner tabs visible and remembers the selected tab.
First 68
BeforeOn a phone, clicking src/calc.py leaves the conversation on screen.AfterThe same file link opens the Files panel and the requested file.
Slash command feedback: bare /btw is silently discarded, and superseded /goal rows keep offering Resume
Submitting /btw without a question silently cleared the composer, and outdated goal rows kept offering Resume. The merged fix explains the missing question and limits Resume to the latest eligible goal row.
First 68
BeforeRecorded CLI output before: /btw clears the composer with no toast or reply.AfterApp capture after: a toast asks the user to provide a question.
The PR’s before image is a rendered capture of real CLI output; the after image is an application screenshot of the same bare-command case.
Sidebar conversation list: misleading download error, and the default layout matches no View preset
A failed conversation download produced two toasts, blamed an unknown ID, and left its menu open. The fix reports the failure once and closes the menu; it also makes a fresh sidebar match the Recent preset.
First 68
BeforeFailed download: two toasts and the menu remains open.AfterFailed download: one error and the menu closes.
Composer "More input actions" Model submenu runs off the right and bottom edges at narrow widths (third profile and LLM Profiles link can't be clicked or tapped)
At narrow widths, the composer’s Model submenu placed profiles and its settings link beyond the screen edge. PR #18115 keeps click- and keyboard-opened submenus inside the viewport.
First 68
BeforeThe Model submenu extends beyond the right and bottom edges.After · PR branchThe same submenu fits inside the narrow viewport.
Send stays disabled after branching from a user message with prefilled text
Branching correctly prefilled the composer but left Send disabled until the user typed. PR #18203 updates Send for restored text and suggestion prompts; the sweep also found the same symptom in Git-action prefills.
Follow-up runs & fixes
BeforeThe original before close-up shows restored branch text with a disabled Send button.After · PR branchProposed fix: restored branch text immediately enables Send.
Original screenshots are preserved at their supplied sizes: the before image is a composer close-up, and the after image is a full page. The PR also verifies suggestion prefills; its Git-menu path was reproduced before the fix but not yet live-tested afterward.
Settings switches are skipped by Tab and invisible to screen readers (SettingsSwitch renders <input hidden>)
Settings switches disappeared from keyboard navigation and the accessibility tree. PR #17963 restores their native checkbox behavior and makes keyboard focus visible.
First 68
After · PR branchProposed fix: a visible focus ring surrounds Sound Notifications.
This after screenshot and the recordings come from draft #17963. Alternative draft #18154 is also open but supplies no screenshots; no before still is available for this comparison.
Automation delete confirmation ignores Escape and is not exposed as a dialog
The automation delete confirmation ignored Escape and lacked dialog semantics. The merged fix makes it an accessible modal and lets Escape dismiss it without deleting the automation.
First 68
BeforeAfter pressing Escape, the delete confirmation remains open.AfterAfter pressing Escape, the dialog closes and the automation remains.
Screen-reader semantics: nested <main> in Settings, no aria-current on Customize, <html lang> never updated
Screen readers received duplicate Settings landmarks, no current-page marker for Customize, and the wrong page language. The merged fix supplies one main landmark and synchronizes navigation, language, and direction semantics.
First 68
BeforeOriginal probe capture: two main landmarks, missing aria-current, and English language metadata.AfterOriginal probe capture: one main landmark, current-page markers, and the selected language and direction.
These are the PR’s rendered captures of browser accessibility/DOM probes, not recordings of screen-reader speech.
Confirmation Continue/Cancel shortcuts are ⌘-only; Ctrl+Enter does nothing on Linux and Windows
Continue and Cancel accepted only Command-key shortcuts, leaving Linux and Windows users without working equivalents. Open PR #18127 accepts Ctrl, shows platform-specific hints, and prevents repeat responses after an action has been submitted.
First 68
EvidenceOriginal issue capture: after Ctrl+Enter on Linux, the action is still waiting and the hints show Command.
Original issue still, captured before the fix. Open PR #18127 provides before/after recordings from a real Linux stack, with a capture annotation showing each keypress; it supplies no static screenshot pair. macOS behavior is covered by unit tests only.
LLM profile row menu is not keyboard-operable: focus stays on the trigger, and arrow keys stop at a disabled item
Opening an LLM profile menu by keyboard left focus on its trigger, and arrow navigation could get stuck at a disabled action. Focus now enters the menu and moves across available actions.
First 68
BeforeKeyboard-opened menu: the focus ring stays on the ⋮ trigger.AfterKeyboard-opened menu: focus moves to Edit.
Icon-only buttons in the composer and conversation drawer have no accessible name
Composer and diff-viewer icon buttons lacked accessible names, so assistive technology could not identify their actions. The fix names those controls; the captured accessibility output makes the change visible.
First 68
BeforeAccessibility output before: unnamed buttons are highlighted in red.AfterAccessibility output after: Stop, Send, Remove, and diff controls have names.
These original PR images render the browser’s accessibility snapshots. The visual button appearance is largely unchanged; the accessible names are the evidence.
Skill detail modal ignores Escape while its '+N' tag popover trigger has focus
Escape behaved incorrectly when the skill detail’s +N tag control had focus. PR #17967 gives the popover its own dismissal step, keeping the detail dialog open until the next Escape.
First 68
BeforeThe PR’s before capture shows the tag overflow still open.AfterThe PR’s after capture shows the popover dismissed and focus back on +2.
Original PR screenshots from a local mock run. These captures remain the original branch evidence.
Agent profile row menu is not keyboard-operable: focus stays on the trigger, and arrow keys stop at the disabled "Set as active"
Opening an agent profile’s menu from the keyboard left focus on its trigger, and arrow navigation got stuck at a disabled item. PR #18112 focuses the first enabled action and returns focus to the trigger when the menu closes.
First 68
BeforeThe menu is open, but the focus ring remains on the three-dot trigger.AfterFocus enters the menu on Edit, ready for keyboard navigation.
Escape does not close the composer "+" menu, the status-dot menu or the drawer tabs menu
Escape left the composer +, status-dot and drawer tabs menus open. The merged fix closes the topmost menu and restores focus without dismissing a layer underneath.
Follow-up runs & fixes
BeforeAfter Escape, the composer’s + menu remains open.AfterAfter Escape, the menu closes and focus returns to +.
Escape does not close the composer "More input actions" overflow menu at narrow widths
At narrow widths, Escape did not close More input actions. The merged fix registers that overflow menu with the shared Escape stack and returns focus to its trigger.
Follow-up runs & fixes
BeforeAt 320px, the More input actions menu remains open after Escape.AfterAt 320px, Escape closes the menu and returns focus to its trigger.
These are the original stills linked from the merged PR. Its newer before/after recordings demonstrate the same Escape behavior.
The Canvas App trust dialog left keyboard focus behind it, allowing Enter to trigger a background Update. PR #18197 moves focus into the confirmation, traps Tab and restores focus on close.
Follow-up runs & fixes
Before · recordingBefore video: Space → Tab → Enter activates background Update while the trust dialog remains open. Full recordingAfter · PR branch · recordingAfter video: the same keys focus and confirm Enable trusted app inside the dialog. Full recording
Original before/after WebM recordings from isolated real Canvas and Agent Server stacks using native keyboard input, without network mocks.
Provider connection menu does not focus its first action
Opening a provider connection’s menu left focus on the trigger, so ArrowDown did nothing. PR #18205 focuses Bulk add after the menu’s portal mounts.
Follow-up runs & fixes
BeforeThe menu opens while keyboard focus remains on the trigger.After · PR branchProposed fix: after opening the menu, ArrowDown reaches Edit with a visible focus ring.
The visible after image is one ArrowDown after opening. The PR’s live activeElement checks establish the initial focus on Bulk add; this capture makes the resulting keyboard navigation visible.
Git user name and email from Settings > Application are never applied to agent commits
Saving a Git name and email did not pass that identity to agent conversations. On a clean host, commits failed with “Author identity unknown”; the issue remains open with no confirmed fix PR.
First 68
EvidenceOriginal issue evidence: the agent cannot commit because Git has no author identity.
Original screenshot from the issue; no confirmed fix PR or after capture was found.
Model Router: choosing "Don't create profiles" snaps back to a provider connection, and Save creates 11 LLM profiles
“Don’t create profiles” immediately reverted to a provider connection, and saving created eleven unwanted profiles. The merged PR preserves the explicit choice; the verified flow creates none.
First 68
BeforeAfter choosing “Don’t create profiles,” the picker has reverted to QA_deepseek.AfterThe picker preserves “Don’t create profiles.”
Model Router list heading and router names are invisible in light color themes
The Model Router list used white text on a white background in light themes. The merged fix makes the list heading and router names readable with the selected theme.
First 68
BeforeIn Light+, the list heading and router names disappear into the background.AfterThe heading and the qa-router names are readable in the same theme.
A failed Save on Settings > Application discards the user's unsaved edits
A failed Application Settings save reset the form and discarded the user’s edits. PR #17945 keeps the inputs mounted and preserves changes so Save can be retried.
First 68
EvidenceOriginal issue capture: the save fails and Sound Notifications reverts to off.
Original issue reproduction capture. PR #17945 and alternative draft #18201 are both open; neither supplies a relevant before/after screenshot pair.
Add LLM Profile keeps the name derived from the initial model after the user picks a different model
A new profile kept its original model’s suggested name after a different model was selected. The name now follows the model until the user enters a custom name.
First 68
BeforeThe chosen model is deepseek-chat, but the suggested name still says gpt-5.6-sol.AfterThe suggested profile name follows deepseek-chat.
Settings save retries a 422 validation error three times before showing it
One settings Save sent the same invalid request three times before showing its validation error. The issue asks for permanent validation failures to surface immediately.
First 68
EvidenceRecorded request log: one Save sends three PATCH requests, each rejected with 422.
This is the issue’s rendered request log, recorded by Playwright—not a browser DevTools screenshot. No associated fix PR was found in the checked links.
Error toasts show raw "HTTP request failed (4xx): {json}" text for /model and profile rename
Profile errors exposed transport status and raw JSON instead of explaining the problem. The merged fix extracts the server’s useful message for model selection, rename, and validation failures.
First 68
BeforeUnknown profile: the toast exposes HTTP status and raw JSON.AfterThe toast says simply “Profile ‘qa-nope’ not found”.
LLM and agent profile editors silently disable Save for a duplicate name, with no message
Entering an existing LLM or agent profile name silently disabled Save with no explanation. PR #18144 shows a duplicate-name message and marks the field invalid while keeping Save disabled.
First 68
BeforeThe duplicate profile name disables Save without explaining why.AfterOn the PR branch, a visible message explains that the name is already taken.
Original PR captures from a production-built local stack with a real Agent Server. The current message reuses the existing “meta-profile” wording.
Add LLM Profile shows the agent's current LLM options (such as temperature) but saves the new profile without them
Add LLM Profile displayed current options such as temperature but silently omitted untouched values when saving. PR #17965 preserves the displayed settings without silently inheriting credentials or provider links.
First 68
EvidenceOriginal issue capture: reopening the newly created profile reveals an empty Temperature field.
The PR has no before/after screenshots; this is the original issue’s reproduction capture.
MCP library lists entries that cannot be installed on a local backend; their install dialog is empty and Install does nothing
The local MCP library offered servers it could not install. The Sentry dialog had no configuration fields and Install did nothing; the issue remains open with no confirmed fix PR.
Skills "Use skill" opens Home with an empty composer instead of pre-filling the skill command
“Use skill” opened Home but left its composer empty. The merged fix carries the chosen command into Home and focuses the composer; Create Automation uses the same repaired path.
First 68
BeforeAfter “Use skill”, Home had an empty composer.AfterThe selected /docker command is ready in Home.
"Show Available Hooks" always says "No hooks configured" on local backends, even when workspace hooks are active
The Available Hooks dialog said “No hooks configured” even when workspace hooks were active. It now loads and displays the workspace hooks, and Refresh fetches changes.
First 68
BeforeThe dialog showed no hooks for a workspace that had one.AfterThe workspace’s Pre Tool Use hook is listed.
Installing a catalog STDIO MCP server twice leaves the second copy without its catalog identity
Installing a catalog STDIO MCP server twice stripped the second copy’s description and badge. Repeated installs now retain their catalog identity, including in search.
First 68
BeforeSecond copies time_1 and memory_1 lose their catalog descriptions and badges.AfterBoth copies retain their catalog descriptions and badges.
A failed STDIO MCP connection test says "Check the URL and server type" and hides the real error
A failed STDIO server test told users to check a URL, even though these servers run commands. The fix shows command-specific guidance and the underlying error.
First 68
BeforeA failed command incorrectly gets “Check the URL and server type.”AfterThe card identifies a server-command failure and shows the command error.
Skill and MCP card toggles can swallow the first click, and keep the new state after a failed save
Skill and MCP toggles could lose the first click, while failed skill saves left an unsaved state looking enabled. PR #17970 stabilizes the click target and rolls failed changes back to the saved state.
First 68
BeforeBefore: after the click and failed-save scenarios, both skills remain shown as enabled.After · PR branchOn the PR branch: the successful disable is retained and the failed enable is rolled back.
Plugins and Apps: install/update errors say "Disconnected (check URL or network)" instead of the server's reason, and plugin errors toast twice
A rejected plugin install was reported twice as a disconnection, hiding the server’s useful explanation. The merged fix shows the server’s reason once and gives clearer guidance for duplicate app installs.
First 68
BeforeOne failed install produces two misleading “Disconnected” toasts.AfterThe same invalid source produces one specific error message.
Plugins and Apps UI: "Launch Launch Plugin" title, "skill" wording on plugin controls, app switch usable while busy, docs point to /extensions
Plugin controls called plugins “skills,” the launch title repeated “Launch,” and busy app switches still responded to the keyboard. The merged fix corrects the labels, disables busy switches, and updates the Apps documentation.
First 68
BeforeThe launch dialog says “Launch Launch Plugin” and “I trust this skill.”AfterThe same dialog says “Launch Plugin” and “I trust this plugin.”
Custom MCP server editor reports a malformed header as an environment-variable error
A malformed MCP authentication header produced an environment-variable error. PR #17966 names the actual field: “Headers must follow KEY=value format.”
First 68
BeforeThe invalid header is incorrectly described as an environment-variable problem.AfterThe same input now receives a header-specific validation message.
Available Hooks dialog shows an empty Commands box for prompt and agent hooks
Prompt and agent hooks showed empty Commands boxes. The merged fix carries their text into the dialog; a remaining type-selection error is tracked separately in #18206.
Follow-up runs & fixes
BeforePrompt-style hook rows contain empty Commands boxes.AfterThe same rows display the hook instructions.
Duplicate plugin install shows the API hint "Use force=true to overwrite." instead of pointing to Update or Uninstall
Installing the same plugin twice showed an API-only force=true hint. The merged fix tells users to use Update or Uninstall; another proposal remains open.
Follow-up runs & fixes
BeforeDuplicate installation tells the user to “Use force=true.”AfterThe message points to the available Update or Uninstall controls.
Slack MCP credential probe reports any failed tool call (e.g. "fetch failed") as "Credential check failed"
A failed Slack probe request was labeled a credential failure even when the network call failed. PR #18166 distinguishes connection failures from rejected credentials.
Follow-up runs & fixes
EvidenceOriginal failure: the Slack dialog calls “fetch failed” a credential problem.
The PR reuses the original failure screenshot; it does not provide an after capture.
A failed MCP server save, toggle or delete shows two identical error toasts
A failed MCP save, toggle or delete produced two identical error toasts. PR #18182 gives each action one error message while preserving feedback from the conversation panel.
Follow-up runs & fixes
BeforeDeleting an already-removed MCP server raises two identical errors.AfterProposed fix: the same failed delete raises one error.
Automations Filters popover closes after every option click, so filters cannot be combined
Choosing one automation filter closed the whole popover, making combinations awkward. The merged fix keeps it open so the next filter can be selected immediately.
First 68
BeforeAfter choosing Disabled, the Filters popover has closed.AfterAfter the same choice, the popover stays open for another filter.
Automation "View logs" shows "(no output)" for runs that created a conversation (local backend)
Completed local automation runs showed “(no output)” even when they had produced logs. The merged fix reads the run’s server-level output, so the Logs dialog displays it.
First 68
BeforeA successful run has an empty Output tab.AfterThe same workflow now displays the run’s actual stdout.
Automation detail and edit: schedule edit silently dropped, wrong run label, slow not-found page, tarball saved as .tar
Automation editing and detail views had four rough edges: a silently discarded schedule edit, a misleading run label, a slow not-found page, and a gzip download named .tar. The issue records reproduction steps for each.
First 68
No screenshot or linked fix PR was available in the issue at the status check.
Git Sync page shows credentials embedded in the repository URL, and its Enable help text is outdated
Git Sync displayed credentials embedded in a repository URL, and its help text described an obsolete startup flag. The merged fix removes credentials from the displayed URL and explains when syncing is enabled.
First 68
BeforeThe original test capture exposes the deliberately fake qa-user:qa-pass credentials.AfterThe same repository label now omits the credentials.
The credentials in this original screenshot are explicit QA fixtures on example.invalid.
Automation detail page never shows Repositories, Plugins or Notification for automations that have them
Automation details and exports dropped repositories and plugins because Canvas read fields the service did not return. The merged fix reads the stored preset metadata and displays the configured sources.
First 68
BeforeThe automation’s repositories and plugins are missing.AfterThe same automation now shows both repositories and both plugins.
The Notification part remains tracked separately in #16644 / PR #16673; this merged PR covers repositories and plugins. Export/import still supports only the first repository.
Automation templates: a search with no matches shows a blank page, and cards say "1 MCPs to connect" for optional integrations
Searching automation templates for a missing term left a blank page, and cards miscounted optional integrations. PR #18153 adds a no-results state with a clear-search action and counts only required integrations with correct singular and plural wording.
First 68
EvidenceOriginal issue capture: a no-match search leaves the template area blank.
The PR reuses this original issue capture and reports automated tests; it does not provide an after-fix screenshot.
Automation setup offers an Agent profile for Prompt and Plugin actions, but choosing one blocks creation with "Extra inputs are not permitted"
Prompt and Plugin automations offered an agent profile that their API could not accept, so choosing one blocked creation. The merged fix offers profiles only for supported Upload/Bundle actions and omits them from preset requests.
This is the PR’s original combined before/after image, captured with mock data. The PR explicitly reports that a live Agent Server/automation backend was unavailable.
Automation setup: Agent profile field goes blank after switching the action through Prompt/Plugin or choosing "Default"
Switching an automation through Prompt/Plugin, or choosing Default, could leave its Agent profile field blank. PR #18111 normalizes the empty value so “Default” is shown and checked.
First 68
BeforeAfter changing actions, the field is blank and no option is checked.AfterThe same flow now displays Default and marks it selected.
Original before/after screenshots from PR #18111. Alternative PR #18187 reuses the original issue screenshot.
Automations: a deleted automation stays on the dashboard, home page and detail page after Run now fails with "Automation not found"
An automation deleted elsewhere stayed visible even after Run now reported it missing. PR #18073 refreshes the list and detail data after that 404, removing the stale dashboard card and updating the count.
First 68
BeforeAfter the failed run, PR Triage Digest remains and the count is seven.After · PR branchThe deleted card disappears and the count drops to six.
Original before/after screenshots from open PR #18073 use Canvas mock mode. The author visually verified the dashboard; home and detail views were not driven separately. Alternative PR #18186 is also open and reuses the original issue screenshot.
Automation run Logs dialog overflows its panel with a long task summary; tabs and output unreachable at phone width
A long automation task summary pushed log tabs and output beyond the dialog, especially on phones. PR #18175 makes the capped panel scroll so the output remains reachable.
Follow-up runs & fixes
BeforeA long task summary spills beyond the phone dialog, leaving the output unreachable.AfterProposed fix: the panel scrolls to the Output and Error tabs and the run log.
Home shows dead automation link and manifest errors when automation interface is absent
Without an admitted automation interface, Home still linked to a missing automation page and displayed two internal manifest errors. PR #18212 hides those surfaces and stops their background queries.
Follow-up runs & fixes
BeforeWith the automation interface deliberately absent, Home shows a dead Schedule a task link and two manifest errors.After · PR branchProposed fix in the same injected state: the link and manifest errors are absent.
The missing interface was created by isolated fault injection in a copied dependency package. The normal pinned package provides it; both captures use the real app and the same healthy Local backend without model calls.
Files tab keeps showing stale file content after Refresh (and even after a reload)
The Files tab kept showing old content even after Refresh or a page reload. The merged fix revalidates text requests and refreshes preview URLs so the visible file matches the current file on disk.
First 68
BeforeAfter an edit and Refresh, the HTML preview is still stuck on “QA page v1.”AfterAfter a further edit and Refresh, the preview shows the current “QA page v3 AGAIN.”
Workspace picker: folder names are invisible at phone width, and the empty state never shows
At phone width, the workspace browser hid every folder name; the empty workspace list also showed no explanation. The merged fix gives names room and restores the “No workspaces yet” state.
First 68
BeforeAt 390px, every entry reads only “Folder.”AfterThe same phone layout now shows distinct folder names.
First-run onboarding and the API-key screen show a "No backend is configured." error toast
First-time setup showed “No backend is configured” while asking the user to add a backend. The merged fix waits until a backend exists before loading its model metadata.
First 68
BeforeFirst-run setup incorrectly raises a missing-backend error.AfterThe same setup step opens without the misleading error.
Launcher never creates TMUX_TMPDIR, so local Agent Canvas instances share one tmux server and reset each other's terminals
Starting or stopping one local Canvas instance could reset another instance’s terminals: both fell back to the same tmux server. Separate launcher socket directories remain to be fixed; the SDK prerequisite for long socket paths has merged.
First 68
The issue documents terminal output and has no attached screenshot or direct launcher fix PR. SDK #5559 merged the long socket-path prerequisite; the launcher issue remains open.
Launcher: --allow-lan-session-key is ignored, the port-in-use error prints twice with a stack trace, and backend-only runtime_services lists a frontend
The launcher ignored its LAN-session-key flag, duplicated a port-in-use failure, and described a frontend service in backend-only mode. Replacement PR #18156 is open; the earlier draft closed without merging.
First 68
Current PR #18156 reports 156 passing launcher tests. Its image attachment returns 404, so there is no available screenshot evidence. Earlier draft #17964 closed without merging.
Switching backend from a Manage backends row leaves a blank conversation page instead of redirecting
Switching backends from Manage backends left the old conversation URL open and the main pane blank. The merged fix redirects away from detail pages to the newly selected backend.
First 68
BeforeQA_Second is selected, but the conversation pane is blank.AfterSelecting QA_Second now opens its home page.
Backends and Cloud sign-in: Escape closes the whole Manage backends stack, raw 404 toast on shared pages, stray about:blank pop-ups
Backend and sign-in flows had three problems: Escape dismissed stacked dialogs together, missing shared conversations produced a duplicate error toast, and failed cloud sign-in left blank tabs. This original capture shows the technical error alongside “Conversation not found.”
First 68
EvidenceOriginal issue capture: the not-found screen also exposes a raw HTTP/JSON error.
Original issue capture. A fresh reproduction still found the duplicate toast, with shorter wording; no linked fix PR was found at this check.
Recovery gate (stored backend answers 503) shows two generic "An error occurred" toasts from the free-models hydrator
A returning user whose saved backend was down saw two generic error toasts on the recovery screen. PR #18164 suppresses background-query toasts where the page already explains the connection failure.
Follow-up runs & fixes
EvidenceOriginal issue evidence: the recovery screen explains Disconnected while an extra generic error toast appears.
Original issue screenshot only. The open PR’s image is a first-run contrast case, not an after capture of this recovery failure.
Docker image exits on hosts without IPv6: static server binds "::" and fails with EAFNOSUPPORT
The Docker image exited on hosts without IPv6 because its frontend bound only to ::. PR #18180 selects an IPv4 bind address when the kernel has no IPv6 support.
Follow-up runs & fixes
AfterProposed fix: the container serves the onboarding page on a host without IPv6.
There is no before page to capture: the original container never served the port and exited. The PR links the EAFNOSUPPORT log; this image is from the patched container.
Edit backend dialog clips its controls at phone width
Edit backend kept a 520px width on phones and clipped its controls. PR #18219 caps the dialog to the viewport and scrolls short-screen forms while keeping the title and close control visible.
Follow-up runs & fixes
BeforeAt 390px, Edit backend extends past both edges and clips its controls.AfterProposed fix: the complete dialog, Close, Cancel and Save fit on the phone.
Command menu misses Model Router and Agent Context, and always shows ⌘K
The command menu could not find Model Router or Agent Context, and showed the Mac shortcut on every platform. The merged fix adds the missing destinations and shows Ctrl+K on Linux and Windows.
First 68
BeforeSearching for Model Router returns “No commands found.”AfterThe same search finds Model Router and shows the Ctrl+K shortcut.
UI copy errors: raw schema text and reStructuredText, untranslated labels, a literal \n, wrong labels and a dead example link
Settings and customization screens exposed schema jargon, wrong labels, untranslated text, and a broken example link. The merged cleanup replaces these with readable UI copy; the condenser settings illustrate the change.
First 68
BeforeCondenser settings show internal names and raw schema descriptions.AfterThe same settings use readable names and explanations.
Generic error toasts still show raw "HTTP request failed (<status>): {json}" text (settings save, MCP delete, shared links)
Generic error toasts exposed HTTP transport text and raw JSON. The merged fix extracts the server’s message and uses readable fallbacks; duplicate MCP toasts are tracked separately in #18181.
Follow-up runs & fixes
BeforeSaving invalid settings exposes HTTP status text and raw JSON.AfterThe same failure now shows the server’s readable message.
Keep sibling Settings and Customize pages reachable on tablets
Tablet layouts hid both the secondary sidebar and phone Back bar, making sibling Settings and Customize pages hard to reach. PR #18200 adds a compact section selector.
Follow-up runs & fixes
BeforeAt tablet width, Application settings offers no direct links to sibling settings pages.AfterProposed fix: a compact Settings selector exposes the sibling pages.
VITE_DO_NOT_TRACK=1 still initializes PostHog and loads remote scripts from the telemetry host
Do Not Track still let the browser attempt requests to the analytics host. Related fix #18086 has merged and stops PostHog from initializing under Do Not Track; this original report and an alternative draft remain open.
First 68
BeforeWith Do Not Track enabled, the original main build attempts three analytics requests.AfterThe same check on the merged fix attempts no analytics requests.
Original before/after captures from merged PR #18086. The verification script added the request-count overlay and blocked the recorded analytics requests; the overlay is not part of Canvas. The original issue #17898 remains open.
condenser_kind is exposed as a user-facing settings field with a reStructuredText developer description
The settings schema exposed an internal condenser selector with developer jargon and raw values. An earlier proposal was closed; a new draft now proposes removing the selector metadata.
First 68
BeforeBefore: the SDK and API publish the internal selector and its developer wording.Proposed changeProposed change: the selector is absent from the exported schema; validation still works.
Original rendered terminal captures from PR #5531, not UI screenshots. That proposal closed without merging, so the after image is historical proposal evidence. New draft #5584 has no screenshots.
Rotating a provider connection's API key does not reach the active LLM settings; new conversations keep using the old key
Changing a provider connection’s API key leaves a stale key in the active settings, so new conversations can keep using the old credential. Two existing proposals address this through the related provider-connection report.
First 68
The related proposals offer read-side resolution or write-side updates. Neither PR includes a before-and-after image pair; their current states are shown in the ledger.
After deleting the active LLM profile, its provider connection cannot be deleted ("referenced by the active agent settings")
Deleting the active LLM profile leaves its provider connection referenced by hidden settings. The connection then refuses deletion even though the UI shows no profiles using it.
First 68
No linked fix PR or before-and-after PR evidence found. The related key-rotation proposals do not cover this deletion guard.
Concurrent plugin requests corrupt .installed.json, so a just-installed plugin is recorded as source "local" and Update/launch then fail
Concurrent plugin requests can corrupt the installation record. A freshly installed plugin is then treated as a local source, so updating or launching it fails.
First 68
PR #5690 is now identified; no before-and-after visual evidence was retained in the original collection. The issue remains open.
POST /api/mcp/test classifies a remote MCP server's HTTP 401 as error_kind "unknown", so Agent Canvas tells a user with a wrong GitHub token "Connection failed"
An invalid GitHub MCP token is reported as a generic connection failure. The server returns no structured credential-error signal for Canvas to use, so the user gets the wrong troubleshooting message.
Follow-up runs & fixes
EvidenceObserved: a rejected dummy token is presented as a connection failure.
Original screenshot from the October 8 feature-map pass. The token was intentionally invalid; no configuration was saved. No fixing PR or after image is available.
Canvas extension manifest drops pages[].nav_label, so Agent Canvas shows the page title in the rail
The server drops a Canvas extension’s navigation label, so the sidebar shows the full page title instead. PR #5521 preserves the label through the manifest and API.
First 68
BeforeBefore: the rail uses the page title, “Hello from an extension”.After · PR branchDraft fix: the rail shows the configured label, “Extension demo”.
Original screenshots from draft PR #5521, using a real Canvas instance and the demo-page extension. All three proposed fixes remain draft; none has merged.
download-trajectory returns 500 when a state file is replaced while the zip is built
Downloading a trajectory while conversation state was being saved could fail with HTTP 500 or include unfinished, unredacted files. The merged fix skips temporary files and archives the same bytes it checks for secrets.
Follow-up runs & fixes
Action
GET /api/file/download-trajectory/{id}
while repeatedly changing confirmation_policy
Original PR reports; different run totals. Output-archive race errors remained in both runs.
The PR provides real-server test output, with no screenshots. Its fix addresses state-save races; the separate output-archive race is tracked by #5539.
download-trajectory fails or returns a broken zip when the same trajectory is downloaded again or concurrently
Repeated downloads reuse the same temporary ZIP, so requests can delete or overwrite one another’s archive. A draft fix gives each download its own file and cleans it up after streaming.
Follow-up runs & fixes
Action
Repeated and concurrent GET /api/file/download-trajectory/{id}
Before
same_path=True
second_valid=False (deleted by first response cleanup)
After · PR branch
5 sequential + 24 concurrent live-server downloads returned valid ZIPs; no temporary ZIPs left.
Original draft PR evidence. The fix has not merged.
The draft PR records a live Uvicorn test and lifecycle output; it contains no before-and-after screenshots.
After a cipher key change, settings/secrets/profile writes answer 200 and null stored keys, not 409
After an encryption-key change, settings and secret writes can permanently replace stored keys with null while reporting success. Rejecting the write would preserve the original values for recovery with the correct key.
Agent Server API · draft map
Action
After restarting with a different OH_SECRET_KEY: PUT /api/settings/secrets, PATCH /api/settings, or POST /api/profiles/{name}/activate
Observed
Writes return 200 and overwrite stored keys with null. Restoring the original key no longer recovers them.
Expected
Refuse writes when stored values cannot be decrypted; preserve the files and active profile.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
PUT /api/settings/secrets with '' or '**********' answers 200 and nulls the secret instead of 422
Saving an empty secret or submitting its masked placeholder reports success but destroys the value. The secret remains listed even though reading it returns 404.
Agent Server API · draft map
Action
PUT /api/settings/secrets
{"name":"MY_TOKEN","value":"**********","description":"edited"}
Observed
200; GET /api/settings/secrets/MY_TOKEN then returns 404, while the name remains in the list.
Expected
Return 422, or preserve the existing secret when updating only its description.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
PATCH provider connection with blank or "**********" api_key returns 200 and replaces the key, not 422
A provider-connection edit accepts blank or masked API keys and breaks every profile linked to that connection. The API rejects null but silently clears a blank key or stores the mask as the real key.
200 with api_key_set: false; activating a linked profile produces api_key: null. A masked value is stored literally.
Expected
Return 422 or keep the existing key; linked profiles continue to use it.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Profile validation ignores the saved provider connection linked by a draft profile. It reports an authentication failure for a configuration that saves and activates successfully with the connection’s key.
Agent Server API · draft map
Action
POST /api/profiles/draft/validate with llm.provider_connection_id pointing to a working saved connection.
Observed
200 with valid: false and LLMAuthenticationError; the same model with the key inline validates successfully.
Expected
200 {"valid":true,"error":null}, using the connection’s key and base URL.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
GET omits secrets added at create or via POST .../secrets until the next save; restart can drop them
Secrets added to a conversation are available in memory but are missing from the saved state until another action triggers a save. A restart or idle eviction can silently discard the last secret added.
Agent Server API · draft map
Action
POST /api/conversations/{id}/secrets
{"secrets":{"QA_LAST":"dummy"}}
GET /api/conversations/{id}
Observed
POST returns 200, but GET omits the secret; a restart before another save loses it.
Expected
The new secret appears immediately in GET and survives a restart.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Pause during a model call not saved: GET says running (expected paused); after SIGKILL it is error
Pausing during a model call updates the live conversation but does not immediately save the pause. Polling clients still see running, and a crash in that window turns the conversation into an error after restart.
Agent Server API · draft map
Action
During an in-flight model call: POST /api/conversations/{id}/pause
GET /api/conversations/{id}
Observed
GET still says running. After SIGKILL and restart it says error, although a PauseEvent exists.
Expected
After the pause returns 200, GET and saved state say paused; a crash and restart preserve that state.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
/goal loop active at a graceful restart stays `running` afterwards instead of `interrupted`
A graceful restart cancels an active goal loop without marking it interrupted. The conversation is paused afterward, but clients keep showing a running goal that the stop endpoint cannot clear.
Agent Server API · draft map
Action
Start POST /api/conversations/{id}/goal; restart gracefully while it is running; inspect goal events.
Observed
The last goal status remains active: true, status: running. POST /goal/stop returns 200 without changing it.
Expected
The last goal status is active: false, status: interrupted, retaining the objective and iteration.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
A second server's first load of a shared conversation reverts meta.json to its stale startup copy
When a second server takes over a shared conversation, it can overwrite current metadata with the stale copy it read at startup. The live reproduction loses a newer title and timestamp during failover.
Agent Server API · draft map
Action
Start server B on shared storage; rename on server A; stop A; load the conversation on B.
Observed
B rewrites meta.json with the startup title, before B, and the older timestamp.
Expected
The takeover preserves the latest title, after B, and updated_at.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Resumed conversation keeps a since-disabled plugin's skills and commands; only its hooks are dropped
Disabling a plugin and restarting removes its hooks from a resumed conversation but leaves its skills and commands active. The resumed conversation behaves differently from a new conversation created after the disable.
Agent Server API · draft map
Action
Disable a loaded plugin, restart, resume the conversation, then send /<plugin>:<command>.
Observed
The command still activates and its skills remain in agent_context.skills; only its hooks are removed.
Expected
The disabled plugin contributes no skills, commands or hooks to the resumed conversation.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Sub-agent file naming a missing skill makes the first message return 500 instead of being skipped
One sub-agent configuration that names a missing skill can make every message in the conversation fail with 500. A globally installed bad file can affect new conversations in every workspace.
Agent Server API · draft map
Action
Load a sub-agent file whose skills list contains a missing name; POST /api/conversations/{id}/events.
Observed
500 on the first and subsequent messages; the sub-agent listing still reports the configuration without a warning.
Expected
Skip the invalid sub-agent with a clear warning, or return a descriptive 4xx.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
POST /api/settings/mcp/{key} with {} returns 201 and new conversations lose all MCP tools; expected 422
Saving an empty MCP server configuration succeeds but can remove every MCP tool from new conversations, including tools from valid neighboring entries. The only explanation appears in a server log.
Agent Server API · draft map
Action
POST /api/settings/mcp/qa-empty
{}
Observed
201 stores {"enabled":true}; a new conversation gets no MCP tools from any configured server.
Expected
422 and no stored entry; valid MCP servers continue to contribute their tools.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Skill/Canvas install with an unknown ref returns 200 when the git source is cached (expected 400)
Installing a skill or Canvas app at a nonexistent Git ref fails correctly on a fresh source but succeeds after the source is cached. The API installs the cached revision and records the missing ref as though it had been requested successfully.
Agent Server API · draft map
Action
After caching a Git source: POST /api/skills/install with ref: no-such-ref and force: true.
Observed
200 installs the cached commit and stores requested_ref: no-such-ref; Canvas installs behave the same way.
Expected
400 with the current installation unchanged, whether or not the source is cached.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
events/search and /count with kind=MessageEvent return nothing; only the full class path matches
Event search accepts the documented short event name but returns an empty history. Only the internal fully qualified class name matches, so clients following the API documentation silently miss events.
Agent Server API · draft map
Action
GET /api/conversations/{id}/events/search?kind=MessageEvent
GET /api/conversations/{id}/events/count?kind=MessageEvent
Observed
200 with items: [] and count 0, even though the conversation contains MessageEvents.
Expected
Return the existing MessageEvent items and their count, matching the full class-path query.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Deleting or sandbox-pausing a conversation leaves its open WebSockets silent instead of closing them
Deleting a conversation or preparing its sandbox for pause leaves existing event sockets open but silent. Clients receive neither a close signal nor further events, so they cannot recover normally.
Agent Server API · draft map
Action
Open a conversation event socket, then DELETE /api/conversations/{id} or POST /api/conversations/prepare-for-sandbox-pause.
Observed
Existing sockets receive no close frame and no further events; after pause, only newly opened sockets receive messages.
Expected
Close the subscribed sockets when the operation completes so clients learn of deletion or reconnect.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Waiting on a locked profile or connection store freezes the whole server; /alive stalls up to 30 s
Waiting for a locked profile or provider-connection store freezes the server event loop. An unrelated health check can stall for nearly 30 seconds even though the server is otherwise healthy.
Agent Server API · draft map
Action
Hold a profile or connection store lock; send a request using that store and concurrently GET /alive.
Observed
The health request waits for the locked request; recorded delays range from 10.5 to 28.1 seconds.
Expected
GET /alive returns 200 within 2 seconds while the other request waits for its lock.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Failed run blocks the event loop while logging a rich traceback: /alive takes ~5 s, not ~100 ms
Formatting a failed model call as a rich console traceback blocks the event loop for several seconds. During that logging work, health checks and delivery of the conversation error all stall.
Agent Server API · draft map
Action
Trigger an LLM connection error with default rich logging and probe GET /alive.
Observed
The health response takes 4.8–6.6 seconds, and socket error frames arrive equally late.
Expected
Health responses stay near 100 ms and below 1.5 seconds; the error reaches socket subscribers promptly.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
App backend status hangs holding the app lock when a child outlives the leader; expected unhealthy
If a Canvas backend exits while a child process still owns its output pipes, the status request hangs while holding the app lock. Stop, logs and other lifecycle actions then queue behind that request.
Agent Server API · draft map
Action
Let a backend leader exit while its child keeps stdout/stderr open; GET /api/canvas-extensions/installed/{name}/backend.
Observed
Status waits for the child to exit; stop, logs and other app lifecycle requests time out behind the lock.
Expected
Return unhealthy immediately so stop can terminate the leftover child process.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Canvas backend outlives a SIGKILLed agent-server; restart reports it stopped and starts a second one
A hard crash of a host-process agent server leaves Canvas backends running while the restarted server reports them stopped. Starting the app again creates a second backend sharing the same data directory.
Agent Server API · draft map
Action
Start a Canvas backend, SIGKILL and restart the agent server, then GET its backend status and POST .../backend/start.
Observed
Status says stopped with pid: null while the old process lives; start creates another process and stop kills only the new one.
Expected
No backend from the old server remains; an app has only one backend process.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Canvas App backend stays 'unhealthy' after one slow health probe instead of returning to 'ready'
One slow health probe can leave a Canvas backend permanently marked unhealthy even after it recovers. Requests through the app bridge then fail until the backend is explicitly stopped and started.
Agent Server API · draft map
Action
Make one backend health probe slow, restore normal health responses, then GET /api/canvas-extensions/installed/{name}/backend.
Observed
State remains unhealthy with the same PID; later calls do not probe again and the bridge returns 503.
Expected
The next successful health probe returns state: ready and restores bridge traffic.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
/app-backends WebSocket bridge negotiates no subprotocol (expected the backend's choice, qa.v1)
The Canvas backend WebSocket bridge drops subprotocol negotiation. Apps whose browser clients require a protocol such as graphql-ws cannot complete their normal handshake through the bridge.
Agent Server API · draft map
Action
Open WS /app-backends/qa-headers/ws offering Sec-WebSocket-Protocol: qa.v1.
Observed
The bridge selects no subprotocol and the backend reports none; a direct backend connection negotiates qa.v1.
Expected
The backend receives qa.v1 and the bridge selects qa.v1 in its 101 response.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
switch_llm returns 200 but keeps the old LLM when the posted usage_id is already registered
Switching models can return success while keeping the old model if the requested usage ID is already registered. This affects the default ID used by saved profiles after a conversation has received a message.
Agent Server API · draft map
Action
After a message registers the default usage ID: POST /api/conversations/{id}/switch_llm with another model using that ID.
Observed
200 {"success":true}, but GET still shows the original agent.llm.model.
Expected
Install the requested model, or return a descriptive 4xx for the occupied usage ID.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
LLM switched in before the first run is not added to ConversationStats; its tokens and cost are lost
Choosing another model before the first message can omit all of that model’s tokens and cost from conversation statistics. The run uses the new model, but cost displays and the run budget do not see its calls.
Agent Server API · draft map
Action
Switch the LLM before the first message or run, run the conversation, then GET /api/conversations/{id}.
Observed
The run uses the new model, but usage_to_metrics contains only condenser; the model’s usage is absent.
Expected
stats.usage_to_metrics contains the switched usage ID with its token usage and cost.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Continued /v1/chat/completions turn reports conversation-total usage (13/7/20), not its own (3/2/5)
A continued chat-completions request returns the token total for the whole conversation instead of that request. Clients that add usage across requests double-count earlier turns, including with streaming usage enabled.
Agent Server API · draft map
Action
POST /v1/chat/completions with X-OpenHands-ServerConversation-ID for a second turn.
Observed
The second response reports 13 prompt, 7 completion, 20 total: the accumulated conversation usage.
Expected
For the recorded 10/5 then 3/2 token calls, the second response reports 3 prompt, 2 completion, 5 total.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
execute_bash_command returns only the last chunk of output over 1 MiB; null exit_code past 100 chunks
The synchronous command endpoint returns only one chunk of large command output without warning. Beyond 100 chunks it can also lose the real exit code, which the TypeScript client then reports as success.
Agent Server API · draft map
Action
POST /api/bash/execute_bash_command
{"command":"yes a | head -c 1100000; echo QA_F13_TAIL","timeout":60}
Observed
Only the last 51,436 characters are returned. Beyond 100 chunks, exit_code can be null even when the command exits 7.
Expected
Return all 1,100,012 output characters and the final exit code.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
download-trajectory returns 404 after POST /api/init sets conversations_path (expected 200 zip)
A warm-pool server initialized with a custom conversation directory can read its conversations but cannot download their trajectories. Both download routes keep looking in the original launch-time directory.
Agent Server API · draft map
Action
POST /api/init with conversations_path; create a conversation; GET /api/file/download-trajectory/{id}.
Observed
404 Conversation not found on both trajectory routes, while GET /api/conversations/{id} returns 200.
Expected
200 with a zip from the configured conversation directory.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
GET /api/git/changes returns [] for a non-repo directory holding a repository, not its changes
The Git changes endpoint misses edits when the workspace directory contains a repository one level below it. A changes view querying the workspace root can therefore appear empty despite uncommitted work.
Agent Server API · draft map
Action
GET /api/git/changes?path=<workspace> where the repository is <workspace>/proj.
Observed
200 []; querying <workspace>/proj directly returns both changes.
Expected
List proj/README.md as UPDATED and proj/notes.txt as ADDED.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
git changes returns 400 for a modified file with a space in its name and mangles non-ASCII names
Modified filenames containing spaces can make the entire Git changes request fail. Non-ASCII names are returned in a corrupted escaped form, so clients cannot reliably identify the changed file.
Agent Server API · draft map
Action
Modify my notes.md or café.txt; GET /api/git/changes.
Observed
A space causes 400 Unexpected git diff output format; non-ASCII names retain Git quoting with corrupted slashes.
Expected
200 with the real filename, such as my notes.md or café.txt.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
The TypeScript conversation manager accepts iteration and stuck-detection options but never forwards them. A client asking for a seven-iteration limit silently gets the server default of 500.
POST and GET contain max_iterations: 500 and stuck_detection: true; calling RemoteConversation.start directly works.
Expected
POST and subsequent GET contain max_iterations: 7 and stuck_detection: false.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
TS RemoteWorkspace.downloadAsBlob/fileDownload return 11 mangled bytes for a 5-byte binary file
The TypeScript remote-workspace download helpers decode binary downloads as text and silently corrupt their bytes. The lower-level FileClient retrieves the same file correctly.
Agent Server API · draft map
Action
RemoteWorkspace.downloadAsBlob(path) for a file containing ff 00 80 fe 41.
Observed
An 11-byte text Blob: [239, 191, 189, 0, 239, 191, 189, 239, 191, 189, 65]; file_size is 11.
Expected
A 5-byte download: [255, 0, 128, 254, 65].
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
ConversationEventStream on Node 22 WebSocket stops after one refused reconnect instead of backing off
On Node 22’s native WebSocket, the TypeScript event stream stops retrying after its first refused reconnect. A routine server restart can therefore leave the client permanently disconnected.
Agent Server API · draft map
Action
Start ConversationEventStream with reconnect enabled on Node 22, restart the server, and let the first reconnect be refused.
Observed
isConnected: false, isReconnecting: true, attemptCount: 1; no further retries occur.
Expected
Keep backing off until the server returns, then reconnect and receive new events.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
RemoteState confirmation_policy/security_analyzer/agent fail or return None on the second read
Reading remote conversation state mutates the cached data, so the next read can throw or return the wrong result. Confirmation mode can change from true to false without any server-side change.
Agent Server API · draft map
Action
Read state.confirmation_policy, state.security_analyzer, and state.agent twice after filling RemoteState’s cache.
Observed
Second policy and agent reads raise ValidationError; the analyzer becomes None and confirmation mode becomes false.
Expected
Both reads return the same valid types; is_confirmation_mode_active stays true.
Source-derived API reproduction from the issue report; no screenshot or new test run. The verification map is PR #5621; its current state is recorded in the downloadable data.
Setting a Git Sync encryption key leaves already-synced automations as plaintext in the repository
Turning on Git Sync encryption left already-synced automation files readable in the repository. The merged fix re-exports them after a key change, so the next sync writes ciphertext.
First 68
BeforeBefore: saving the key leaves existing files as plaintext; the regression tests fail.AfterAfter: the next sync encrypts both files in one commit; the regression tests pass.
Original rendered terminal captures from the real-service test in PR #563, not UI screenshots. This fixes the next sync after saving a key; the PR records follow-ups for a cycle already in progress and environment-only key changes.
The first 68 come from #17961’s “Issues filed” inventory: 62 in Agent Canvas, five in the SDK, and one in automation. Later Canvas findings are included when their issue bodies or run reports tie them to a feature-map verification run. Each distinct issue is counted once, including follow-up reports of a recurring symptom; existing issues rechecked by later runs keep their original group. Themes are editorial groupings.
The separate verify-agent-server PR #5621 reports 31 new API bugs, catalogued in its umbrella issue. Those 31 have their own run label. The umbrella and requests to build the maps are not counted as additional bugs. Older issues merely reproduced by later runs are not added again.
Issue titles and states were read from GitHub. Fixing PRs were checked through closure events, explicit closing references, PR descriptions and issue discussions. General mentions and the feature-map PR itself are excluded. Related alternatives and partial fixes are labeled. “No fix PR found” means no fixing PR was identified in those sources at the check time.
PR state and issue state are separate. “Merged” describes the PR; it does not claim that every part of an issue was fixed. Check indicators use the newest run of each app, workflow and job on an open PR’s latest head. Older failed runs superseded by a successful run do not count as current failures; the original GitHub rollup remains in the downloadable data. A red indicator records a failing check, not a judgment about its cause. Closed PRs are not treated as merged.
The fix-fleet report records how work already claimed by contributors was left to them. This page tracks the fixes together, without attributing every PR to that fleet.