Every session route resolves its transcript first, and that lookup parsed
the whole file to rank candidates even when only one file matched. The
desktop pages a long session's history request by request, so reopening
it cost one full parse per page and grew quadratically with file size.
Skip the content check when a single file matches, and stop a multi-file
check at the first conversation record. Ranking and API responses are
unchanged.
Long Agent Teams runs lost members for good: a truncated provider stream
ended a member's turn with nobody to wake it, the desktop Stop button and
every lead restart killed all members and marked the plan interrupted,
mail sent to a stopped member landed in an inbox nothing read, and a lead
kept inside one long turn never saw member reports. Aligned with the
official CLI 2.1.284 and verified with DeepSeek Flash through a
fault-injecting proxy.
Stream recovery
- Re-send a stream that breaks before any tool ran (proxy truncation,
transport errors), with the existing retry budget and backoff; the
desktop drops the discarded attempt's tool cards and todo update.
Desktop team runtime (teamPlanRuntime)
- The server supervises members: a stopped member restarts from its own
transcript when messaged; transient failures continue automatically
(15s/45s/2m/5m/10m) and only exhausted retries reach the lead; ready
dependent tasks wake their owner; a crash-loop guard ignores user stops.
- Stop pauses the team instead of ending it; the lead's next user message
is followed by a notice listing the stopped members and their open
tasks. Lead restarts (model/permission switch, crash) keep members;
server restarts re-own the team. Teams end on /clear or session delete.
- Approving a plan no longer races a concurrent plan read into
"Launch ownership was lost".
Mailbox and messaging
- Atomic inbox writes, identity-based read marking, read history files,
idle notifications with result/failureReason, and write failures
reported instead of "Message sent". External builds keep the official
between-turn delivery to the lead.
- SendMessage resumes non-running in-process teammates from their
transcript, notes restarting desktop members, queues mail for members
of a plan awaiting approval, and rejects unknown names.
CLI in-process teammates
- Compaction uses the teammate's own controller and real history and no
longer kills it on error; failed turns are classified and continued;
the turn-end mailbox drains as one batch; one durable transcript per
teammate.
Lead behaviour
- An unmet /goal ends the lead turn while members work, so member reports
arrive; WaitSessions on own team members returns immediately.
Desktop UI
- Member states for stopped, auto-retrying and failed, with reason,
countdown and recovery hint in all five locales.
Tests and tooling
- Regression tests for every behaviour above; module mocks in four test
files are restored after use so the single-process coverage run is not
polluted; the desktop smoke asserts the new Stop semantics.
Open the skills market on a bundled catalog of 398 curated ClawHub and
SkillHub skills in 13 categories instead of querying both registries live.
Live search stays available as an explicit "search all markets" scope
whose results are marked as not curated.
- Server: catalog scope (default) with category filter, filtering before
pagination and batched install state; catalog metadata overlaid on live
detail; per-scanner ClawHub reports, changelog and page URL; manual
catalog refresh script.
- Pin ClawHub reads and installs to the card's owner so a same-slug copy
by another author is never shown or installed in its place.
- Desktop: category chips, curated cards with tags, locale-aware summaries,
detail page with stats, security report, changelog, capability panel and
triggers; install confirmation requires acknowledgement for unaudited or
flagged skills.
- Docs: rewrite the skills market section and refresh screenshots.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397
Page transcript and trace reads, bound UI caches and retained task records,
and replace full-file background polling with incremental projections.
Preserve recovery and ownership semantics across pages and cancel stale work.
OpenCode Go binds the wire format to the URL path and translates nothing
between formats, so one provider record serves /chat/completions, /messages
and /responses depending on the model, each accepting a different credential
header. A record carries only one apiFormat, so presets can now declare
ordered per-model prefix rules (modelApiFormats) and the proxy resolves the
effective format from the request body. Only the exceptions are listed;
anything unmatched keeps the provider's format, which is the endpoint with the
broadest compatibility. A preset with rules is authoritative for apiFormat,
because a value written by a cc-switch import or the edit form would otherwise
silently disable every rule and point the CLI straight at the upstream.
The gateway also rejects any request without a stable per-conversation
x-opencode-session, so presets can declare upstreamHeaders with $SESSION_ID
and $VERSION placeholders. The id is the one the CLI already sends; it is
redacted in traces the same way the credential is, since it also names the
local transcript files.
Title generation builds its own upstream request and bypassed all of the
above, which left AI titles failing for every provider that needs local
request handling; it now goes through the same proxy the CLI uses.
Verified against the live gateway: glm/kimi reach /chat/completions,
minimax/qwen/union-alpha reach /messages, grok/gpt reach /responses, each
with the credential that endpoint accepts.
README.md carried the Chinese version while README.en.md held the English one,
so the GitHub landing page opened in Chinese. Swap them: README.md is now
English and the Chinese version lives in README.zh-CN.md, with both language
switchers pointing at the new paths.
Route native-only changes through macOS checks and verify relocated cursor resources in final packages. Reject resources that escape the app and exercise visible click feedback against a disposable native receiver.
Connect the regressions to required checks and capture shell fixture output through temporary files.
Revert efe5cae19 so existing sessions remain usable and models can be
switched without protocol admission checks. Retain the additive index
schema for already-upgraded caches and rebuild their summaries.
Add a persistent isolated JavaScript worker for native app actions and batch
known operations without a model round trip between each input. Preserve
per-cell context, native errors, screenshot coordinates, and image types.
Align macOS gesture, key, inventory, capture, scroll, and clipboard behavior;
include a signed native receiver fixture and compiled sidecar regression tests.
Keep the Windows pixel route and fix cancellation with session-owned mouse
cleanup, lock revalidation, and portable signing-fixture tests.
- Add 新功能支持 issue template, refine bug/question templates with a
pointer to the user group
- Replace the Feishu user group QR with the WeChat Work group QR in both
READMEs, include contact info for enterprise/Agent customization
- Make the Chinese README the default README.md, keep the English one as
README.en.md, update language badges on both
- Update PR change policy, pr-triage docs set and site docs check for the
renamed README files
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
The instruction files told every coding agent to reach for agent-browser
whenever a change needed browser-level evidence: copilot-instructions
listed "E2E or agent-browser smoke" as the remedy for cross-boundary
flows, and both contributing guides repeated it. That wording outlived
the tool. With the agent-browser skill uninstalled, agents still parsed
those lines as a recommendation and went looking for the binary instead
of using the browser skill that is actually installed.
Deleting the references would have made the docs wrong. agent-browser is
still a real dependency: check:desktop-ui-smoke spawns it on Linux CI,
and seven maintainer-run e2e scripts under desktop/scripts drive it
directly. It cannot be swapped for ego-browser either — ego lite is a
macOS-only GUI app with no headless mode and a one-time interactive
onboarding, so it cannot run on ubuntu-latest at all.
So the lanes keep the binary and the prose loses the recommendation.
agent-browser is now described as an implementation detail of those two
call sites, and ad-hoc browser work — manual verification, screenshots,
exploratory UI checks — is pointed at the ego-browser skill.
The quality contract asserted the old string, so it would have failed
closed on the reworded line. It now pins the replacement plus the new
routing rule; flipping either sentence turns the test red.
The smoke waited on `textarea`, but the composer became a ProseMirror
contenteditable (MentionComposer) some time ago, so the selector could
never match. The lane hung for its full 30s timeout and failed on every
run — the screenshot it captured on the way out showed the app rendered
and idle, which is what makes this easy to misread as a UI break.
Target `[data-composer-editor]`, the attribute the editor puts on its
editable node and the same hook composerTestUtils already drives.
Two things, in one commit because both touch `src/server/api/computer-use.ts`
and the halves cannot be split without rewriting the file twice.
## The regression
`desktop/package.json` had lost `"claude-sidecar-[^/]+$"` from `mac.signIgnore`.
That entry is not a hardening nicety — it is load-bearing for attestation:
1. `build-sidecars.ts` signs the sidecar with an explicit
`--identifier com.claude-code-haha.desktop.sidecar`
2. `ClientAttestation.swift` compares that identifier EXACTLY in
`validDesktopChain`
3. without the exclusion, electron-builder re-signs the sidecar and drops the
flag, so codesign derives the identifier from the file name
(`claude-sidecar-aarch64-apple-darwin`), which never matches
Measured on the shipped 0.5.3 build: host and helper identifiers were correct,
the sidecar's was not. `validDesktopChain` backs BOTH `authorizeOneShot` and
`authorizeDaemon`, so this was not merely a broken permission probe — every
Computer Use call in that build failed closed with `unauthorized_client`.
The existing guard was a literal `toEqual` on the whole signIgnore array, which
does not survive an edit that changes the array and the expectation together —
exactly how the entry was lost. The replacement asserts the behaviour instead:
real sidecar file names must match some exclusion pattern, and the identifier
constants in `sign-identity.ts` and `ClientAttestation.swift` must agree.
`checkCuHelperPermissions` swallowed the failure into nulls, which the settings
page renders as a permanent "checking…" — indistinguishable from a probe still
in flight, and the only symptom this bug ever produced. It now records an
error-level diagnostic: the helper binary is present, so the user did nothing
wrong and the check still did not complete.
## App icons
Rows in the picker and the authorized list showed a letter tile. They now show
the application's own icon, resolved the way Finder does it: `CFBundleIconFile`
from Info.plist, the `.icns` under `Contents/Resources`, rasterised with `sips`.
`openTargetService` already does this, but its resolver is keyed on a
`TargetDefinition` and cannot answer for an arbitrary installed app, so the
path-only half lives in `macAppIcon.ts` and that service is left alone.
The endpoint takes a bundle id and resolves the path itself. There is
deliberately no parameter that names a file — it rasterises and returns bytes,
so its input surface is a security property, and a test drives paths at it.
Enumeration is shared across concurrent lookups: opening the picker fires one
icon request per visible row while the cache is still cold, and a plain
check-then-fill cache would walk every application root once per row.
macOS only, matching where this engine exists. The Windows list renders no icon
slot at all, and Linux is not a supported Computer Use platform.
## Verification
- 208 installed applications on the dev machine: 204 icons resolved; the 4
misses are background bundles (Adobe sync extension, a URL handler, an
updater, a token host) that ship no icon and correctly fall back
- check:server 3345 pass, desktop 4101 pass, check:policy 243 pass, lint clean
(the 2 failures in each suite are `*.golden.test.ts`, pre-existing on main)
- mutation-checked that the new guards actually fail: removing the signIgnore
entry, dropping the icon `onError` fallback, and breaking the in-flight share
each turn a test red
Rebuilt against main so the branch carries the Computer Use work and no other
divergence. Three unrelated efforts had been sitting uncommitted in this
worktree and were swept into an earlier commit; they are preserved on
cu-worktree-full-backup and belong on their own branches — adapter control
credentials, Electron asar sealing, and the sidecar code-loading audit. Every
file outside Computer Use now matches main exactly.
The engine
A Swift helper drives apps through the accessibility tree, with coordinate
actuation for the Chromium and Electron apps whose tree is a bare window
frame. Ten primitives matching the shape Codex uses, so an app's guidance and
the model's habits transfer.
Coordinate actions resolve their target window once and refuse when none can
be named. The unbound event they used to fall back to is discarded by custom
renderers, so a minimized target produced a whole session of "Action
completed" with nothing behind it.
Input acceptance is established for typing and key presses as well as clicks:
each MCP call is seconds apart, so the keyboard cannot inherit the focus a
click established. The synthetic focus notification is gated on the target
not already being active — sent unconditionally it names window 0 at an app
that already owns a key window, and nine window-bound clicks were discarded
with the traffic lights fully lit.
State the model can trust
An off-screen target says so, and says which tools still reach it: element
actions need no on-screen geometry, so an app with a real tree can still be
driven from the Dock. A fully covered window is recovered once, then left
alone — burying it again is the user wanting their screen back. A repeated
capture is reported with the cause that actually applies rather than both,
because coverage is something we compute.
Signing
The helper is signed under a stable identity before electron-builder sees it,
and excluded from re-signing: macOS ties Accessibility and Screen Recording
grants to the signing identity, so rotating it drops both on every update.
Discoverability
The desktop slash menu falls back to a directory scan while a session's CLI
has not started, which is when the menu is first opened. Built-ins and
bundled skills live in the binary, so /computer-use was absent until after
the first message.