Long Agent Teams runs lost members for good: a truncated provider stream
ended a member's turn with nobody to wake it, the desktop Stop button and
every lead restart killed all members and marked the plan interrupted,
mail sent to a stopped member landed in an inbox nothing read, and a lead
kept inside one long turn never saw member reports. Aligned with the
official CLI 2.1.284 and verified with DeepSeek Flash through a
fault-injecting proxy.
Stream recovery
- Re-send a stream that breaks before any tool ran (proxy truncation,
transport errors), with the existing retry budget and backoff; the
desktop drops the discarded attempt's tool cards and todo update.
Desktop team runtime (teamPlanRuntime)
- The server supervises members: a stopped member restarts from its own
transcript when messaged; transient failures continue automatically
(15s/45s/2m/5m/10m) and only exhausted retries reach the lead; ready
dependent tasks wake their owner; a crash-loop guard ignores user stops.
- Stop pauses the team instead of ending it; the lead's next user message
is followed by a notice listing the stopped members and their open
tasks. Lead restarts (model/permission switch, crash) keep members;
server restarts re-own the team. Teams end on /clear or session delete.
- Approving a plan no longer races a concurrent plan read into
"Launch ownership was lost".
Mailbox and messaging
- Atomic inbox writes, identity-based read marking, read history files,
idle notifications with result/failureReason, and write failures
reported instead of "Message sent". External builds keep the official
between-turn delivery to the lead.
- SendMessage resumes non-running in-process teammates from their
transcript, notes restarting desktop members, queues mail for members
of a plan awaiting approval, and rejects unknown names.
CLI in-process teammates
- Compaction uses the teammate's own controller and real history and no
longer kills it on error; failed turns are classified and continued;
the turn-end mailbox drains as one batch; one durable transcript per
teammate.
Lead behaviour
- An unmet /goal ends the lead turn while members work, so member reports
arrive; WaitSessions on own team members returns immediately.
Desktop UI
- Member states for stopped, auto-retrying and failed, with reason,
countdown and recovery hint in all five locales.
Tests and tooling
- Regression tests for every behaviour above; module mocks in four test
files are restored after use so the single-process coverage run is not
polluted; the desktop smoke asserts the new Stop semantics.
Open the skills market on a bundled catalog of 398 curated ClawHub and
SkillHub skills in 13 categories instead of querying both registries live.
Live search stays available as an explicit "search all markets" scope
whose results are marked as not curated.
- Server: catalog scope (default) with category filter, filtering before
pagination and batched install state; catalog metadata overlaid on live
detail; per-scanner ClawHub reports, changelog and page URL; manual
catalog refresh script.
- Pin ClawHub reads and installs to the card's owner so a same-slug copy
by another author is never shown or installed in its place.
- Desktop: category chips, curated cards with tags, locale-aware summaries,
detail page with stats, security report, changelog, capability panel and
triggers; install confirmation requires acknowledgement for unaudited or
flagged skills.
- Docs: rewrite the skills market section and refresh screenshots.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397
The "Session running" dialog treats a running background task as the session
being busy, but Stop & Close only called stopGeneration, which interrupts the
foreground turn and Agent tasks. A background shell command (and its CLI) kept
running, so the reopened session spun again and asked the same question on the
next close.
Stop & Close now also stops every running background task, before the socket is
closed. "Is the session running" and "which tasks to stop" share one predicate
in backgroundTasks, so the dialog cannot offer a Stop that leaves the task that
raised it alive. Keep running sends nothing.
Refs #1398
Everything the assistant did between a <task-notification> turn and the next
real user message was hidden, so the reply a model gave once a background
command finished (and any tool work after it) never reached the chat, live or
after a reload, even though it was written to the transcript.
Drop that suppression from every session history projection (full, paged,
recovery, sub-agent lookup) and from the desktop store (history mapping and the
live-stream flag). Only the injected notification prompt stays hidden; the
background task cards still come from its notification data. The suppression
was added for #886 to quiet replies to stale notifications; the CLI already
avoids those at the source when the model has read the task's result.
Refs #1389
Add mobile provider and General settings while preserving desktop behavior.
Harden remote credential handling, ngrok session ownership and consent upgrades.
Add a persistent isolated JavaScript worker for native app actions and batch
known operations without a model round trip between each input. Preserve
per-cell context, native errors, screenshot coordinates, and image types.
Align macOS gesture, key, inventory, capture, scroll, and clipboard behavior;
include a signed native receiver fixture and compiled sidecar regression tests.
Keep the Windows pixel route and fix cancellation with session-owned mouse
cleanup, lock revalidation, and portable signing-fixture tests.
Add shared project and session selection across IM adapters, preserve bindings on
failed restoration, and synchronize permissions with desktop clients.
Fixes#1286
- Add 新功能支持 issue template, refine bug/question templates with a
pointer to the user group
- Replace the Feishu user group QR with the WeChat Work group QR in both
READMEs, include contact info for enterprise/Agent customization
- Make the Chinese README the default README.md, keep the English one as
README.en.md, update language badges on both
- Update PR change policy, pr-triage docs set and site docs check for the
renamed README files
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Connecting Feishu meant creating a bot by hand on the open platform and
pasting an App ID and App Secret back. Feishu also exposes an RFC 8628
device-authorization flow, so the desktop can now render a QR code, and
confirming it in the app creates the bot and stores its credentials
directly. `adapters/feishu/registration.ts` implements that protocol
rather than importing `registerApp` from `@larksuiteoapi/node-sdk@1.73`:
the repository pins 1.60 for the chat client, and the SDK runs the whole
poll inside one un-cancellable promise where the desktop needs the
stateless begin/poll pair the DingTalk registration already uses. The
scan is create-only, so it can never rewrite the configuration of a bot
the user already runs, and it pre-fills exactly the scopes, events and
callbacks this adapter calls. International tenants finish on Lark's
domain, which is now persisted and honoured by the client.
WeCom, QQ and Slack join the same session model. WeCom and QQ bind by
scanning; Slack has no scan flow, so it uses an app manifest that
pre-fills the scopes and Socket Mode. All three run over long
connections, so no public callback URL is needed, and all three accept
private chats only — pairing authorizes one person, and answering in a
group would extend that authorization to everyone else in the room.
They are built on a new `adapters/common/chat-runtime.ts` instead of a
fourth copy of the loop the five existing adapters each carry. A platform
supplies a `ChatPort` — how to say something, how to open a streaming
reply, optionally how to send an image — and the runtime owns pairing,
command routing, session restore, permission bookkeeping and the
translation of the server's stream. The existing five are deliberately
left on their own copies; migrating them is a separate change with its
own regression surface.
Attachments are downloaded through a deferred loader that runs after the
pairing gate and inside the per-chat queue. Resolving them eagerly would
let an unpaired stranger make the adapter fetch bytes and write them
under ~/.claude/im-downloads — on Slack with the bot token attached —
and would let a slow attachment overtake a text message sent after it.
The sidecar launcher's per-adapter branches become one table. It is
declared above the mode dispatch on purpose: `runAdapters` is hoisted and
runs at module top level, so a table declared below it is still in its
temporal dead zone when the adapters mode reads it — which type checks,
lints and unit tests all miss, and only the compiled binary reveals.
Verified with the checks `check:impact` selects: adapters, server,
desktop, electron, policy, chat-contract, agent-flow, docs, native
(sidecar compile, packaging and an adapters-mode smoke against the real
binary) and coverage. The scan flows themselves are not verified against
live platforms — that needs real WeCom, QQ and Slack accounts and would
create real bots.
Claude-Session: https://claude.ai/code/session_01CCGoP316AK7wdQG3Ms6Uwq
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
Add Atlas Cloud to the sponsor tables in both README.md and
README.zh-CN.md, using a cc-haha-specific campaign link. Ship
light/dark logo variants so the wordmark stays visible in both
GitHub themes.
The instruction files told every coding agent to reach for agent-browser
whenever a change needed browser-level evidence: copilot-instructions
listed "E2E or agent-browser smoke" as the remedy for cross-boundary
flows, and both contributing guides repeated it. That wording outlived
the tool. With the agent-browser skill uninstalled, agents still parsed
those lines as a recommendation and went looking for the binary instead
of using the browser skill that is actually installed.
Deleting the references would have made the docs wrong. agent-browser is
still a real dependency: check:desktop-ui-smoke spawns it on Linux CI,
and seven maintainer-run e2e scripts under desktop/scripts drive it
directly. It cannot be swapped for ego-browser either — ego lite is a
macOS-only GUI app with no headless mode and a one-time interactive
onboarding, so it cannot run on ubuntu-latest at all.
So the lanes keep the binary and the prose loses the recommendation.
agent-browser is now described as an implementation detail of those two
call sites, and ad-hoc browser work — manual verification, screenshots,
exploratory UI checks — is pointed at the ego-browser skill.
The quality contract asserted the old string, so it would have failed
closed on the reworded line. It now pins the replacement plus the new
routing rule; flipping either sentence turns the test red.
Built-in agents pin their own models — Explore and claude-code-guide run
on Haiku, statusline-setup on Sonnet — and there was no way to change
that. The only override mechanism was a same-named user agent, which
replaces the definition wholesale: AGENT_SLUG_PATTERN is lowercase-only
so `Explore` and `Plan` cannot even be created, and a replacement loses
the runtime getSystemPrompt and the built-in tool privileges keyed off
`source`. CLAUDE_CODE_SUBAGENT_MODEL is the only other lever and forces
every subagent onto one model.
Add `builtInAgentOverrides` to settings.json, carrying model and effort
per agentType. It is applied in getBuiltInAgents(), the single choke
point every consumer goes through, so the effective value reaches
spawning, `/agents` and the desktop list without any of them knowing an
override exists — and `source` stays `built-in`, preserving the prompt
and tool privileges.
Three constraints are load-bearing rather than stylistic:
- `model: "inherit"` is a real value, not a reset. Built-in defaults
differ per agent and per build, so clearing means deleting the field;
the entry and then the key are removed once empty.
- getSystemPrompt is never wrapped. serializeActiveAgent branches on
`.length === 0`, and claude-code-guide declares one parameter while
the others declare none, so a wrapper makes it destructure undefined.
- The settings write is a read-modify-write inside the file lock.
updateUserSettings is a top-level shallow merge and would replace the
whole record, losing an update when two agents are changed in quick
succession.
strictPluginOnlyCustomization is enforced when resolving, not only when
writing, since settings.json is user-editable by definition.
Also surface edit and delete on the agent list rows. Both already
existed but were reachable only after opening an agent's detail page.
Rows become a div with the primary button and the actions as siblings
rather than nested buttons, keep focus-within so the controls are not
tabbable while invisible, and stay visible on touch where hover never
fires. Built-in rows get the override entry point and no delete.
The store's create/update/override paths now share runAgentMutation
instead of a third hand-copy of the out-of-order guards; the existing
store tests pass unchanged.
The glowing border framed the app being driven. It was purely decorative and
the user asked for it to go — the animated cursor already shows what Claude is
acting on.
It was not, as suspected, interfering with clicks: the overlay is
`ignoresMouseEvents`, sits at `CGShieldingWindowLevel()`, and every window
enumeration in `WindowGeometry` filters to `layer == 0`, so it could not enter
any hit-test or occlusion decision. It went because it earns nothing, not
because it broke anything.
`WindowFrameTracker` goes with it: its own header says it exists to keep the
glow glued to a moving window, and nothing else referenced it. That also
resolves the "30Hz CGWindowList polling" entry on the redesign doc's cut list —
the polling was the tracker's fallback path, so it is gone rather than reduced.
`NonAnimatedGradientLayer` moved into `VirtualCursor.swift`. It lived in the
glow file but the cursor's orb needs it too (without it Core Animation's
implicit 0.25s actions smear every frame change), and the compiler caught the
dangling reference.
`overlay_show` now only reveals the cursor, so `playSound` and
`stopOverlaySession(animated:)` lost their only consumers. The TypeScript side
had eleven comments describing a border that no longer exists; they now describe
what the code does.
restoreAvailable conflated two questions: whether the files this checkpoint
reports can be put back, and whether the checkpoint saw every file the turn
touched. Any tool outside an 18-name allowlist — Bash, PowerShell, TaskCreate,
every MCP tool — forced the second to false, and one such call anywhere from
the target turn onward disabled undo for the whole range. In practice that is
every real turn, so v0.5.3 shipped with undo effectively dead.
Split the two. restoreAvailable now answers only the first question.
The second becomes unverifiedChangeSources: tool names whose file effects the
checkpoint could not capture. Undo stays available and restores exactly the
files it lists, and the card, the confirmation, and the completion toast each
name what it is leaving behind. An unrecognized tool now costs a warning
instead of the feature. A transcript that cannot be read still blocks, because
then even the reported file list may be wrong.
Bash calls are classified against the existing read-only allowlist via a new
recordedCommandIsReadOnly, which drops the sandbox/cwd checks that describe the
live process rather than the replayed session. `git status` no longer warns at
all, so the warning means something when it appears.
Rewind also takes a mode. `conversation` skips the file restore entirely, so a
turn whose files cannot be restored no longer costs the user the ability to
back out of the prompt — matching how upstream keeps Restore conversation
independent of Restore code.
Files written by shell commands were never recoverable here; the checkpoint
only ever covered the structured file tools. Reporting that is honest, and
docs now say so.
Since v0.5.1 every IM channel listed only the default project. All five
adapters passed the default work dir to AdapterHttpClient as the sole
allowed project root, so listRecentProjects filtered out everything else;
matchProject, listSessions, sessionExists, createSession and listSkills
were clamped the same way. Feishu is where it was reported, but telegram,
wechat, dingtalk and whatsapp were identical.
defaultWorkDir is documented as where a new IM session starts, not as an
access boundary. Using it as the boundary failed both ways: configured, it
hid every other project; blank, it falls back to PWD/cwd(), which is "/"
for a GUI-launched sidecar, so the boundary allowed the whole filesystem.
Split the two concepts. allowedProjectRoots is now its own setting (global,
per-platform, or ADAPTER_ALLOWED_PROJECT_ROOTS), resolved together with the
work dir by resolveAdapterWorkspace so the default project is always inside
the boundary and /new cannot fail on inconsistent config. The default is the
home directory; it refuses to inherit "/" or any ancestor of home. Pairing
remains the primary authorization control, so unusable roots warn and fall
back rather than locking the bot out.
All five entrypoints now build their client through createAdapterClient
instead of repeating the wiring, which is what let one defect appear in five
places at once.
Known gap, left for a follow-up: a project outside the boundary is still
reported as "not found" rather than "outside the allowed directories".
Rebuilt against main so the branch carries the Computer Use work and no other
divergence. Three unrelated efforts had been sitting uncommitted in this
worktree and were swept into an earlier commit; they are preserved on
cu-worktree-full-backup and belong on their own branches — adapter control
credentials, Electron asar sealing, and the sidecar code-loading audit. Every
file outside Computer Use now matches main exactly.
The engine
A Swift helper drives apps through the accessibility tree, with coordinate
actuation for the Chromium and Electron apps whose tree is a bare window
frame. Ten primitives matching the shape Codex uses, so an app's guidance and
the model's habits transfer.
Coordinate actions resolve their target window once and refuse when none can
be named. The unbound event they used to fall back to is discarded by custom
renderers, so a minimized target produced a whole session of "Action
completed" with nothing behind it.
Input acceptance is established for typing and key presses as well as clicks:
each MCP call is seconds apart, so the keyboard cannot inherit the focus a
click established. The synthetic focus notification is gated on the target
not already being active — sent unconditionally it names window 0 at an app
that already owns a key window, and nine window-bound clicks were discarded
with the traffic lights fully lit.
State the model can trust
An off-screen target says so, and says which tools still reach it: element
actions need no on-screen geometry, so an app with a real tree can still be
driven from the Dock. A fully covered window is recovered once, then left
alone — burying it again is the user wanting their screen back. A repeated
capture is reported with the cause that actually applies rather than both,
because coverage is something we compute.
Signing
The helper is signed under a stable identity before electron-builder sees it,
and excluded from re-signing: macOS ties Accessibility and Screen Recording
grants to the signing identity, so rotating it drops both on every update.
Discoverability
The desktop slash menu falls back to a directory scan while a session's CLI
has not started, which is when the menu is first opened. Built-ins and
bundled skills live in the binary, so /computer-use was absent until after
the first message.
One conflict, same structural cause as the last merge: git offers the pre-split
Settings.tsx against the 183-line shell. All four of main's hunks belong to
ProviderFormModal, which now lives in settings/ProviderSettings.tsx — the useId
import, the addToast on a successful model fetch, and the base-URL field growing
an explicit label plus a help tooltip. Ported there; tsc caught the one import
(useUIStore) the move needed.
Everything else auto-merged, including the two files both sides changed:
chatStore.ts keeps main's pushAssistantHistoryThinking alongside the removal of
the content-equality replay guard — different functions, history mapping versus
the live path — and the five locales land at 2521 keys each.
Note: six generalSettings tests fail after this merge and they fail identically
on main. 62f648cb2 added an IconButton labelled 'Base URL help' next to the base
URL input without updating generalSettings.test.tsx, so getByLabelText(/Base
URL/i) now matches both. Confirmed against a temporary worktree at main: 6
failed / 104 passed there too. Not introduced here, and fixed separately.
4f9fec876 added this workflow with `cron: '0 18 * * *'`. That was the wrong call
to make unilaterally: the repository had no scheduled workflow at all before it,
so this was not one more cron among several but the introduction of recurring CI
spend — about ninety minutes per run — on a schedule nobody asked for.
The reasoning for the sweep still holds: a per-PR gate only covers what the diff
reaches, so it is blind to checks no recent PR selected and to failures that only
appear when the whole suite runs together. Keeping the workflow on
`workflow_dispatch` keeps that one click away without deciding for the maintainer
when to spend the time.
pr-quality-workflow.test.ts now asserts the absence of `schedule:` and `cron:`
rather than their presence, so a schedule cannot drift back in unnoticed —
verified by adding the cron back and watching the test go red. The docs' four-tier
table renames the tier accordingly; calling it "Nightly" when nothing runs nightly
is exactly the kind of comment that outlives its code.
- Add both gateways as featured presets pointing at their Anthropic-compatible
roots (https://api.fenno.ai, https://api.qnaigc.com) with auth_token strategy
- Ship no default model ids: the available catalog depends on the plan the user
bought, so they fetch the live list and pick one instead
- Keep modelContextWindows, which stays useful per picked model id and covers
ids the built-in table cannot resolve (claude-opus-5, namespaced 七牛云 ids)
- Add both as sponsors in the English and Chinese READMEs, and list them in the
preset docs