* fix(agent-teams): keep a stopped team stopped and wake members after any failure
Stop killed every member, but the lead kept polling its mailbox: reports
sent before the Stop started hidden lead turns, and a single SendMessage
from one of them unpaused the whole team. Stop now sends team_plan_pause
so the lead holds teammate mail until the user's next message, and the
server only lets lead instructions written after that message resume the
team. A restart in flight when Stop arrives is stopped as well, and the
idle Stop button hides once every member is stopped.
Members that stopped on a usage limit, billing, network or crash failure
were only resumed if the lead happened to message them. The user's next
message now lists every stopped member with its failure and open tasks,
once per failure, whether or not Stop was pressed.
* feat(desktop): show where each team member stands on its tasks
A member that stopped or failed left its task reading "In progress" with
an animated bar, and its card said only "Stopped" or "Error". Tasks whose
owner stopped, failed or waits to retry now show that state without
animation, and the card says which task the member stopped on.
The member card chips tell done, current and next tasks apart, with the
subject on hover. The member drawer groups tasks into now, up next and
done, names the unfinished dependencies of blocked work, shows done as
3/5, and no longer prints +0:00 for a duration polling could not measure.
Sessions get a Chat / Trajectory switch in the header. The trajectory is a
dense one-line-per-event ledger (system prompt, user, injected context such
as skills and system reminders, assistant responses, tool calls) with a
three-lane minimap, turn folding, search, and a detail panel per row.
- CLI: record a deduplicated prompt snapshot sidecar for desktop sessions
(system prompt, tool catalog, user context) without touching the transcript.
- Server: project transcript records into trajectory rows with bounded,
cursor-paged reads, live appends, a turn index, row detail by byte range,
snapshot blobs, and a time-window lookup into the raw trace capture.
- Desktop: windowed ledger, minimap, detail panel with a raw request summary,
chat <-> trajectory navigation, subagent drill-in.
- Remove the Settings trace list, trace tabs and the standalone trace
window; migrate persisted trace tabs to session tabs.
* feat(desktop): edit a sent prompt and rerun from there (#1343)
Hovering a prompt the rewind API can already target now offers "Edit and
resend". The bubble turns into an inline editor; sending dry-runs the
existing rewind, confirms when later turns or restorable files are at
stake, rewinds with the same conversation/both modes as undo, reloads
history and sends the edited prompt. A failed rewind changes nothing and
keeps the draft; if the edit cannot be sent after a successful rewind it
is handed back to the composer.
Undo and edit share one rewind routine, the unused per-message
rewindAction prop is removed, and TextArea forwards its ref.
* fix(desktop): preserve edit-resend session and context
Long Agent Teams runs lost members for good: a truncated provider stream
ended a member's turn with nobody to wake it, the desktop Stop button and
every lead restart killed all members and marked the plan interrupted,
mail sent to a stopped member landed in an inbox nothing read, and a lead
kept inside one long turn never saw member reports. Aligned with the
official CLI 2.1.284 and verified with DeepSeek Flash through a
fault-injecting proxy.
Stream recovery
- Re-send a stream that breaks before any tool ran (proxy truncation,
transport errors), with the existing retry budget and backoff; the
desktop drops the discarded attempt's tool cards and todo update.
Desktop team runtime (teamPlanRuntime)
- The server supervises members: a stopped member restarts from its own
transcript when messaged; transient failures continue automatically
(15s/45s/2m/5m/10m) and only exhausted retries reach the lead; ready
dependent tasks wake their owner; a crash-loop guard ignores user stops.
- Stop pauses the team instead of ending it; the lead's next user message
is followed by a notice listing the stopped members and their open
tasks. Lead restarts (model/permission switch, crash) keep members;
server restarts re-own the team. Teams end on /clear or session delete.
- Approving a plan no longer races a concurrent plan read into
"Launch ownership was lost".
Mailbox and messaging
- Atomic inbox writes, identity-based read marking, read history files,
idle notifications with result/failureReason, and write failures
reported instead of "Message sent". External builds keep the official
between-turn delivery to the lead.
- SendMessage resumes non-running in-process teammates from their
transcript, notes restarting desktop members, queues mail for members
of a plan awaiting approval, and rejects unknown names.
CLI in-process teammates
- Compaction uses the teammate's own controller and real history and no
longer kills it on error; failed turns are classified and continued;
the turn-end mailbox drains as one batch; one durable transcript per
teammate.
Lead behaviour
- An unmet /goal ends the lead turn while members work, so member reports
arrive; WaitSessions on own team members returns immediately.
Desktop UI
- Member states for stopped, auto-retrying and failed, with reason,
countdown and recovery hint in all five locales.
Tests and tooling
- Regression tests for every behaviour above; module mocks in four test
files are restored after use so the single-process coverage run is not
polluted; the desktop smoke asserts the new Stop semantics.
Open the skills market on a bundled catalog of 398 curated ClawHub and
SkillHub skills in 13 categories instead of querying both registries live.
Live search stays available as an explicit "search all markets" scope
whose results are marked as not curated.
- Server: catalog scope (default) with category filter, filtering before
pagination and batched install state; catalog metadata overlaid on live
detail; per-scanner ClawHub reports, changelog and page URL; manual
catalog refresh script.
- Pin ClawHub reads and installs to the card's owner so a same-slug copy
by another author is never shown or installed in its place.
- Desktop: category chips, curated cards with tags, locale-aware summaries,
detail page with stats, security report, changelog, capability panel and
triggers; install confirmation requires acknowledgement for unaudited or
flagged skills.
- Docs: rewrite the skills market section and refresh screenshots.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397
The "Session running" dialog treats a running background task as the session
being busy, but Stop & Close only called stopGeneration, which interrupts the
foreground turn and Agent tasks. A background shell command (and its CLI) kept
running, so the reopened session spun again and asked the same question on the
next close.
Stop & Close now also stops every running background task, before the socket is
closed. "Is the session running" and "which tasks to stop" share one predicate
in backgroundTasks, so the dialog cannot offer a Stop that leaves the task that
raised it alive. Keep running sends nothing.
Refs #1398
Everything the assistant did between a <task-notification> turn and the next
real user message was hidden, so the reply a model gave once a background
command finished (and any tool work after it) never reached the chat, live or
after a reload, even though it was written to the transcript.
Drop that suppression from every session history projection (full, paged,
recovery, sub-agent lookup) and from the desktop store (history mapping and the
live-stream flag). Only the injected notification prompt stays hidden; the
background task cards still come from its notification data. The suppression
was added for #886 to quiet replies to stale notifications; the CLI already
avoids those at the source when the model has read the task's result.
Refs #1389
Add mobile provider and General settings while preserving desktop behavior.
Harden remote credential handling, ngrok session ownership and consent upgrades.
Built-in agents pin their own models — Explore and claude-code-guide run
on Haiku, statusline-setup on Sonnet — and there was no way to change
that. The only override mechanism was a same-named user agent, which
replaces the definition wholesale: AGENT_SLUG_PATTERN is lowercase-only
so `Explore` and `Plan` cannot even be created, and a replacement loses
the runtime getSystemPrompt and the built-in tool privileges keyed off
`source`. CLAUDE_CODE_SUBAGENT_MODEL is the only other lever and forces
every subagent onto one model.
Add `builtInAgentOverrides` to settings.json, carrying model and effort
per agentType. It is applied in getBuiltInAgents(), the single choke
point every consumer goes through, so the effective value reaches
spawning, `/agents` and the desktop list without any of them knowing an
override exists — and `source` stays `built-in`, preserving the prompt
and tool privileges.
Three constraints are load-bearing rather than stylistic:
- `model: "inherit"` is a real value, not a reset. Built-in defaults
differ per agent and per build, so clearing means deleting the field;
the entry and then the key are removed once empty.
- getSystemPrompt is never wrapped. serializeActiveAgent branches on
`.length === 0`, and claude-code-guide declares one parameter while
the others declare none, so a wrapper makes it destructure undefined.
- The settings write is a read-modify-write inside the file lock.
updateUserSettings is a top-level shallow merge and would replace the
whole record, losing an update when two agents are changed in quick
succession.
strictPluginOnlyCustomization is enforced when resolving, not only when
writing, since settings.json is user-editable by definition.
Also surface edit and delete on the agent list rows. Both already
existed but were reachable only after opening an agent's detail page.
Rows become a div with the primary button and the actions as siblings
rather than nested buttons, keep focus-within so the controls are not
tabbable while invisible, and stay visible on touch where hover never
fires. Built-in rows get the override entry point and no delete.
The store's create/update/override paths now share runAgentMutation
instead of a third hand-copy of the out-of-order guards; the existing
store tests pass unchanged.
restoreAvailable conflated two questions: whether the files this checkpoint
reports can be put back, and whether the checkpoint saw every file the turn
touched. Any tool outside an 18-name allowlist — Bash, PowerShell, TaskCreate,
every MCP tool — forced the second to false, and one such call anywhere from
the target turn onward disabled undo for the whole range. In practice that is
every real turn, so v0.5.3 shipped with undo effectively dead.
Split the two. restoreAvailable now answers only the first question.
The second becomes unverifiedChangeSources: tool names whose file effects the
checkpoint could not capture. Undo stays available and restores exactly the
files it lists, and the card, the confirmation, and the completion toast each
name what it is leaving behind. An unrecognized tool now costs a warning
instead of the feature. A transcript that cannot be read still blocks, because
then even the reported file list may be wrong.
Bash calls are classified against the existing read-only allowlist via a new
recordedCommandIsReadOnly, which drops the sandbox/cwd checks that describe the
live process rather than the replayed session. `git status` no longer warns at
all, so the warning means something when it appears.
Rewind also takes a mode. `conversation` skips the file restore entirely, so a
turn whose files cannot be restored no longer costs the user the ability to
back out of the prompt — matching how upstream keeps Restore conversation
independent of Restore code.
Files written by shell commands were never recoverable here; the checkpoint
only ever covered the structured file tools. Reporting that is honest, and
docs now say so.
Since v0.5.1 every IM channel listed only the default project. All five
adapters passed the default work dir to AdapterHttpClient as the sole
allowed project root, so listRecentProjects filtered out everything else;
matchProject, listSessions, sessionExists, createSession and listSkills
were clamped the same way. Feishu is where it was reported, but telegram,
wechat, dingtalk and whatsapp were identical.
defaultWorkDir is documented as where a new IM session starts, not as an
access boundary. Using it as the boundary failed both ways: configured, it
hid every other project; blank, it falls back to PWD/cwd(), which is "/"
for a GUI-launched sidecar, so the boundary allowed the whole filesystem.
Split the two concepts. allowedProjectRoots is now its own setting (global,
per-platform, or ADAPTER_ALLOWED_PROJECT_ROOTS), resolved together with the
work dir by resolveAdapterWorkspace so the default project is always inside
the boundary and /new cannot fail on inconsistent config. The default is the
home directory; it refuses to inherit "/" or any ancestor of home. Pairing
remains the primary authorization control, so unusable roots warn and fall
back rather than locking the bot out.
All five entrypoints now build their client through createAdapterClient
instead of repeating the wiring, which is what let one defect appear in five
places at once.
Known gap, left for a follow-up: a project outside the boundary is still
reported as "not found" rather than "outside the allowed directories".
The site had drifted from the product. Every screenshot predated the
v0.5.0 UI redesign, the reading experience shipped no search and no
syntax highlighting, and a third of the pages were internal process
artefacts — migration task lists addressed to agentic workers, a
release runbook, a proposal marked "historical".
Reorganise around the only two people who read this: someone getting
the desktop app running for the first time, and someone reading the
source. Five sections replace nine — start / desktop / im / cli /
internals — and the pages that served neither reader are gone.
Site rewrite:
- Palette lifted from the desktop app's 「纸·墨·印」 themes, so the
site and the product read as one thing. Light mirrors 纯白, dark
mirrors 墨夜, and dark mode exists at all now.
- Fonts are self-hosted. The old @import from Google Fonts is
unreachable from mainland China, which left every heading in a
fallback serif; it also only requested weight 600 while the CSS
asked for 900, so Latin and CJK in the same heading disagreed.
- Docs were shipped as one 968KB manifest downloaded on every page
view. Split into a 32KB index plus one lazily imported chunk per
page; the entry bundle is now 101KB gzipped.
- Add search, syntax highlighting, per-route meta with canonical and
hreflang, a sitemap, and an error boundary. Replace the 44vh
mobile sidebar with a drawer.
- Image dimensions are read at build time and written into the tag,
so lazy images reserve their space instead of collapsing.
Screenshots are recaptured from a real v0.5.0 build against a clean
demo project, with tokens, QR codes and paired accounts redacted.
The previous set is deleted rather than kept alongside.
Routes follow file paths, so the restructure would have broken every
inbound link; 37 old paths redirect, in both languages. The PR policy
gate and CODEOWNERS also hardcoded docs/guide/contributing.md.
Verified: check:docs 78 pages / 323 links / 0 problems, check:policy
127 pass. Walked every route at 1440 and 390 in both themes for
overflow, contrast, keyboard reachability and focus management.
Importing an animated pet required a file that was exactly 1536x2288, laid
out as 88 seamless cells, with the last two rows holding sixteen distinct
gaze angles. No image model emits that. Whatever a user got back from Jimeng
or ChatGPT was some fixed size like 1024x1536, so the path ended at "the
animation atlas must be exactly 1536x2288 pixels" every time. The third card
was worse: "AI-generate full animation" was hardcoded `disabled`, so the one
entry point named after what people actually wanted to do was dead.
The fix was already in the tree. `scripts/assemble-generated-pet-atlas.py`
landed in the same commit as the four built-in pets, which is to say the
built-ins were produced this way — it takes an action sheet at any size,
slices it on an 8x9 grid, fits each cell to 192x208, mirrors the run row to
make run-left, and reuses rows to reach eleven. That capability was never
wired to anything a user could reach.
`petAtlasNormalize.ts` reimplements it on a canvas in the renderer, so an
author draws nine rows and the app derives the rest. Verified against the
reference assembler by reversing dada-code's atlas into a nine-row sheet and
re-normalizing it: every difference lands on semi-transparent antialiased
edges (2314 pixels, max channel delta 14/255) and opaque regions are
identical. That residue is canvas premultiplied-alpha round-tripping, not a
slicing bug.
Three contract details worth stating. Row frame counts are now derived from
`PET_ANIMATION_DEFINITIONS` rather than typed out a fourth time; they come
out equal to the assembler's `(6,8,8,4,5,8,6,6,6,8,8)`. A sheet already at
1536x2288 passes through byte-for-byte instead of being resliced, because
resampling finished artwork buys nothing. And since the validator never
inspects the alpha channel, a flattened white background used to import
happily and render as a rectangle on the desktop — the renderer now rejects
sheets whose atlas is under 5% transparent (the built-ins sit near 78%) with
a message that names the actual problem.
The copy stops describing the implementation. "Animate one image" and
"Import professional animation atlas / exact 1536x2288 v2 PNG" become "use a
picture you already have" and "I already have an action sheet"; the dead AI
card becomes a three-step walkthrough carrying a copyable prompt, a labelled
8x9 reference grid that can be saved locally, and the checks that catch the
common failures. Reference images are generated by a script rather than hand-
placed, in both languages. All five locales move together.
Caught while reviewing the real dialog in Electron: after finishing the
walkthrough the form heading fell through to the atlas branch and announced
"I already have an action sheet" to someone who had just been walked through
drawing one. Covered by a test now.
Not done: docs/images/desktop_ui/15_pet_create_methods.png still shows the
old dialog and needs a fresh capture from a running app to match the styling
of the shots around it.
Prepare the v0.4.6 desktop release note, bump the desktop package version, and refresh README/docs guidance for signed releases and updater validation.
Tested: bun run scripts/release.ts 0.4.6 --dry
Tested: bun test scripts/pr/release-workflow.test.ts scripts/release-update-metadata.test.ts scripts/quality-gate/package-smoke/index.test.ts
Tested: bun run check:policy
Tested: bun run check:docs
Not-tested: bun run verify; release prep was validated with docs and release-focused gates only.
Confidence: high
Scope-risk: narrow
Tested: bun test scripts/pr/release-workflow.test.ts scripts/release-update-metadata.test.ts scripts/quality-gate/package-smoke/index.test.ts
Tested: bun run check:policy
Tested: bun run check:docs
Tested: workflow YAML parse and git diff --check
Scope-risk: moderate
Real-world testing showed the unsigned app just needs the original 0.3.2
one-liner after the "damaged" prompt (System Settings "Open Anyway" also
works). Drop the Sequoia / "don't trash it" / clear-DMG-first / script
walkthrough noise and match the historical release-note style.
- release-notes/v0.4.0.md: macOS = drag in + `xattr -cr ...`, Windows one
line, Linux kept short.
- docs/desktop/04-installation.md: same trim, remove duplicated FAQ entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Rewrite release-notes/v0.4.0.md to match the historical style (no emoji,
Highlights/Fixes/Notes), keep it user-facing only (drop release-process
notes), and correct Linux (supported since the Tauri builds, not new).
- README.md / README.en.md: replace Tauri references with Electron, list
macOS / Windows / Linux, and point first-launch approval to the guide.
- Rewrite docs/desktop/04-installation.md for Electron: Electron asset
names, Linux section, and the unsigned-macOS flow (clear the DMG quarantine
before double-clicking; drop the WebView2/right-click-open leftovers).
- install-macos-unsigned.sh: keep the same-folder DMG flow, no online download.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Complete the Electron replacement boundary before merging by removing the renderer-side Tauri host fallback, tightening H5/browser access so only desktop navigation is tokenless, and moving desktop release publication to a tag-driven GitHub Actions matrix with a single final publish job.
Constraint: H5/browser capability access must not gain tokenless access through localhost or retired Tauri origins
Constraint: Desktop release artifacts must be built by GitHub Actions from version tags, not treated as local build outputs
Rejected: Keep localhost browser origins trusted for convenience | local browser contexts can access loopback services and must use the H5 token path
Rejected: Publish from each matrix job | partial releases can be created before all platforms finish
Confidence: high
Scope-risk: broad
Directive: Do not reintroduce Tauri origins or localhost browser origins into the trusted desktop origin set without a reviewed security design
Tested: bun test src/server/__tests__/h5-access-policy.test.ts src/server/__tests__/h5-access-auth.test.ts src/server/__tests__/diagnostics-service.test.ts src/server/middleware/cors.test.ts
Tested: bun test scripts/pr/release-workflow.test.ts scripts/release-update-metadata.test.ts
Tested: bun run check:desktop
Tested: bun run check:native
Tested: git diff --check
Not-tested: bun run check:server is blocked by expired quarantine entries server:cron-scheduler, server:providers-real, server:tasks, server:e2e:business-flow, server:e2e:full-flow
Introduce the Electron desktop shell alongside the existing React renderer and local Bun server boundary. The migration keeps the DesktopHost contract explicit across Tauri, Electron, and browser runtimes while adding Electron main/preload services for dialogs, shell, notifications, updates, tray/window lifecycle, terminal, preview WebContentsView, app mode, and release/package validation.
The commit also carries the latest local main desktop command updates, including agent slash entries and hidden-by-default markdown thinking details, so the packaged Electron build matches the current main UX surface.
Constraint: React renderer, local Bun server, REST/WebSocket, and sidecar boundaries must remain reusable during the migration
Constraint: macOS dev packages are ad-hoc signed and cannot prove Developer ID notarization or Gatekeeper release launch
Rejected: Browser-only smoke validation | it cannot exercise native dialogs, keychain prompts, notification behavior, or packaged app startup
Confidence: medium
Scope-risk: broad
Directive: Do not remove Tauri host support until signed Electron release artifacts pass native OS smoke on macOS, Windows, and Linux
Tested: bun run check:desktop
Tested: cd desktop && bun run check:electron
Tested: CSC_IDENTITY_AUTO_DISCOVERY=false bun run electron:package:dir
Tested: bun run test:package-smoke --platform macos --package-kind dir --artifacts-dir desktop/build-artifacts/electron
Tested: Computer Use read packaged Electron app window at desktop/build-artifacts/electron/mac-arm64/Claude Code Haha.app
Not-tested: Developer ID signed/notarized Gatekeeper launch
Not-tested: Real OS notification click-to-session action
Not-tested: Windows and Linux packaged app smoke on real hosts
The desktop docs now explain the intended H5 setup path for personal and team use: enable H5 in Settings, generate a one-time token, configure allowed origins, and use LAN or a reverse proxy to open the mobile browser chat surface.
Constraint: H5 is opt-in browser access, not a public multi-tenant auth system.
Rejected: Hide the setup details in the implementation spec only | users need operational guidance for token handling, CORS, and reverse proxy routing.
Confidence: high
Scope-risk: narrow
Directive: Keep this page aligned with Settings labels and the token/CORS behavior before documenting broader public hosting.
Tested: bun run check:docs
Tested: Live local smoke with temporary HOME: /health 200, H5 verify without token 401, H5 verify with token 200, configured Origin echoed by CORS.
- security: XSS sanitization with DOMPurify in Markdown/Mermaid/PermissionDialog;
path whitelist in filesystem API; fake keys in test/config files
- perf: fine-grained Zustand selectors in Sidebar/StatusBar/ContentRouter;
50ms throttle on streaming deltas; React.memo + useMemo in MessageList;
useRef for frequent keyboard shortcut state; AbortController 30s timeout
- leaks: WS session TTL timers (5-min cleanup on close); batch splice for
sdkMessages/stderrLines; folderPath validation in cronScheduler
- quality: optimistic update rollback in settingsStore; error state in
providerStore/teamStore; i18n for all hardcoded English strings
- docs: desktop architecture and features docs updated; VitePress nav fixed
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>