main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
Windows has no equivalent of `CGEvent.postToPid`. `pyautogui` bottoms out in
`SendInput`, which injects into the one system-wide input stream and warps the
one real cursor — so on Windows the agent shares the mouse and keyboard with
the user, and `SendInput` reports success unconditionally whether or not
anything acted on the events.
That combination produced the same lie the macOS engine was just fixed for:
click a point behind another window and it lands on that window; click one
off-screen and it lands nowhere; either way the helper answered "Action
completed".
So the helper now refuses instead of guessing:
* `ForegroundLease` samples `GetLastInputInfo` around every mutating command.
Interference before the action is `user_interference` — nothing ran, retry
is safe. Interference during it is `user_interference_result_unknown`,
because injection already went out and a retry could double-apply it. On a
play/pause toggle those two differ by exactly one wrong outcome.
* `ensure_point_on_screen` and `ensure_target_window_reachable` reject
coordinates outside every display and targets whose windows are minimized
or hidden, before anything is sent.
`GetLastInputInfo` is the signal because it needs no privileges and is not
advanced by `SendInput`, so the agent cannot trip its own detector. Both guards
fail open on an unreadable reading: a safety layer that turns the feature off
is not safety.
The guard set lives in one place rather than in each dispatcher branch — an
eleventh verb wired like the ten before it would otherwise be silently
unguarded, which is the bug class this pass removes. All ten mutating branches
now return through `_finish`, so no branch can write its own success response
and skip the post-action check.
Also adds a Windows cursor badge. It deliberately does NOT mirror the macOS
virtual cursor: there the real pointer never moves, so the drawn one is the
only cursor and replaces it. Here the real pointer does move, and a second
fake pointer would just be two cursors with one of them lying about where the
click lands. The badge annotates instead — it answers "is this me or the
agent?", which matters because grabbing the mouse mid-action is what makes the
two input streams interleave.
Retires `runtime/mac_helper.py` and its pyobjc requirements: macOS routes every
command to the signed native daemon and `helperBridge` refuses to fall back, so
both were unreachable. Renames the `callPythonHelper` alias to `callHelper`,
which is what it has actually imported since the native engine landed.
Verified with mutation testing — 13 injected regressions across both languages
(dropped lease, bypassed `_finish`, collapsed interference codes, guard reusing
the filtered `list_windows`, badge losing click-through), all caught.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
Three failures shared one root cause: the engine reported success for input it
had no way to deliver or verify.
A fully covered Chromium/CEF window stops drawing, so every screenshot returns
the last frame it painted while each action still answers "Action completed".
One session read a frozen image for three and a half minutes, pressed play four
times, and reported a song playing that the final capture showed paused.
Mutating commands now fail closed with `window_occluded` instead.
The physical-input monitor counted `.mouseMoved`, so moving the cursor anywhere
on screen aborted background automation — the one thing the feature exists to
allow. Movement is not an interaction with any app; presses, drags, keys and
scroll still are.
Focus notifications went out on NSEvent type 13 alone. The target accepted them,
ignored them, and reported nothing: 24 actions, 1 effective. Key-focus subtypes
travel on type 21, and `keyFocusReturned` (0x8000) needs a sign-preserving
truncation or it collapses to subtype 0 — a notification the target accepts and
discards.
The stale-capture notice claimed input does not depend on visibility. Measured
behaviour contradicts it, so it now says actions are refused while covered, and
says it without asking the model to consult the user about window management —
an earlier wording did, and the model handed back a three-step task after one
action.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
ImageEdit trusted only two per-session directories: the bridge upload dir
and its own generated-images dir. Neither is where a user's image actually
lands. An @-mentioned file keeps its original path anywhere on disk, and a
pasted or dropped image goes to ~/.claude/image-cache/<sessionId>/, so every
image a user supplied was refused.
The tool description made it worse by pointing the model at
[Image source: ...] paths, which are exactly the image-cache paths the check
then rejected. The model followed the description, got refused, and asked
users to re-attach a file they had just attached.
Record images the user names explicitly in a session-scoped registry, and
trust the paste directory by location. Paths the model found on its own — a
Glob hit, a path read out of a file — stay refused, which is what the
original root-dir check was protecting against.
Claude-Session: https://claude.ai/code/session_01ArehnJt4QLb83xkEukqNFX
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.
`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.
What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.
Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.
The rest of the menu:
- The system-default row is named after the application it will actually open,
but still launches through the system-default target. Only that path carries
the guard that refuses to hand an executable to the shell.
- A file type nothing can edit no longer offers an editor. Reusing the
workspace preview gate rather than writing a third extension table.
- One editor is guaranteed a slot and the rest fill what the applications
leave. On Windows and Linux there is no application list, so a hard cap of
one would have dropped installed editors for nothing.
- The file manager is named per platform instead of interpolating the server's
label, which produced "在 Explorer 中显示" — an English name inside a Chinese
sentence, and not what Windows calls it. Same shape as #1236.
- `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
scroll inside the overlay as a viewport change: the listener is on capture,
so the menu closed the instant the user reached for its own scrollbar.
Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.
And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.
Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
Requesting worktree isolation in a workspace with no git repository and
no WorktreeCreate hook failed every subagent it was asked for, so a
whole workflow run came back with nothing. Isolation is best-effort, not
a precondition for running: the agent now runs in the workspace
directory instead, and the skip is reported rather than hidden — once
per run in the workflow log, and in the Agent tool's result.
Explicit entry points that ask for a worktree (EnterWorktree, --worktree,
bridge worktree mode) still fail loudly. Deliberate divergence from
upstream, which hard-fails this path in both the Agent tool and the
workflow harness.
Materialize the question card directly from a permission request when the
streamed tool block has not arrived yet. Upsert by tool-use id so either event
order produces one visible, answerable card without requiring a refresh.
Remove model-bound thinking and redacted thinking blocks from the in-memory
request history when a resumed session changes models. Preserve the persisted
transcript and leave same-model histories untouched.
Two stacked issues kept grok-4.6 requests at high effort even when the
UI selected xhigh: the effort capability table had no Grok entries, so
resolveAppliedEffort clamped xhigh to high; and an explicit --effort
flag never overrode CLAUDE_CODE_EFFORT_LEVEL env for main-loop queries.
Resolve Grok capabilities from the bundled Grok catalog and wire
effortValueOverridesEnv through the REPL for explicit CLI effort.
normalizeRuntimeSelection dropped the effort level whenever the selected
Grok model was missing from the bundled desktop catalog (e.g. grok-4.6),
so the reasoning effort slider snapped back to the default on every drag.
Preserve effort for live-catalog models and let the server validate it,
and sync the bundled desktop catalog with grok-4.6 as the new default.
Add grok-4.6 to the bundled fallback catalog with specs from the live
/v1/models endpoint (500k context, high default effort, xhigh..low
effort levels) and make it the default main model for the Grok Official
provider.
The workbench described a team run in ways the underlying data did not
support, so a member marked "已完成" led to an agent that was still
streaming, and the feed reported no communication for a team that had
just handed out all of its work.
A teammate rewrites its entire transcript every turn, leaving a chain of
`agent-<id>.jsonl` fragments where each repeats its predecessor. The
reader concatenated all of them, and `fragmentScopedId` gave the same
entry a different id per fragment, so id-based deduplication downstream
could not collapse them: one member replayed its work up to eight times
and kept growing. Drop a fragment whose entries are a strict prefix of
another's, ignoring only the four fields a rewrite restamps (agentId,
slug, cwd, promptId). Independent resumes reuse entry ids for different
work, so identity comes from content and those fragments all survive.
The activity path scoped tool ids per fragment before deduplicating,
which split a call from its result across two scopes; it now folds
first.
Member activity came from owning an `in_progress` task. A teammate marks
a task started and can then end its turn, and an umbrella task stays open
across every turn beneath it, so every member read as permanently
working. Report the runner's own turn markers as a `TeamMemberActivity`
separate from `status`, falling back to a transcript-write probe only for
backends that record no markers, and let the workbench say `unknown`
rather than guess.
The remaining corrections follow from the same principle -- show what the
data says:
- A member is drawn once, on the task it most recently started, instead
of on every task it owns; other cards carry an owner chip.
- Opening a task card describes the task; reaching its owner is now a
deliberate second click.
- Cards lead with dependency depth. A task id records the order the lead
wrote tasks down, so review work planned first showed "#1" while
hanging off the bottom of the graph.
- Task handovers are communication, not lifecycle noise, and self-claims
read differently from assignments. Repeated idle notices fold into a
count.
- A blocker missing from the task list counts as resolved, matching
`claimTask`, instead of stranding its dependents in `blocked`.
- Batch-completed tasks record their owner. Automatic ownership only
triggered on `in_progress`, so work closed without ever starting had
none; announcing an already-finished task is skipped so an idle
teammate is not woken for it.
- Transcript pages expose the `TaskUpdate` calls that bound each task,
so a member conversation can be read as the tasks it worked through.
Atlas Cloud now sponsors the project, so give it the same treatment as
the other sponsors: mark the preset featured to move it into the sponsor
row, and add the cc-haha campaign link behind the "get API key" button.
Reorder the preset to sit last in that row, matching the README.
Add Atlas Cloud to the sponsor tables in both README.md and
README.zh-CN.md, using a cc-haha-specific campaign link. Ship
light/dark logo variants so the wordmark stays visible in both
GitHub themes.