Track real and synthetic focus with process-bound monitor receipts.
Bind keyboard activation to the snapshot window and align targeted
pointer and keyboard bursts with the reference event sequence.
Wait for pasteboard data delivery before restoring the clipboard,
and add focus/stream diagnostics with transition regression coverage.
Validated with check:native and six covered NetEase focus cycles.
A traced session already recorded everything needed to explain a model
request, but the presentation kept it out of reach. The system prompt lived
inside each individual call, so reading it meant picking a call first. The
context the harness assembles — CLAUDE.md, system-reminder blocks,
deferred-tool rosters — was indistinguishable from what the person typed,
because the provider receives all of it as user-role text.
- The session overview names the system prompt and tool catalog, read once
from the first model call rather than hunted down per request. This is the
opening header, not a session-wide invariant: late tool registration and a
mid-session model change rewrite it for later requests, which keep their own
header on their own detail.
- A model call's detail separates injected context from the exchange, one row
per injection labelled by its own content instead of its wrapper tag.
- Tool rows carry their input summary and model rows their token counts, so
scanning the tree distinguishes one call from the next.
- An assistant turn that only reasoned says so, including when the provider
withheld the reasoning, instead of rendering as a bare label.
Recognition of injected context is by wrapper tag against a closed list.
Position cannot stand in for it: hoistToolResults moves every tool result to
the front of a merged user message, so an attachment the harness appended and
an instruction typed after interrupting a tool arrive in the same shape. The
ambiguity is resolved toward the conversation — unrecognized text stays the
person's message, because looking for what you said and not finding it is the
worse failure.
requestParse also learns the OpenAI Responses wire format, which about a
quarter of traced sessions use and which previously yielded an empty message
list and no system prompt. Responses splits a tool round trip into sibling
function_call / function_call_output entries; across 150 real trace files
those outnumber plain messages 1199 to 878, so they are mapped onto the
tool_use / tool_result vocabulary rather than skipped. Its flat tool
`parameters` and `instructions` spellings are read too, the latter with `||`
so an empty `system` array cannot shadow it.
All new surfaces use design tokens, so the six themes need no per-theme work.
Claude-Session: https://claude.ai/code/session_0119s8U5VzUvVpNgWA7B3pSg
The sidebar could only group sessions by workspace. With many workspaces
open that is orthogonal to the question people actually ask — "which of
the tasks I just started are done?" — so a finished session had to be
hunted for across a dozen collapsed project groups.
A bell in the sidebar header now switches the list to a flat task view:
running sessions first, then sessions bucketed by calendar day (today,
yesterday, previous 7/30 days, earlier). Each row carries its title and
its owning directory, which is exactly the pairing the project grouping
folds away.
The bell shares one state with the existing "organize sidebar -> by
time" menu item, which promised this and still grouped by workspace.
That reuses the `projectOrganization` preference already persisted to
localStorage and to the desktop UI preferences service, so no new
storage key and no migration.
A session stopped on a permission prompt shows a warning dot rather than
the running spinner: it counts as unfinished, but it is waiting on the
user, not working.
Buckets use calendar-day boundaries, not elapsed hours — a task finished
at 23:50 belongs to "yesterday" when read at 00:10, and a mutation to
elapsed-hours arithmetic turns the guard test red.
Verified: check:desktop green (318 files, 4617 tests, lint + tsc +
build); the flat view, bell round trip and preference flip exercised in
a real browser against a sandboxed server with seeded sessions;
group/title/workspace contrast measured >= 4.83:1 across all six themes;
no horizontal overflow at the 240px minimum sidebar width.
Claude-Session: https://claude.ai/code/session_018mbioW3EEr7KbrWVgBvUvJ
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
Windows has no equivalent of `CGEvent.postToPid`. `pyautogui` bottoms out in
`SendInput`, which injects into the one system-wide input stream and warps the
one real cursor — so on Windows the agent shares the mouse and keyboard with
the user, and `SendInput` reports success unconditionally whether or not
anything acted on the events.
That combination produced the same lie the macOS engine was just fixed for:
click a point behind another window and it lands on that window; click one
off-screen and it lands nowhere; either way the helper answered "Action
completed".
So the helper now refuses instead of guessing:
* `ForegroundLease` samples `GetLastInputInfo` around every mutating command.
Interference before the action is `user_interference` — nothing ran, retry
is safe. Interference during it is `user_interference_result_unknown`,
because injection already went out and a retry could double-apply it. On a
play/pause toggle those two differ by exactly one wrong outcome.
* `ensure_point_on_screen` and `ensure_target_window_reachable` reject
coordinates outside every display and targets whose windows are minimized
or hidden, before anything is sent.
`GetLastInputInfo` is the signal because it needs no privileges and is not
advanced by `SendInput`, so the agent cannot trip its own detector. Both guards
fail open on an unreadable reading: a safety layer that turns the feature off
is not safety.
The guard set lives in one place rather than in each dispatcher branch — an
eleventh verb wired like the ten before it would otherwise be silently
unguarded, which is the bug class this pass removes. All ten mutating branches
now return through `_finish`, so no branch can write its own success response
and skip the post-action check.
Also adds a Windows cursor badge. It deliberately does NOT mirror the macOS
virtual cursor: there the real pointer never moves, so the drawn one is the
only cursor and replaces it. Here the real pointer does move, and a second
fake pointer would just be two cursors with one of them lying about where the
click lands. The badge annotates instead — it answers "is this me or the
agent?", which matters because grabbing the mouse mid-action is what makes the
two input streams interleave.
Retires `runtime/mac_helper.py` and its pyobjc requirements: macOS routes every
command to the signed native daemon and `helperBridge` refuses to fall back, so
both were unreachable. Renames the `callPythonHelper` alias to `callHelper`,
which is what it has actually imported since the native engine landed.
Verified with mutation testing — 13 injected regressions across both languages
(dropped lease, bypassed `_finish`, collapsed interference codes, guard reusing
the filtered `list_windows`, badge losing click-through), all caught.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
Three failures shared one root cause: the engine reported success for input it
had no way to deliver or verify.
A fully covered Chromium/CEF window stops drawing, so every screenshot returns
the last frame it painted while each action still answers "Action completed".
One session read a frozen image for three and a half minutes, pressed play four
times, and reported a song playing that the final capture showed paused.
Mutating commands now fail closed with `window_occluded` instead.
The physical-input monitor counted `.mouseMoved`, so moving the cursor anywhere
on screen aborted background automation — the one thing the feature exists to
allow. Movement is not an interaction with any app; presses, drags, keys and
scroll still are.
Focus notifications went out on NSEvent type 13 alone. The target accepted them,
ignored them, and reported nothing: 24 actions, 1 effective. Key-focus subtypes
travel on type 21, and `keyFocusReturned` (0x8000) needs a sign-preserving
truncation or it collapses to subtype 0 — a notification the target accepts and
discards.
The stale-capture notice claimed input does not depend on visibility. Measured
behaviour contradicts it, so it now says actions are refused while covered, and
says it without asking the model to consult the user about window management —
an earlier wording did, and the model handed back a three-step task after one
action.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
ImageEdit trusted only two per-session directories: the bridge upload dir
and its own generated-images dir. Neither is where a user's image actually
lands. An @-mentioned file keeps its original path anywhere on disk, and a
pasted or dropped image goes to ~/.claude/image-cache/<sessionId>/, so every
image a user supplied was refused.
The tool description made it worse by pointing the model at
[Image source: ...] paths, which are exactly the image-cache paths the check
then rejected. The model followed the description, got refused, and asked
users to re-attach a file they had just attached.
Record images the user names explicitly in a session-scoped registry, and
trust the paste directory by location. Paths the model found on its own — a
Glob hit, a path read out of a file — stay refused, which is what the
original root-dir check was protecting against.
Claude-Session: https://claude.ai/code/session_01ArehnJt4QLb83xkEukqNFX
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.
`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.
What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.
Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.
The rest of the menu:
- The system-default row is named after the application it will actually open,
but still launches through the system-default target. Only that path carries
the guard that refuses to hand an executable to the shell.
- A file type nothing can edit no longer offers an editor. Reusing the
workspace preview gate rather than writing a third extension table.
- One editor is guaranteed a slot and the rest fill what the applications
leave. On Windows and Linux there is no application list, so a hard cap of
one would have dropped installed editors for nothing.
- The file manager is named per platform instead of interpolating the server's
label, which produced "在 Explorer 中显示" — an English name inside a Chinese
sentence, and not what Windows calls it. Same shape as #1236.
- `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
scroll inside the overlay as a viewport change: the listener is on capture,
so the menu closed the instant the user reached for its own scrollbar.
Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.
And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.
Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
Requesting worktree isolation in a workspace with no git repository and
no WorktreeCreate hook failed every subagent it was asked for, so a
whole workflow run came back with nothing. Isolation is best-effort, not
a precondition for running: the agent now runs in the workspace
directory instead, and the skip is reported rather than hidden — once
per run in the workflow log, and in the Agent tool's result.
Explicit entry points that ask for a worktree (EnterWorktree, --worktree,
bridge worktree mode) still fail loudly. Deliberate divergence from
upstream, which hard-fails this path in both the Agent tool and the
workflow harness.