main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.
`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.
What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.
Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.
The rest of the menu:
- The system-default row is named after the application it will actually open,
but still launches through the system-default target. Only that path carries
the guard that refuses to hand an executable to the shell.
- A file type nothing can edit no longer offers an editor. Reusing the
workspace preview gate rather than writing a third extension table.
- One editor is guaranteed a slot and the rest fill what the applications
leave. On Windows and Linux there is no application list, so a hard cap of
one would have dropped installed editors for nothing.
- The file manager is named per platform instead of interpolating the server's
label, which produced "在 Explorer 中显示" — an English name inside a Chinese
sentence, and not what Windows calls it. Same shape as #1236.
- `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
scroll inside the overlay as a viewport change: the listener is on capture,
so the menu closed the instant the user reached for its own scrollbar.
Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.
And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.
Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
Materialize the question card directly from a permission request when the
streamed tool block has not arrived yet. Upsert by tool-use id so either event
order produces one visible, answerable card without requiring a refresh.
normalizeRuntimeSelection dropped the effort level whenever the selected
Grok model was missing from the bundled desktop catalog (e.g. grok-4.6),
so the reasoning effort slider snapped back to the default on every drag.
Preserve effort for live-catalog models and let the server validate it,
and sync the bundled desktop catalog with grok-4.6 as the new default.
The workbench described a team run in ways the underlying data did not
support, so a member marked "已完成" led to an agent that was still
streaming, and the feed reported no communication for a team that had
just handed out all of its work.
A teammate rewrites its entire transcript every turn, leaving a chain of
`agent-<id>.jsonl` fragments where each repeats its predecessor. The
reader concatenated all of them, and `fragmentScopedId` gave the same
entry a different id per fragment, so id-based deduplication downstream
could not collapse them: one member replayed its work up to eight times
and kept growing. Drop a fragment whose entries are a strict prefix of
another's, ignoring only the four fields a rewrite restamps (agentId,
slug, cwd, promptId). Independent resumes reuse entry ids for different
work, so identity comes from content and those fragments all survive.
The activity path scoped tool ids per fragment before deduplicating,
which split a call from its result across two scopes; it now folds
first.
Member activity came from owning an `in_progress` task. A teammate marks
a task started and can then end its turn, and an umbrella task stays open
across every turn beneath it, so every member read as permanently
working. Report the runner's own turn markers as a `TeamMemberActivity`
separate from `status`, falling back to a transcript-write probe only for
backends that record no markers, and let the workbench say `unknown`
rather than guess.
The remaining corrections follow from the same principle -- show what the
data says:
- A member is drawn once, on the task it most recently started, instead
of on every task it owns; other cards carry an owner chip.
- Opening a task card describes the task; reaching its owner is now a
deliberate second click.
- Cards lead with dependency depth. A task id records the order the lead
wrote tasks down, so review work planned first showed "#1" while
hanging off the bottom of the graph.
- Task handovers are communication, not lifecycle noise, and self-claims
read differently from assignments. Repeated idle notices fold into a
count.
- A blocker missing from the task list counts as resolved, matching
`claimTask`, instead of stranding its dependents in `blocked`.
- Batch-completed tasks record their owner. Automatic ownership only
triggered on `in_progress`, so work closed without ever starting had
none; announcing an already-finished task is skipped so an idle
teammate is not woken for it.
- Transcript pages expose the `TaskUpdate` calls that bound each task,
so a member conversation can be read as the tasks it worked through.
One conflict, in SubagentRunPage.test.tsx: both sides rewrote the
`../api/subagents` mock. main added hoisted teammate mocks for the new
TeamMemberRunPage tests; this branch made the factory spread the real module
so the agent-id ref helpers stay real rather than being stubbed into a
fiction. Both are kept — the helpers decide which endpoint the page calls,
and the teammate tests need their handles.
A workflow is a JS script the model writes in the moment and hands to the
Workflow tool, which runs it in a locked-down `node:vm` and orchestrates
subagents through `agent()`/`parallel()`/`pipeline()`/`phase()`. Saving one
as a `/name` command is the secondary path; the inline script is the point.
Runtime: cross-realm value marshalling so a script cannot reach the host
`Function`, determinism guards on `Date.now()`/`Math.random()` (they would
make a resume replay diverge), a FIFO concurrency gate, and a journal that
lets an interrupted run resume from its longest unchanged prefix.
Desktop: the run shows up as a `workflow` section in the existing activity
panel — phases as headings, their agents beneath. A workflow agent is an
ordinary subagent run by the same runner, so its row opens the existing
subagent page rather than a parallel viewer of its own; that needed a
`by-agent` lookup, because these agents have no parent `Agent` tool call to
key off. Finished runs are rebuilt from the per-agent sidecars when a
session is reopened, since the live progress stream does not outlive the
process.
Opening a running subagent's record showed its prompt and nothing else,
while the parent transcript streamed that same agent's tool calls. The
detail route resolves an agent id three ways — a regex over the parent's
tool_result, a task id passed by the client, and a task notification —
and all three only exist once the agent has finished, so a synchronously
dispatched subagent had no transcript to read for its entire run. The
sidecar the CLI writes before the query loop starts already carries the
spawning tool_use id; resolving through that first is the only hint that
exists while the run is still in flight.
Rewinding a real session to the moment that agent was still running: no
agent id and 0 messages before, 38 messages and 14 tool calls after.
That page also offered a composer for every agent. A one-shot subagent
has no inbox — send_agent_message falls through to resumeAgentBackground
and forks a detached copy whose output has nowhere to land. The response
now reports whether an inbox exists, and only named teammates and
in-flight background agents get a composer.
The Agent Teams panel had the opposite problem: at docked width it tried
to be the whole workbench. The dependency map, the member drill-down and
a 210px message strip left no room for the summary that answers whether
the run worked; the seam between transcript and panel only appeared on
hover; and the header progress bar spanned the full width because
Progress is w-full internally and cx does not merge Tailwind classes. The
docked side is a run report now — roster, task list, outcome — while the
map, feed, member views and history scrubbing stay in the workbench tab
that already existed.
Dropped the toolbar's team toggle. It predates AgentTeamsStrip, renders
under exactly the same condition, and does the same thing with less
context.
Desktop suite over the touched areas: 249 passed. Server-side subagent
tests: 30 passed. Lint and tsc clean.
Two conflicts, both adjacent inserts where each side kept its own line:
- MessageList.tsx: TEAM_CARD_HEIGHT landed next to the recalibrated
ACTIVITY_GROUP_COLLAPSED_HEIGHT. Kept both, and restated the team card
at 86 — 78 was measured against the pre-refactor box, which did not yet
carry the turn's 8px padding-bottom.
- MessageList.test.tsx: both sides appended a describe block at EOF.
Follow-ups the merge needed but could not produce on its own:
- UserMessage: the new teammate bubble kept an `mb-5` that the transcript
no longer uses. `.chat-turn-rail` establishes a block formatting
context, so that margin would have been trapped inside the measured box
and made every teammate row 20px taller than a prompt.
- globals.css: the note about containers spacing transcript blocks
themselves was written when SubagentRunPage assembled its own render
model. It renders MessageList now, as does AgentTeamsMemberView, so
both inherit the spacing.
- A rail case for `team_card`: it is the one render item that is neither a
message nor a tool group, so it is the one that could fall out of the
turn walk unnoticed.
- SubagentRunPage's live-refresh test unfolded the activity summary before
reaching for a row. A running run plays open now, so that click folded
the rows away instead. Assert it is already open and go straight to the
row; what the test guards — expansion surviving a refresh — is unchanged.
Local desktop suite: 304 files, 4226 passed. Build and desktop-ui-smoke
both green.
The transcript put every assistant reply in its own bordered card, the
same border and radius the tool-activity bar used, so prose and machinery
looked identical and a scrolled session read as a stack of boxes. A
one-line reply cost 112px to show 22px of text.
Alignment and width already say who is speaking — the prompt hugs its
text on the right, the reply takes the full column on the left — so the
border was repeating that at the cost of the space. Drop it, and separate
the layers by tone instead: prose stays primary, tool rows go 12px
tertiary.
Activity now plays open while its turn is still producing into it and
folds into a counted digest once that turn moves on:
thought 5 times, read 2 files, ran 2 commands 8.0s
Counts are what make the line worth having; the older "ran some commands"
summary said strictly less than the rows it hid. Clicking pins the
reader's choice so a run they opened is not closed under them later.
Also:
- Bash rows show the model's own description of a command rather than the
raw pipeline; the command itself moves to the terminal that ran it.
- Rows gain a leading icon, reusing activitySegmentIcon.
- ThinkingBlock loses its second, larger form. A thought that happened not
to be followed by a tool call rendered as a bare label with none of its
content — the least informative thing on screen, for a reason that was
never about the thought. Its preview follows the tail while streaming.
- Message actions render only on the reply that closes a turn. The bar
reserved 36px on every reply whether or not it was hovered, which on a
one-line reply outweighed the text. Mid-turn text stays copyable by
selection.
- Completed background tasks stop drawing a card; the activity panel
already lists them, and a team session emits dozens. Failures still
interrupt.
- Virtualizer estimates are recalibrated to the new box model, and item
height is measured from the border box so the turn gap is counted.
VIRTUAL_MIN_ITEM_HEIGHT clamps measurements too, so it had to drop below
the shortest real row or those rows were recorded too tall forever.
27c20389a made a nested tool_use_complete keep the parent's chatState
rather than dropping the session back to `thinking`, so a subagent's
inner tool call no longer clears the parent Task's progress indicator.
That was the intended fix, but the golden fixture still pinned the old
`thinking` value and the two subagent scenarios have failed since —
masked until now because localStorage was taking the whole file down
first.
Regenerating touches exactly one line, and produces an identical file
from a clean HEAD, confirming this is fixture staleness rather than a
behaviour change of ours.
Local desktop suite is now fully green: 304 files, 4202 passed.