A traced session already recorded everything needed to explain a model
request, but the presentation kept it out of reach. The system prompt lived
inside each individual call, so reading it meant picking a call first. The
context the harness assembles — CLAUDE.md, system-reminder blocks,
deferred-tool rosters — was indistinguishable from what the person typed,
because the provider receives all of it as user-role text.
- The session overview names the system prompt and tool catalog, read once
from the first model call rather than hunted down per request. This is the
opening header, not a session-wide invariant: late tool registration and a
mid-session model change rewrite it for later requests, which keep their own
header on their own detail.
- A model call's detail separates injected context from the exchange, one row
per injection labelled by its own content instead of its wrapper tag.
- Tool rows carry their input summary and model rows their token counts, so
scanning the tree distinguishes one call from the next.
- An assistant turn that only reasoned says so, including when the provider
withheld the reasoning, instead of rendering as a bare label.
Recognition of injected context is by wrapper tag against a closed list.
Position cannot stand in for it: hoistToolResults moves every tool result to
the front of a merged user message, so an attachment the harness appended and
an instruction typed after interrupting a tool arrive in the same shape. The
ambiguity is resolved toward the conversation — unrecognized text stays the
person's message, because looking for what you said and not finding it is the
worse failure.
requestParse also learns the OpenAI Responses wire format, which about a
quarter of traced sessions use and which previously yielded an empty message
list and no system prompt. Responses splits a tool round trip into sibling
function_call / function_call_output entries; across 150 real trace files
those outnumber plain messages 1199 to 878, so they are mapped onto the
tool_use / tool_result vocabulary rather than skipped. Its flat tool
`parameters` and `instructions` spellings are read too, the latter with `||`
so an empty `system` array cannot shadow it.
All new surfaces use design tokens, so the six themes need no per-theme work.
Claude-Session: https://claude.ai/code/session_0119s8U5VzUvVpNgWA7B3pSg
The sidebar could only group sessions by workspace. With many workspaces
open that is orthogonal to the question people actually ask — "which of
the tasks I just started are done?" — so a finished session had to be
hunted for across a dozen collapsed project groups.
A bell in the sidebar header now switches the list to a flat task view:
running sessions first, then sessions bucketed by calendar day (today,
yesterday, previous 7/30 days, earlier). Each row carries its title and
its owning directory, which is exactly the pairing the project grouping
folds away.
The bell shares one state with the existing "organize sidebar -> by
time" menu item, which promised this and still grouped by workspace.
That reuses the `projectOrganization` preference already persisted to
localStorage and to the desktop UI preferences service, so no new
storage key and no migration.
A session stopped on a permission prompt shows a warning dot rather than
the running spinner: it counts as unfinished, but it is waiting on the
user, not working.
Buckets use calendar-day boundaries, not elapsed hours — a task finished
at 23:50 belongs to "yesterday" when read at 00:10, and a mutation to
elapsed-hours arithmetic turns the guard test red.
Verified: check:desktop green (318 files, 4617 tests, lint + tsc +
build); the flat view, bell round trip and preference flip exercised in
a real browser against a sandboxed server with seeded sessions;
group/title/workspace contrast measured >= 4.83:1 across all six themes;
no horizontal overflow at the 240px minimum sidebar width.
Claude-Session: https://claude.ai/code/session_018mbioW3EEr7KbrWVgBvUvJ
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.
`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.
What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.
Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.
The rest of the menu:
- The system-default row is named after the application it will actually open,
but still launches through the system-default target. Only that path carries
the guard that refuses to hand an executable to the shell.
- A file type nothing can edit no longer offers an editor. Reusing the
workspace preview gate rather than writing a third extension table.
- One editor is guaranteed a slot and the rest fill what the applications
leave. On Windows and Linux there is no application list, so a hard cap of
one would have dropped installed editors for nothing.
- The file manager is named per platform instead of interpolating the server's
label, which produced "在 Explorer 中显示" — an English name inside a Chinese
sentence, and not what Windows calls it. Same shape as #1236.
- `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
scroll inside the overlay as a viewport change: the listener is on capture,
so the menu closed the instant the user reached for its own scrollbar.
Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.
And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.
Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
normalizeRuntimeSelection dropped the effort level whenever the selected
Grok model was missing from the bundled desktop catalog (e.g. grok-4.6),
so the reasoning effort slider snapped back to the default on every drag.
Preserve effort for live-catalog models and let the server validate it,
and sync the bundled desktop catalog with grok-4.6 as the new default.
The workbench described a team run in ways the underlying data did not
support, so a member marked "已完成" led to an agent that was still
streaming, and the feed reported no communication for a team that had
just handed out all of its work.
A teammate rewrites its entire transcript every turn, leaving a chain of
`agent-<id>.jsonl` fragments where each repeats its predecessor. The
reader concatenated all of them, and `fragmentScopedId` gave the same
entry a different id per fragment, so id-based deduplication downstream
could not collapse them: one member replayed its work up to eight times
and kept growing. Drop a fragment whose entries are a strict prefix of
another's, ignoring only the four fields a rewrite restamps (agentId,
slug, cwd, promptId). Independent resumes reuse entry ids for different
work, so identity comes from content and those fragments all survive.
The activity path scoped tool ids per fragment before deduplicating,
which split a call from its result across two scopes; it now folds
first.
Member activity came from owning an `in_progress` task. A teammate marks
a task started and can then end its turn, and an umbrella task stays open
across every turn beneath it, so every member read as permanently
working. Report the runner's own turn markers as a `TeamMemberActivity`
separate from `status`, falling back to a transcript-write probe only for
backends that record no markers, and let the workbench say `unknown`
rather than guess.
The remaining corrections follow from the same principle -- show what the
data says:
- A member is drawn once, on the task it most recently started, instead
of on every task it owns; other cards carry an owner chip.
- Opening a task card describes the task; reaching its owner is now a
deliberate second click.
- Cards lead with dependency depth. A task id records the order the lead
wrote tasks down, so review work planned first showed "#1" while
hanging off the bottom of the graph.
- Task handovers are communication, not lifecycle noise, and self-claims
read differently from assignments. Repeated idle notices fold into a
count.
- A blocker missing from the task list counts as resolved, matching
`claimTask`, instead of stranding its dependents in `blocked`.
- Batch-completed tasks record their owner. Automatic ownership only
triggered on `in_progress`, so work closed without ever starting had
none; announcing an already-finished task is skipped so an idle
teammate is not woken for it.
- Transcript pages expose the `TaskUpdate` calls that bound each task,
so a member conversation can be read as the tasks it worked through.
One conflict, in SubagentRunPage.test.tsx: both sides rewrote the
`../api/subagents` mock. main added hoisted teammate mocks for the new
TeamMemberRunPage tests; this branch made the factory spread the real module
so the agent-id ref helpers stay real rather than being stubbed into a
fiction. Both are kept — the helpers decide which endpoint the page calls,
and the teammate tests need their handles.
A workflow is a JS script the model writes in the moment and hands to the
Workflow tool, which runs it in a locked-down `node:vm` and orchestrates
subagents through `agent()`/`parallel()`/`pipeline()`/`phase()`. Saving one
as a `/name` command is the secondary path; the inline script is the point.
Runtime: cross-realm value marshalling so a script cannot reach the host
`Function`, determinism guards on `Date.now()`/`Math.random()` (they would
make a resume replay diverge), a FIFO concurrency gate, and a journal that
lets an interrupted run resume from its longest unchanged prefix.
Desktop: the run shows up as a `workflow` section in the existing activity
panel — phases as headings, their agents beneath. A workflow agent is an
ordinary subagent run by the same runner, so its row opens the existing
subagent page rather than a parallel viewer of its own; that needed a
`by-agent` lookup, because these agents have no parent `Agent` tool call to
key off. Finished runs are rebuilt from the per-agent sidecars when a
session is reopened, since the live progress stream does not outlive the
process.
Opening a running subagent's record showed its prompt and nothing else,
while the parent transcript streamed that same agent's tool calls. The
detail route resolves an agent id three ways — a regex over the parent's
tool_result, a task id passed by the client, and a task notification —
and all three only exist once the agent has finished, so a synchronously
dispatched subagent had no transcript to read for its entire run. The
sidecar the CLI writes before the query loop starts already carries the
spawning tool_use id; resolving through that first is the only hint that
exists while the run is still in flight.
Rewinding a real session to the moment that agent was still running: no
agent id and 0 messages before, 38 messages and 14 tool calls after.
That page also offered a composer for every agent. A one-shot subagent
has no inbox — send_agent_message falls through to resumeAgentBackground
and forks a detached copy whose output has nowhere to land. The response
now reports whether an inbox exists, and only named teammates and
in-flight background agents get a composer.
The Agent Teams panel had the opposite problem: at docked width it tried
to be the whole workbench. The dependency map, the member drill-down and
a 210px message strip left no room for the summary that answers whether
the run worked; the seam between transcript and panel only appeared on
hover; and the header progress bar spanned the full width because
Progress is w-full internally and cx does not merge Tailwind classes. The
docked side is a run report now — roster, task list, outcome — while the
map, feed, member views and history scrubbing stay in the workbench tab
that already existed.
Dropped the toolbar's team toggle. It predates AgentTeamsStrip, renders
under exactly the same condition, and does the same thing with less
context.
Desktop suite over the touched areas: 249 passed. Server-side subagent
tests: 30 passed. Lint and tsc clean.
Two conflicts, both adjacent inserts where each side kept its own line:
- MessageList.tsx: TEAM_CARD_HEIGHT landed next to the recalibrated
ACTIVITY_GROUP_COLLAPSED_HEIGHT. Kept both, and restated the team card
at 86 — 78 was measured against the pre-refactor box, which did not yet
carry the turn's 8px padding-bottom.
- MessageList.test.tsx: both sides appended a describe block at EOF.
Follow-ups the merge needed but could not produce on its own:
- UserMessage: the new teammate bubble kept an `mb-5` that the transcript
no longer uses. `.chat-turn-rail` establishes a block formatting
context, so that margin would have been trapped inside the measured box
and made every teammate row 20px taller than a prompt.
- globals.css: the note about containers spacing transcript blocks
themselves was written when SubagentRunPage assembled its own render
model. It renders MessageList now, as does AgentTeamsMemberView, so
both inherit the spacing.
- A rail case for `team_card`: it is the one render item that is neither a
message nor a tool group, so it is the one that could fall out of the
turn walk unnoticed.
- SubagentRunPage's live-refresh test unfolded the activity summary before
reaching for a row. A running run plays open now, so that click folded
the rows away instead. Assert it is already open and go straight to the
row; what the test guards — expansion surviving a refresh — is unchanged.
Local desktop suite: 304 files, 4226 passed. Build and desktop-ui-smoke
both green.
The transcript put every assistant reply in its own bordered card, the
same border and radius the tool-activity bar used, so prose and machinery
looked identical and a scrolled session read as a stack of boxes. A
one-line reply cost 112px to show 22px of text.
Alignment and width already say who is speaking — the prompt hugs its
text on the right, the reply takes the full column on the left — so the
border was repeating that at the cost of the space. Drop it, and separate
the layers by tone instead: prose stays primary, tool rows go 12px
tertiary.
Activity now plays open while its turn is still producing into it and
folds into a counted digest once that turn moves on:
thought 5 times, read 2 files, ran 2 commands 8.0s
Counts are what make the line worth having; the older "ran some commands"
summary said strictly less than the rows it hid. Clicking pins the
reader's choice so a run they opened is not closed under them later.
Also:
- Bash rows show the model's own description of a command rather than the
raw pipeline; the command itself moves to the terminal that ran it.
- Rows gain a leading icon, reusing activitySegmentIcon.
- ThinkingBlock loses its second, larger form. A thought that happened not
to be followed by a tool call rendered as a bare label with none of its
content — the least informative thing on screen, for a reason that was
never about the thought. Its preview follows the tail while streaming.
- Message actions render only on the reply that closes a turn. The bar
reserved 36px on every reply whether or not it was hovered, which on a
one-line reply outweighed the text. Mid-turn text stays copyable by
selection.
- Completed background tasks stop drawing a card; the activity panel
already lists them, and a team session emits dozens. Failures still
interrupt.
- Virtualizer estimates are recalibrated to the new box model, and item
height is measured from the border box so the turn gap is counted.
VIRTUAL_MIN_ITEM_HEIGHT clamps measurements too, so it had to drop below
the shortest real row or those rows were recorded too tall forever.
The workbench took over the right-hand slot the moment a team was
discovered: it opened by default, evicted the workspace panel, hid the
workspace and activity toolbar entries outright, and compacted the
transcript down from its 900px reading measure. The main session was
left with the coordination tools filtered out and nothing in their
place, so a team run read as the lead talking to itself.
Replace that with three graduated surfaces. A header strip and an
in-transcript card (rendered where TeamCreate happened) are the whole
main-session footprint; the docked panel now opens on an explicit
gesture; and a `__team__` tab detaches the workbench full screen with
the communication feed in its own column. The panel and the workspace
share one slot but keep independent entries — opening one evicts the
other rather than making it disappear.
Inside the workbench:
- Protocol payloads are narrated instead of printed. The feed was
emitting raw `{"type":"idle_notification",...}` even though
AGENT_LIFECYCLE_TYPES already filters these out of chat; they now
read as sentences and collapse behind a toggle.
- Rows carry their own send time. Every row previously rendered
`T+{snapshotIndex}`, identical for every message in a snapshot.
- The member drawer becomes a workbench view. A full-panel overlay was
unusable at 440px; a teammate's run now renders through MessageList,
so thinking blocks and grouped tool calls match the main session.
Two data-layer causes of the flat member transcript:
- Teammate turns mapped to anonymous `user_text`, making a lead's
instruction indistinguishable from the operator's own prompt. They
now carry `teammateFrom` and render left-aligned and attributed.
- `transcriptMessageFromEntry` dropped `toolUseResult` (and teamStore's
`asEntries` dropped it again), degrading structured tool output to
plain text.