Materialize the question card directly from a permission request when the
streamed tool block has not arrived yet. Upsert by tool-use id so either event
order produces one visible, answerable card without requiring a refresh.
Remove model-bound thinking and redacted thinking blocks from the in-memory
request history when a resumed session changes models. Preserve the persisted
transcript and leave same-model histories untouched.
Two stacked issues kept grok-4.6 requests at high effort even when the
UI selected xhigh: the effort capability table had no Grok entries, so
resolveAppliedEffort clamped xhigh to high; and an explicit --effort
flag never overrode CLAUDE_CODE_EFFORT_LEVEL env for main-loop queries.
Resolve Grok capabilities from the bundled Grok catalog and wire
effortValueOverridesEnv through the REPL for explicit CLI effort.
normalizeRuntimeSelection dropped the effort level whenever the selected
Grok model was missing from the bundled desktop catalog (e.g. grok-4.6),
so the reasoning effort slider snapped back to the default on every drag.
Preserve effort for live-catalog models and let the server validate it,
and sync the bundled desktop catalog with grok-4.6 as the new default.
Add grok-4.6 to the bundled fallback catalog with specs from the live
/v1/models endpoint (500k context, high default effort, xhigh..low
effort levels) and make it the default main model for the Grok Official
provider.
The workbench described a team run in ways the underlying data did not
support, so a member marked "已完成" led to an agent that was still
streaming, and the feed reported no communication for a team that had
just handed out all of its work.
A teammate rewrites its entire transcript every turn, leaving a chain of
`agent-<id>.jsonl` fragments where each repeats its predecessor. The
reader concatenated all of them, and `fragmentScopedId` gave the same
entry a different id per fragment, so id-based deduplication downstream
could not collapse them: one member replayed its work up to eight times
and kept growing. Drop a fragment whose entries are a strict prefix of
another's, ignoring only the four fields a rewrite restamps (agentId,
slug, cwd, promptId). Independent resumes reuse entry ids for different
work, so identity comes from content and those fragments all survive.
The activity path scoped tool ids per fragment before deduplicating,
which split a call from its result across two scopes; it now folds
first.
Member activity came from owning an `in_progress` task. A teammate marks
a task started and can then end its turn, and an umbrella task stays open
across every turn beneath it, so every member read as permanently
working. Report the runner's own turn markers as a `TeamMemberActivity`
separate from `status`, falling back to a transcript-write probe only for
backends that record no markers, and let the workbench say `unknown`
rather than guess.
The remaining corrections follow from the same principle -- show what the
data says:
- A member is drawn once, on the task it most recently started, instead
of on every task it owns; other cards carry an owner chip.
- Opening a task card describes the task; reaching its owner is now a
deliberate second click.
- Cards lead with dependency depth. A task id records the order the lead
wrote tasks down, so review work planned first showed "#1" while
hanging off the bottom of the graph.
- Task handovers are communication, not lifecycle noise, and self-claims
read differently from assignments. Repeated idle notices fold into a
count.
- A blocker missing from the task list counts as resolved, matching
`claimTask`, instead of stranding its dependents in `blocked`.
- Batch-completed tasks record their owner. Automatic ownership only
triggered on `in_progress`, so work closed without ever starting had
none; announcing an already-finished task is skipped so an idle
teammate is not woken for it.
- Transcript pages expose the `TaskUpdate` calls that bound each task,
so a member conversation can be read as the tasks it worked through.
Atlas Cloud now sponsors the project, so give it the same treatment as
the other sponsors: mark the preset featured to move it into the sponsor
row, and add the cc-haha campaign link behind the "get API key" button.
Reorder the preset to sit last in that row, matching the README.
Add Atlas Cloud to the sponsor tables in both README.md and
README.zh-CN.md, using a cc-haha-specific campaign link. Ship
light/dark logo variants so the wordmark stays visible in both
GitHub themes.
One conflict, in SubagentRunPage.test.tsx: both sides rewrote the
`../api/subagents` mock. main added hoisted teammate mocks for the new
TeamMemberRunPage tests; this branch made the factory spread the real module
so the agent-id ref helpers stay real rather than being stubbed into a
fiction. Both are kept — the helpers decide which endpoint the page calls,
and the teammate tests need their handles.
`getTaskListId()` fell back to the session id, and a subagent runs
in-process, so every task one created landed in its parent session's list
and rendered in the UI as if the assistant had planned it. Observed with a
workflow whose scan agent filed three review tasks into the main session.
The reach was the bug, not the capability: an agent tracking its own work
is what lets it hold a goal across a long run. So the list is scoped by
agent rather than the tools being withheld. `TodoWriteTool` has always
keyed on `context.agentId ?? getSessionId()`; this brings the four task
tools in line with it.
Teammates are resolved before the agent branch and keep sharing the
leader's list, which is the point of a team. Also corrects the docstring's
claim that `CLAUDE_CODE_TEAM_NAME` is consulted — `getTeamName()` never
reads it.
A workflow is a JS script the model writes in the moment and hands to the
Workflow tool, which runs it in a locked-down `node:vm` and orchestrates
subagents through `agent()`/`parallel()`/`pipeline()`/`phase()`. Saving one
as a `/name` command is the secondary path; the inline script is the point.
Runtime: cross-realm value marshalling so a script cannot reach the host
`Function`, determinism guards on `Date.now()`/`Math.random()` (they would
make a resume replay diverge), a FIFO concurrency gate, and a journal that
lets an interrupted run resume from its longest unchanged prefix.
Desktop: the run shows up as a `workflow` section in the existing activity
panel — phases as headings, their agents beneath. A workflow agent is an
ordinary subagent run by the same runner, so its row opens the existing
subagent page rather than a parallel viewer of its own; that needed a
`by-agent` lookup, because these agents have no parent `Agent` tool call to
key off. Finished runs are rebuilt from the per-agent sidecars when a
session is reopened, since the live progress stream does not outlive the
process.
The instruction files told every coding agent to reach for agent-browser
whenever a change needed browser-level evidence: copilot-instructions
listed "E2E or agent-browser smoke" as the remedy for cross-boundary
flows, and both contributing guides repeated it. That wording outlived
the tool. With the agent-browser skill uninstalled, agents still parsed
those lines as a recommendation and went looking for the binary instead
of using the browser skill that is actually installed.
Deleting the references would have made the docs wrong. agent-browser is
still a real dependency: check:desktop-ui-smoke spawns it on Linux CI,
and seven maintainer-run e2e scripts under desktop/scripts drive it
directly. It cannot be swapped for ego-browser either — ego lite is a
macOS-only GUI app with no headless mode and a one-time interactive
onboarding, so it cannot run on ubuntu-latest at all.
So the lanes keep the binary and the prose loses the recommendation.
agent-browser is now described as an implementation detail of those two
call sites, and ad-hoc browser work — manual verification, screenshots,
exploratory UI checks — is pointed at the ego-browser skill.
The quality contract asserted the old string, so it would have failed
closed on the reworded line. It now pins the replacement plus the new
routing rule; flipping either sentence turns the test red.
Opening a running subagent's record showed its prompt and nothing else,
while the parent transcript streamed that same agent's tool calls. The
detail route resolves an agent id three ways — a regex over the parent's
tool_result, a task id passed by the client, and a task notification —
and all three only exist once the agent has finished, so a synchronously
dispatched subagent had no transcript to read for its entire run. The
sidecar the CLI writes before the query loop starts already carries the
spawning tool_use id; resolving through that first is the only hint that
exists while the run is still in flight.
Rewinding a real session to the moment that agent was still running: no
agent id and 0 messages before, 38 messages and 14 tool calls after.
That page also offered a composer for every agent. A one-shot subagent
has no inbox — send_agent_message falls through to resumeAgentBackground
and forks a detached copy whose output has nowhere to land. The response
now reports whether an inbox exists, and only named teammates and
in-flight background agents get a composer.
The Agent Teams panel had the opposite problem: at docked width it tried
to be the whole workbench. The dependency map, the member drill-down and
a 210px message strip left no room for the summary that answers whether
the run worked; the seam between transcript and panel only appeared on
hover; and the header progress bar spanned the full width because
Progress is w-full internally and cx does not merge Tailwind classes. The
docked side is a run report now — roster, task list, outcome — while the
map, feed, member views and history scrubbing stay in the workbench tab
that already existed.
Dropped the toolbar's team toggle. It predates AgentTeamsStrip, renders
under exactly the same condition, and does the same thing with less
context.
Desktop suite over the touched areas: 249 passed. Server-side subagent
tests: 30 passed. Lint and tsc clean.
Two conflicts, both adjacent inserts where each side kept its own line:
- MessageList.tsx: TEAM_CARD_HEIGHT landed next to the recalibrated
ACTIVITY_GROUP_COLLAPSED_HEIGHT. Kept both, and restated the team card
at 86 — 78 was measured against the pre-refactor box, which did not yet
carry the turn's 8px padding-bottom.
- MessageList.test.tsx: both sides appended a describe block at EOF.
Follow-ups the merge needed but could not produce on its own:
- UserMessage: the new teammate bubble kept an `mb-5` that the transcript
no longer uses. `.chat-turn-rail` establishes a block formatting
context, so that margin would have been trapped inside the measured box
and made every teammate row 20px taller than a prompt.
- globals.css: the note about containers spacing transcript blocks
themselves was written when SubagentRunPage assembled its own render
model. It renders MessageList now, as does AgentTeamsMemberView, so
both inherit the spacing.
- A rail case for `team_card`: it is the one render item that is neither a
message nor a tool group, so it is the one that could fall out of the
turn walk unnoticed.
- SubagentRunPage's live-refresh test unfolded the activity summary before
reaching for a row. A running run plays open now, so that click folded
the rows away instead. Assert it is already open and go straight to the
row; what the test guards — expansion surviving a refresh — is unchanged.
Local desktop suite: 304 files, 4226 passed. Build and desktop-ui-smoke
both green.
The transcript put every assistant reply in its own bordered card, the
same border and radius the tool-activity bar used, so prose and machinery
looked identical and a scrolled session read as a stack of boxes. A
one-line reply cost 112px to show 22px of text.
Alignment and width already say who is speaking — the prompt hugs its
text on the right, the reply takes the full column on the left — so the
border was repeating that at the cost of the space. Drop it, and separate
the layers by tone instead: prose stays primary, tool rows go 12px
tertiary.
Activity now plays open while its turn is still producing into it and
folds into a counted digest once that turn moves on:
thought 5 times, read 2 files, ran 2 commands 8.0s
Counts are what make the line worth having; the older "ran some commands"
summary said strictly less than the rows it hid. Clicking pins the
reader's choice so a run they opened is not closed under them later.
Also:
- Bash rows show the model's own description of a command rather than the
raw pipeline; the command itself moves to the terminal that ran it.
- Rows gain a leading icon, reusing activitySegmentIcon.
- ThinkingBlock loses its second, larger form. A thought that happened not
to be followed by a tool call rendered as a bare label with none of its
content — the least informative thing on screen, for a reason that was
never about the thought. Its preview follows the tail while streaming.
- Message actions render only on the reply that closes a turn. The bar
reserved 36px on every reply whether or not it was hovered, which on a
one-line reply outweighed the text. Mid-turn text stays copyable by
selection.
- Completed background tasks stop drawing a card; the activity panel
already lists them, and a team session emits dozens. Failures still
interrupt.
- Virtualizer estimates are recalibrated to the new box model, and item
height is measured from the border box so the turn gap is counted.
VIRTUAL_MIN_ITEM_HEIGHT clamps measurements too, so it had to drop below
the shortest real row or those rows were recorded too tall forever.
The smoke waited on `textarea`, but the composer became a ProseMirror
contenteditable (MentionComposer) some time ago, so the selector could
never match. The lane hung for its full 30s timeout and failed on every
run — the screenshot it captured on the way out showed the app rendered
and idle, which is what makes this easy to misread as a UI break.
Target `[data-composer-editor]`, the attribute the editor puts on its
editable node and the same hook composerTestUtils already drives.
27c20389a made a nested tool_use_complete keep the parent's chatState
rather than dropping the session back to `thinking`, so a subagent's
inner tool call no longer clears the parent Task's progress indicator.
That was the intended fix, but the golden fixture still pinned the old
`thinking` value and the two subagent scenarios have failed since —
masked until now because localStorage was taking the whole file down
first.
Regenerating touches exactly one line, and produces an identical file
from a clean HEAD, confirming this is fixture staleness rather than a
behaviour change of ours.
Local desktop suite is now fully green: 304 files, 4202 passed.
Node 22+ defines a global `localStorage` accessor that stays inert
unless the process was started with `--localstorage-file`. vitest's
jsdom environment only copies a window key onto `globalThis` when
nothing is there yet, so Node's inert stub wins over jsdom's real
Storage and `window.localStorage` reads back undefined.
Every suite whose `beforeEach` calls `localStorage.clear()` therefore
died at setup, taking the whole file with it — 12 files and 396 tests
on a clean tree, none of them a real failure. Install a Storage
implementation from a setup file so browser semantics hold regardless
of Node version.
Local desktop suite: 396 failures across 12 files -> 2 across 1.
The workbench took over the right-hand slot the moment a team was
discovered: it opened by default, evicted the workspace panel, hid the
workspace and activity toolbar entries outright, and compacted the
transcript down from its 900px reading measure. The main session was
left with the coordination tools filtered out and nothing in their
place, so a team run read as the lead talking to itself.
Replace that with three graduated surfaces. A header strip and an
in-transcript card (rendered where TeamCreate happened) are the whole
main-session footprint; the docked panel now opens on an explicit
gesture; and a `__team__` tab detaches the workbench full screen with
the communication feed in its own column. The panel and the workspace
share one slot but keep independent entries — opening one evicts the
other rather than making it disappear.
Inside the workbench:
- Protocol payloads are narrated instead of printed. The feed was
emitting raw `{"type":"idle_notification",...}` even though
AGENT_LIFECYCLE_TYPES already filters these out of chat; they now
read as sentences and collapse behind a toggle.
- Rows carry their own send time. Every row previously rendered
`T+{snapshotIndex}`, identical for every message in a snapshot.
- The member drawer becomes a workbench view. A full-panel overlay was
unusable at 440px; a teammate's run now renders through MessageList,
so thinking blocks and grouped tool calls match the main session.
Two data-layer causes of the flat member transcript:
- Teammate turns mapped to anonymous `user_text`, making a lead's
instruction indistinguishable from the operator's own prompt. They
now carry `teammateFrom` and render left-aligned and attributed.
- `transcriptMessageFromEntry` dropped `toolUseResult` (and teamStore's
`asEntries` dropped it again), degrading structured tool output to
plain text.
Built-in agents pin their own models — Explore and claude-code-guide run
on Haiku, statusline-setup on Sonnet — and there was no way to change
that. The only override mechanism was a same-named user agent, which
replaces the definition wholesale: AGENT_SLUG_PATTERN is lowercase-only
so `Explore` and `Plan` cannot even be created, and a replacement loses
the runtime getSystemPrompt and the built-in tool privileges keyed off
`source`. CLAUDE_CODE_SUBAGENT_MODEL is the only other lever and forces
every subagent onto one model.
Add `builtInAgentOverrides` to settings.json, carrying model and effort
per agentType. It is applied in getBuiltInAgents(), the single choke
point every consumer goes through, so the effective value reaches
spawning, `/agents` and the desktop list without any of them knowing an
override exists — and `source` stays `built-in`, preserving the prompt
and tool privileges.
Three constraints are load-bearing rather than stylistic:
- `model: "inherit"` is a real value, not a reset. Built-in defaults
differ per agent and per build, so clearing means deleting the field;
the entry and then the key are removed once empty.
- getSystemPrompt is never wrapped. serializeActiveAgent branches on
`.length === 0`, and claude-code-guide declares one parameter while
the others declare none, so a wrapper makes it destructure undefined.
- The settings write is a read-modify-write inside the file lock.
updateUserSettings is a top-level shallow merge and would replace the
whole record, losing an update when two agents are changed in quick
succession.
strictPluginOnlyCustomization is enforced when resolving, not only when
writing, since settings.json is user-editable by definition.
Also surface edit and delete on the agent list rows. Both already
existed but were reachable only after opening an agent's detail page.
Rows become a div with the primary button and the actions as siblings
rather than nested buttons, keep focus-within so the controls are not
tabbable while invisible, and stay visible on touch where hover never
fires. Built-in rows get the override entry point and no delete.
The store's create/update/override paths now share runAgentMutation
instead of a third hand-copy of the out-of-order guards; the existing
store tests pass unchanged.