Commit Graph

1220 Commits

Author SHA1 Message Date
Relakkes Yang e5d385dac2 fix(computer-use): harden Windows desktop automation 2026-08-24 02:32:23 +08:00
程序员阿江(Relakkes) dd4edc0efe merge: bring main into the computer-use worktree
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.

Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.

`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.

`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.

`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.

Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:39:47 +08:00
程序员阿江(Relakkes) 27c31bdc15 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-23 17:21:53 +08:00
程序员阿江(Relakkes) ae6e11eeac chore(release): prepare v0.5.5 2026-08-23 02:33:24 +08:00
程序员阿江(Relakkes) ba8a5cb832 fix(desktop): show generated files once per turn 2026-08-23 02:18:50 +08:00
程序员阿江(Relakkes) 4575887630 fix(desktop): make the open-with menu a short, named list
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.

`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.

What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.

Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.

The rest of the menu:

  - The system-default row is named after the application it will actually open,
    but still launches through the system-default target. Only that path carries
    the guard that refuses to hand an executable to the shell.
  - A file type nothing can edit no longer offers an editor. Reusing the
    workspace preview gate rather than writing a third extension table.
  - One editor is guaranteed a slot and the rest fill what the applications
    leave. On Windows and Linux there is no application list, so a hard cap of
    one would have dropped installed editors for nothing.
  - The file manager is named per platform instead of interpolating the server's
    label, which produced "在 Explorer 中显示" — an English name inside a Chinese
    sentence, and not what Windows calls it. Same shape as #1236.
  - `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
    scroll inside the overlay as a viewport change: the listener is on capture,
    so the menu closed the instant the user reached for its own scrollbar.

Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.

And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.

Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
2026-08-22 21:08:33 +08:00
程序员阿江(Relakkes) 6fb8105439 fix: resolve reported provider and composer issues 2026-08-21 17:51:27 +08:00
程序员阿江(Relakkes) b09e252fcc fix(desktop): stabilize chat image rendering
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
2026-08-18 17:52:01 +08:00
程序员阿江(Relakkes) a0e63ccae7 fix(desktop): adapt file manager labels by platform #1236 2026-08-18 17:21:49 +08:00
程序员阿江(Relakkes) 783f539b84 feat(desktop): add model-aware default reasoning effort 2026-08-16 16:35:20 +08:00
程序员阿江(Relakkes) 1d221f9f8d fix(desktop): open generated files with native apps #1233 2026-08-16 15:33:33 +08:00
程序员阿江(Relakkes) 2792406815 fix(provider): make Tool Search opt-in #1229 2026-08-16 15:32:27 +08:00
程序员阿江(Relakkes) b5d60a77ae fix(desktop): stabilize virtual row measurement #1223 2026-08-16 01:42:46 +08:00
程序员阿江(Relakkes) e2e302e1b3 fix(desktop): reconcile stale H5 task state #1214 2026-08-16 01:40:56 +08:00
程序员阿江(Relakkes) 7eb406243e fix(desktop): warn on proxy-only settings takeover #1219 2026-08-16 00:33:34 +08:00
程序员阿江(Relakkes) 5001502882 fix(desktop): recover safely from sidecar crashes #1227 2026-08-16 00:18:09 +08:00
程序员阿江(Relakkes) e72c55264f fix(rewind): prevent long-session checkpoint stalls 2026-08-15 23:59:18 +08:00
程序员阿江(Relakkes) d52bbec707 chore(release): prepare v0.5.4 2026-08-14 18:11:23 +08:00
程序员阿江(Relakkes) 9f5c1951a2 fix(desktop): reuse model picker for agent settings 2026-08-14 16:43:44 +08:00
程序员阿江(Relakkes) a8c900b591 fix(auth): align Claude OAuth model and recovery 2026-08-14 14:59:44 +08:00
程序员阿江(Relakkes) 8703165e01 fix(h5): show AskUserQuestion before tool stream completes #1220
Materialize the question card directly from a permission request when the
streamed tool block has not arrived yet. Upsert by tool-use id so either event
order produces one visible, answerable card without requiring a refresh.
2026-08-13 23:12:08 +08:00
程序员阿江(Relakkes) bb74c9246d fix(desktop): protect unloaded sessions from tab cleanup #1217 2026-08-13 22:54:05 +08:00
程序员阿江(Relakkes) 8673f2092a chore(sponsor): remove TeamoRouter sponsor and provider preset 2026-08-13 08:36:50 +08:00
程序员阿江(Relakkes) 59408abaae fix(provider): keep Grok effort editable for live-catalog models
normalizeRuntimeSelection dropped the effort level whenever the selected
Grok model was missing from the bundled desktop catalog (e.g. grok-4.6),
so the reasoning effort slider snapped back to the default on every drag.
Preserve effort for live-catalog models and let the server validate it,
and sync the bundled desktop catalog with grok-4.6 as the new default.
2026-08-13 04:18:06 +08:00
程序员阿江(Relakkes) 963c09bf6e feat(teams): rebuild the collaborative workbench 2026-08-12 14:36:16 +08:00
程序员阿江(Relakkes) c0f4ba5f84 fix(teams): report real member and task state in the workbench
The workbench described a team run in ways the underlying data did not
support, so a member marked "已完成" led to an agent that was still
streaming, and the feed reported no communication for a team that had
just handed out all of its work.

A teammate rewrites its entire transcript every turn, leaving a chain of
`agent-<id>.jsonl` fragments where each repeats its predecessor. The
reader concatenated all of them, and `fragmentScopedId` gave the same
entry a different id per fragment, so id-based deduplication downstream
could not collapse them: one member replayed its work up to eight times
and kept growing. Drop a fragment whose entries are a strict prefix of
another's, ignoring only the four fields a rewrite restamps (agentId,
slug, cwd, promptId). Independent resumes reuse entry ids for different
work, so identity comes from content and those fragments all survive.
The activity path scoped tool ids per fragment before deduplicating,
which split a call from its result across two scopes; it now folds
first.

Member activity came from owning an `in_progress` task. A teammate marks
a task started and can then end its turn, and an umbrella task stays open
across every turn beneath it, so every member read as permanently
working. Report the runner's own turn markers as a `TeamMemberActivity`
separate from `status`, falling back to a transcript-write probe only for
backends that record no markers, and let the workbench say `unknown`
rather than guess.

The remaining corrections follow from the same principle -- show what the
data says:

- A member is drawn once, on the task it most recently started, instead
  of on every task it owns; other cards carry an owner chip.
- Opening a task card describes the task; reaching its owner is now a
  deliberate second click.
- Cards lead with dependency depth. A task id records the order the lead
  wrote tasks down, so review work planned first showed "#1" while
  hanging off the bottom of the graph.
- Task handovers are communication, not lifecycle noise, and self-claims
  read differently from assignments. Repeated idle notices fold into a
  count.
- A blocker missing from the task list counts as resolved, matching
  `claimTask`, instead of stranding its dependents in `blocked`.
- Batch-completed tasks record their owner. Automatic ownership only
  triggered on `in_progress`, so work closed without ever starting had
  none; announcing an already-finished task is skipped so an idle
  teammate is not woken for it.
- Transcript pages expose the `TaskUpdate` calls that bound each task,
  so a member conversation can be read as the tasks it worked through.
2026-08-11 20:17:07 +08:00
程序员阿江(Relakkes) ab0d89e16e fix(agents): stream child runs in shared UI 2026-08-11 14:46:28 +08:00
程序员阿江(Relakkes) b5a4d23af1 fix(teams): stabilize activity, history, and lifecycle 2026-08-11 05:57:15 +08:00
程序员阿江(Relakkes) bf695f21d7 fix(desktop): isolate Agent Teams activity and share session UI 2026-08-11 02:42:10 +08:00
程序员阿江(Relakkes) a0afcc7efa feat(desktop): add custom project display names 2026-08-11 02:38:55 +08:00
程序员阿江(Relakkes) c23632d2c4 fix(teams): synchronize member tasks and activity 2026-08-10 19:43:04 +08:00
程序员阿江(Relakkes) 795120fbca fix(teams): reconcile Activity with workbench state 2026-08-10 19:13:29 +08:00
程序员阿江(Relakkes) 3a9ba776bb fix(workflows): save desktop runs and surface cached activity 2026-08-10 18:19:06 +08:00
程序员阿江(Relakkes) 9c9a0885b9 fix(teams): align task status and slash entry 2026-08-10 17:41:49 +08:00
程序员阿江(Relakkes) eee09fe3f4 fix(desktop): add Ultracode keyword controls 2026-08-10 17:31:45 +08:00
程序员阿江(Relakkes) 38de706963 fix(desktop): retain newly created branch for session launch 2026-08-10 17:02:01 +08:00
程序员阿江(Relakkes) 3719cd76f5 fix(agents): reload built-in overrides in active sessions 2026-08-10 16:52:31 +08:00
程序员阿江(Relakkes) fdead6b27a fix(desktop): show undo for unverified-only turns 2026-08-10 16:48:47 +08:00
程序员阿江(Relakkes) e64fc58b4d fix(activity): scope agent activity to its owning run 2026-08-10 16:36:45 +08:00
程序员阿江(Relakkes) 0b13dc6f71 fix(sessions): keep worktree paths authoritative 2026-08-10 16:23:44 +08:00
程序员阿江(Relakkes) 2504ffd877 fix(desktop): render teammate Markdown and sync member activity 2026-08-10 04:37:25 +08:00
程序员阿江(Relakkes) 17e06146ec fix(workflows): deduplicate resumed activity runs 2026-08-10 03:36:50 +08:00
程序员阿江(Relakkes) 4644804afd Merge main into the dynamic workflows branch
One conflict, in SubagentRunPage.test.tsx: both sides rewrote the
`../api/subagents` mock. main added hoisted teammate mocks for the new
TeamMemberRunPage tests; this branch made the factory spread the real module
so the agent-id ref helpers stay real rather than being stubbed into a
fiction. Both are kept — the helpers decide which endpoint the page calls,
and the teammate tests need their handles.
2026-08-10 00:21:03 +08:00
程序员阿江(Relakkes) 3ce7c67822 feat(workflows): dynamic workflow orchestration end to end
A workflow is a JS script the model writes in the moment and hands to the
Workflow tool, which runs it in a locked-down `node:vm` and orchestrates
subagents through `agent()`/`parallel()`/`pipeline()`/`phase()`. Saving one
as a `/name` command is the secondary path; the inline script is the point.

Runtime: cross-realm value marshalling so a script cannot reach the host
`Function`, determinism guards on `Date.now()`/`Math.random()` (they would
make a resume replay diverge), a FIFO concurrency gate, and a journal that
lets an interrupted run resume from its longest unchanged prefix.

Desktop: the run shows up as a `workflow` section in the existing activity
panel — phases as headings, their agents beneath. A workflow agent is an
ordinary subagent run by the same runner, so its row opens the existing
subagent page rather than a parallel viewer of its own; that needed a
`by-agent` lookup, because these agents have no parent `Agent` tool call to
key off. Finished runs are rebuilt from the per-agent sidecars when a
session is reopened, since the live progress stream does not outlive the
process.
2026-08-10 00:17:03 +08:00
程序员阿江(Relakkes) be9c539223 fix(desktop): streamline AgentTeams workbench 2026-08-10 00:00:44 +08:00
程序员阿江(Relakkes) b0178a717a refactor(desktop): compact completed-turn footer 2026-08-10 00:00:16 +08:00
程序员阿江(Relakkes) 0f3a517085 fix(desktop): recover live subagent runs and split the team report from the workbench
Opening a running subagent's record showed its prompt and nothing else,
while the parent transcript streamed that same agent's tool calls. The
detail route resolves an agent id three ways — a regex over the parent's
tool_result, a task id passed by the client, and a task notification —
and all three only exist once the agent has finished, so a synchronously
dispatched subagent had no transcript to read for its entire run. The
sidecar the CLI writes before the query loop starts already carries the
spawning tool_use id; resolving through that first is the only hint that
exists while the run is still in flight.

Rewinding a real session to the moment that agent was still running: no
agent id and 0 messages before, 38 messages and 14 tool calls after.

That page also offered a composer for every agent. A one-shot subagent
has no inbox — send_agent_message falls through to resumeAgentBackground
and forks a detached copy whose output has nowhere to land. The response
now reports whether an inbox exists, and only named teammates and
in-flight background agents get a composer.

The Agent Teams panel had the opposite problem: at docked width it tried
to be the whole workbench. The dependency map, the member drill-down and
a 210px message strip left no room for the summary that answers whether
the run worked; the seam between transcript and panel only appeared on
hover; and the header progress bar spanned the full width because
Progress is w-full internally and cx does not merge Tailwind classes. The
docked side is a run report now — roster, task list, outcome — while the
map, feed, member views and history scrubbing stay in the workbench tab
that already existed.

Dropped the toolbar's team toggle. It predates AgentTeamsStrip, renders
under exactly the same condition, and does the same thing with less
context.

Desktop suite over the touched areas: 249 passed. Server-side subagent
tests: 30 passed. Lint and tsc clean.
2026-08-09 05:03:38 +08:00
程序员阿江(Relakkes) 17d231a6bf Merge agent-teams workbench into the transcript refactor
Two conflicts, both adjacent inserts where each side kept its own line:

- MessageList.tsx: TEAM_CARD_HEIGHT landed next to the recalibrated
  ACTIVITY_GROUP_COLLAPSED_HEIGHT. Kept both, and restated the team card
  at 86 — 78 was measured against the pre-refactor box, which did not yet
  carry the turn's 8px padding-bottom.
- MessageList.test.tsx: both sides appended a describe block at EOF.

Follow-ups the merge needed but could not produce on its own:

- UserMessage: the new teammate bubble kept an `mb-5` that the transcript
  no longer uses. `.chat-turn-rail` establishes a block formatting
  context, so that margin would have been trapped inside the measured box
  and made every teammate row 20px taller than a prompt.
- globals.css: the note about containers spacing transcript blocks
  themselves was written when SubagentRunPage assembled its own render
  model. It renders MessageList now, as does AgentTeamsMemberView, so
  both inherit the spacing.
- A rail case for `team_card`: it is the one render item that is neither a
  message nor a tool group, so it is the one that could fall out of the
  turn walk unnoticed.
- SubagentRunPage's live-refresh test unfolded the activity summary before
  reaching for a row. A running run plays open now, so that click folded
  the rows away instead. Assert it is already open and go straight to the
  row; what the test guards — expansion surviving a refresh — is unchanged.

Local desktop suite: 304 files, 4226 passed. Build and desktop-ui-smoke
both green.
2026-08-09 03:00:44 +08:00
程序员阿江(Relakkes) c896347b8e refactor(desktop): fold settled tool activity into a counted summary
The transcript put every assistant reply in its own bordered card, the
same border and radius the tool-activity bar used, so prose and machinery
looked identical and a scrolled session read as a stack of boxes. A
one-line reply cost 112px to show 22px of text.

Alignment and width already say who is speaking — the prompt hugs its
text on the right, the reply takes the full column on the left — so the
border was repeating that at the cost of the space. Drop it, and separate
the layers by tone instead: prose stays primary, tool rows go 12px
tertiary.

Activity now plays open while its turn is still producing into it and
folds into a counted digest once that turn moves on:

    thought 5 times, read 2 files, ran 2 commands        8.0s

Counts are what make the line worth having; the older "ran some commands"
summary said strictly less than the rows it hid. Clicking pins the
reader's choice so a run they opened is not closed under them later.

Also:
- Bash rows show the model's own description of a command rather than the
  raw pipeline; the command itself moves to the terminal that ran it.
- Rows gain a leading icon, reusing activitySegmentIcon.
- ThinkingBlock loses its second, larger form. A thought that happened not
  to be followed by a tool call rendered as a bare label with none of its
  content — the least informative thing on screen, for a reason that was
  never about the thought. Its preview follows the tail while streaming.
- Message actions render only on the reply that closes a turn. The bar
  reserved 36px on every reply whether or not it was hovered, which on a
  one-line reply outweighed the text. Mid-turn text stays copyable by
  selection.
- Completed background tasks stop drawing a card; the activity panel
  already lists them, and a team session emits dozens. Failures still
  interrupt.
- Virtualizer estimates are recalibrated to the new box model, and item
  height is measured from the border box so the turn gap is counted.
  VIRTUAL_MIN_ITEM_HEIGHT clamps measurements too, so it had to drop below
  the shortest real row or those rows were recorded too tall forever.
2026-08-09 02:47:34 +08:00
程序员阿江(Relakkes) 8288d1e42e test(chat): realign the stale subagent golden snapshot
27c20389a made a nested tool_use_complete keep the parent's chatState
rather than dropping the session back to `thinking`, so a subagent's
inner tool call no longer clears the parent Task's progress indicator.
That was the intended fix, but the golden fixture still pinned the old
`thinking` value and the two subagent scenarios have failed since —
masked until now because localStorage was taking the whole file down
first.

Regenerating touches exactly one line, and produces an identical file
from a clean HEAD, confirming this is fixture staleness rather than a
behaviour change of ours.

Local desktop suite is now fully green: 304 files, 4202 passed.
2026-08-08 23:57:42 +08:00