Commit Graph

1894 Commits

Author SHA1 Message Date
Relakkes Yang 90a4a11d28 fix(computer-use): remove per-app approval prompts 2026-09-05 03:09:15 +08:00
Relakkes Yang d93d8248c9 fix(computer-use): align Windows virtual cursor with macOS 2026-09-04 23:35:10 +08:00
程序员阿江(Relakkes) e28d0a6534 fix(computer-use): resolve duplicate AX groups 2026-09-02 11:27:51 +08:00
程序员阿江(Relakkes) 5520c62b35 fix(computer-use): keep labeled controls actionable 2026-09-02 11:15:15 +08:00
程序员阿江(Relakkes) 19e88db8ea fix(computer-use): keep startup enumeration out of turns 2026-09-02 10:43:12 +08:00
程序员阿江(Relakkes) 3e3ae6510a fix(desktop): await sidecar cleanup before quit 2026-09-02 10:18:27 +08:00
程序员阿江(Relakkes) 7dd09a4964 fix(computer-use): retire daemon on client shutdown 2026-09-02 09:53:45 +08:00
程序员阿江(Relakkes) e1616ddfc7 fix(computer-use): deminiaturize offscreen macOS targets 2026-09-02 09:26:12 +08:00
程序员阿江(Relakkes) 79f15a08a4 fix(trace): preserve semantics for truncated requests 2026-09-02 09:05:52 +08:00
程序员阿江(Relakkes) f3d23631cf fix(rewind): scope checkpoints to replacement turns 2026-09-02 09:04:36 +08:00
程序员阿江(Relakkes) 4926a00232 fix(desktop): clarify cached token usage 2026-09-02 09:00:42 +08:00
程序员阿江(Relakkes) a7fe10581b test(desktop): stabilize macOS build install check 2026-09-02 08:55:53 +08:00
程序员阿江(Relakkes) df0a566ed2 fix(computer-use): restore offscreen macOS targets 2026-09-02 08:51:02 +08:00
程序员阿江(Relakkes) 1440f0acd8 fix(computer-use): resolve exact app bundle paths 2026-09-02 08:47:33 +08:00
程序员阿江(Relakkes) e1f7cc9d81 fix(desktop): count unique edited files in activity summary 2026-09-02 08:44:48 +08:00
程序员阿江(Relakkes) 08929e88ee fix(desktop): preserve project cwd after session creation 2026-09-02 08:44:48 +08:00
程序员阿江(Relakkes) d354840c3b fix(desktop): install adapter dependencies for macOS builds 2026-09-02 08:44:48 +08:00
程序员阿江(Relakkes) 5ec0947bf8 docs: record post-v0.5.5 issue triage 2026-09-01 21:57:23 +08:00
程序员阿江(Relakkes) 8d8169ea70 fix(models): add GLM 5.3 reasoning capabilities (#1283) 2026-09-01 21:56:49 +08:00
程序员阿江(Relakkes) cf9d6c41ab fix(proxy): preserve Computer Use tool images (#1277) 2026-09-01 21:35:15 +08:00
程序员阿江(Relakkes) aa237f2c22 feat(desktop): expose conversation-only rewind (#1273) 2026-09-01 21:33:55 +08:00
程序员阿江(Relakkes) ece7c54c45 fix(proxy): preserve DeepSeek reasoning content by model (#1269) 2026-09-01 21:33:18 +08:00
程序员阿江(Relakkes) 5eecf5cbc5 fix(api): tolerate progressing tool input streams (#1271) 2026-09-01 21:32:42 +08:00
程序员阿江(Relakkes) c0f3318d9f fix(telegram): preserve complete streamed replies (#1264) 2026-09-01 21:29:09 +08:00
程序员阿江(Relakkes) b1ecce7812 fix(desktop): preserve inline image placement (#1240) 2026-09-01 21:21:24 +08:00
程序员阿江(Relakkes) ee1efaff04 fix(desktop): clarify long rate-limit retries (#1266) 2026-09-01 21:13:25 +08:00
程序员阿江(Relakkes) 0fd6904a95 merge: integrate Computer Use into local main 2026-09-01 20:14:24 +08:00
程序员阿江(Relakkes) 8a37865536 fix(computer-use): harden native macOS automation runtime 2026-09-01 20:06:25 +08:00
程序员阿江(Relakkes) ad7321f7cc fix(computer-use): align long-running macOS automation with Codex 2026-09-01 17:31:17 +08:00
程序员阿江(Relakkes) 8751e558e5 fix(computer-use): preserve background input across focus changes
Track real and synthetic focus with process-bound monitor receipts.
Bind keyboard activation to the snapshot window and align targeted
pointer and keyboard bursts with the reference event sequence.
Wait for pasteboard data delivery before restoring the clipboard,
and add focus/stream diagnostics with transition regression coverage.

Validated with check:native and six covered NetEase focus cycles.
2026-08-31 23:03:54 +08:00
程序员阿江(Relakkes) 217867fff1 fix(computer-use): capture fresh aligned frames for covered windows 2026-08-31 12:30:58 +08:00
程序员阿江(Relakkes) 42d7cb1c93 fix(computer-use): maintain long-lived window capture streams 2026-08-31 11:25:38 +08:00
程序员阿江(Relakkes) 6ace7f865b fix(computer-use): restore background macOS automation 2026-08-26 20:38:47 +08:00
程序员阿江(Relakkes) 66206cd8a3 feat(desktop): surface the system prompt and injected context in trace
A traced session already recorded everything needed to explain a model
request, but the presentation kept it out of reach. The system prompt lived
inside each individual call, so reading it meant picking a call first. The
context the harness assembles — CLAUDE.md, system-reminder blocks,
deferred-tool rosters — was indistinguishable from what the person typed,
because the provider receives all of it as user-role text.

- The session overview names the system prompt and tool catalog, read once
  from the first model call rather than hunted down per request. This is the
  opening header, not a session-wide invariant: late tool registration and a
  mid-session model change rewrite it for later requests, which keep their own
  header on their own detail.
- A model call's detail separates injected context from the exchange, one row
  per injection labelled by its own content instead of its wrapper tag.
- Tool rows carry their input summary and model rows their token counts, so
  scanning the tree distinguishes one call from the next.
- An assistant turn that only reasoned says so, including when the provider
  withheld the reasoning, instead of rendering as a bare label.

Recognition of injected context is by wrapper tag against a closed list.
Position cannot stand in for it: hoistToolResults moves every tool result to
the front of a merged user message, so an attachment the harness appended and
an instruction typed after interrupting a tool arrive in the same shape. The
ambiguity is resolved toward the conversation — unrecognized text stays the
person's message, because looking for what you said and not finding it is the
worse failure.

requestParse also learns the OpenAI Responses wire format, which about a
quarter of traced sessions use and which previously yielded an empty message
list and no system prompt. Responses splits a tool round trip into sibling
function_call / function_call_output entries; across 150 real trace files
those outnumber plain messages 1199 to 878, so they are mapped onto the
tool_use / tool_result vocabulary rather than skipped. Its flat tool
`parameters` and `instructions` spellings are read too, the latter with `||`
so an empty `system` array cannot shadow it.

All new surfaces use design tokens, so the six themes need no per-theme work.

Claude-Session: https://claude.ai/code/session_0119s8U5VzUvVpNgWA7B3pSg
2026-08-26 19:17:24 +08:00
程序员阿江(Relakkes) a834434284 feat(desktop): add a sidebar task view grouped by run state and day
The sidebar could only group sessions by workspace. With many workspaces
open that is orthogonal to the question people actually ask — "which of
the tasks I just started are done?" — so a finished session had to be
hunted for across a dozen collapsed project groups.

A bell in the sidebar header now switches the list to a flat task view:
running sessions first, then sessions bucketed by calendar day (today,
yesterday, previous 7/30 days, earlier). Each row carries its title and
its owning directory, which is exactly the pairing the project grouping
folds away.

The bell shares one state with the existing "organize sidebar -> by
time" menu item, which promised this and still grouped by workspace.
That reuses the `projectOrganization` preference already persisted to
localStorage and to the desktop UI preferences service, so no new
storage key and no migration.

A session stopped on a permission prompt shows a warning dot rather than
the running spinner: it counts as unfinished, but it is waiting on the
user, not working.

Buckets use calendar-day boundaries, not elapsed hours — a task finished
at 23:50 belongs to "yesterday" when read at 00:10, and a mutation to
elapsed-hours arithmetic turns the guard test red.

Verified: check:desktop green (318 files, 4617 tests, lint + tsc +
build); the flat view, bell round trip and preference flip exercised in
a real browser against a sandboxed server with seeded sessions;
group/title/workspace contrast measured >= 4.83:1 across all six themes;
no horizontal overflow at the 240px minimum sidebar width.

Claude-Session: https://claude.ai/code/session_018mbioW3EEr7KbrWVgBvUvJ
2026-08-26 19:17:24 +08:00
程序员阿江(Relakkes) 7858c46bd2 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-26 19:17:24 +08:00
Relakkes Yang e5d385dac2 fix(computer-use): harden Windows desktop automation 2026-08-24 02:32:23 +08:00
程序员阿江-Relakkes a78b4e73ce feat(release): sign Windows artifacts with SignPath (#1265)
feat(release): sign Windows artifacts with SignPath
2026-08-23 19:57:27 +08:00
Relakkes Yang a8cb0c509d test(policy): budget the full dead-import sweep 2026-08-23 19:55:11 +08:00
Relakkes Yang 5fafabfca1 fix(release): isolate draft validation from published assets 2026-08-23 19:36:21 +08:00
Relakkes Yang 417cfdfd69 fix(release): preserve Windows updater config before signing 2026-08-23 19:17:29 +08:00
Relakkes Yang c7a90f2cd1 fix(release): separate SignPath test and release policies 2026-08-23 18:55:10 +08:00
程序员阿江(Relakkes) dd4edc0efe merge: bring main into the computer-use worktree
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.

Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.

`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.

`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.

`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.

Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:39:47 +08:00
程序员阿江(Relakkes) 96e31ede14 fix(computer-use): make the Windows helper refuse what it cannot deliver
Windows has no equivalent of `CGEvent.postToPid`. `pyautogui` bottoms out in
`SendInput`, which injects into the one system-wide input stream and warps the
one real cursor — so on Windows the agent shares the mouse and keyboard with
the user, and `SendInput` reports success unconditionally whether or not
anything acted on the events.

That combination produced the same lie the macOS engine was just fixed for:
click a point behind another window and it lands on that window; click one
off-screen and it lands nowhere; either way the helper answered "Action
completed".

So the helper now refuses instead of guessing:

  * `ForegroundLease` samples `GetLastInputInfo` around every mutating command.
    Interference before the action is `user_interference` — nothing ran, retry
    is safe. Interference during it is `user_interference_result_unknown`,
    because injection already went out and a retry could double-apply it. On a
    play/pause toggle those two differ by exactly one wrong outcome.
  * `ensure_point_on_screen` and `ensure_target_window_reachable` reject
    coordinates outside every display and targets whose windows are minimized
    or hidden, before anything is sent.

`GetLastInputInfo` is the signal because it needs no privileges and is not
advanced by `SendInput`, so the agent cannot trip its own detector. Both guards
fail open on an unreadable reading: a safety layer that turns the feature off
is not safety.

The guard set lives in one place rather than in each dispatcher branch — an
eleventh verb wired like the ten before it would otherwise be silently
unguarded, which is the bug class this pass removes. All ten mutating branches
now return through `_finish`, so no branch can write its own success response
and skip the post-action check.

Also adds a Windows cursor badge. It deliberately does NOT mirror the macOS
virtual cursor: there the real pointer never moves, so the drawn one is the
only cursor and replaces it. Here the real pointer does move, and a second
fake pointer would just be two cursors with one of them lying about where the
click lands. The badge annotates instead — it answers "is this me or the
agent?", which matters because grabbing the mouse mid-action is what makes the
two input streams interleave.

Retires `runtime/mac_helper.py` and its pyobjc requirements: macOS routes every
command to the signed native daemon and `helperBridge` refuses to fall back, so
both were unreachable. Renames the `callPythonHelper` alias to `callHelper`,
which is what it has actually imported since the native engine landed.

Verified with mutation testing — 13 injected regressions across both languages
(dropped lease, bypassed `_finish`, collapsed interference codes, guard reusing
the filtered `list_windows`, badge losing click-through), all caught.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:25:21 +08:00
程序员阿江(Relakkes) f3a6e02903 fix(computer-use): refuse to act through a window that cannot receive input
Three failures shared one root cause: the engine reported success for input it
had no way to deliver or verify.

A fully covered Chromium/CEF window stops drawing, so every screenshot returns
the last frame it painted while each action still answers "Action completed".
One session read a frozen image for three and a half minutes, pressed play four
times, and reported a song playing that the final capture showed paused.
Mutating commands now fail closed with `window_occluded` instead.

The physical-input monitor counted `.mouseMoved`, so moving the cursor anywhere
on screen aborted background automation — the one thing the feature exists to
allow. Movement is not an interaction with any app; presses, drags, keys and
scroll still are.

Focus notifications went out on NSEvent type 13 alone. The target accepted them,
ignored them, and reported nothing: 24 actions, 1 effective. Key-focus subtypes
travel on type 21, and `keyFocusReturned` (0x8000) needs a sign-preserving
truncation or it collapses to subtype 0 — a notification the target accepts and
discards.

The stale-capture notice claimed input does not depend on visibility. Measured
behaviour contradicts it, so it now says actions are refused while covered, and
says it without asking the model to consult the user about window management —
an earlier wording did, and the model handed back a three-step task after one
action.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:24:47 +08:00
Relakkes Yang af4454f38a feat(release): sign Windows artifacts with SignPath 2026-08-23 18:14:52 +08:00
程序员阿江(Relakkes) 27c31bdc15 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-23 17:21:53 +08:00
程序员阿江(Relakkes) ae6e11eeac chore(release): prepare v0.5.5 v0.5.5 2026-08-23 02:33:24 +08:00
程序员阿江(Relakkes) ba859a28ae fix(imagegen): accept user-attached images as edit inputs
ImageEdit trusted only two per-session directories: the bridge upload dir
and its own generated-images dir. Neither is where a user's image actually
lands. An @-mentioned file keeps its original path anywhere on disk, and a
pasted or dropped image goes to ~/.claude/image-cache/<sessionId>/, so every
image a user supplied was refused.

The tool description made it worse by pointing the model at
[Image source: ...] paths, which are exactly the image-cache paths the check
then rejected. The model followed the description, got refused, and asked
users to re-attach a file they had just attached.

Record images the user names explicitly in a session-scoped registry, and
trust the paste directory by location. Paths the model found on its own — a
Glob hit, a path read out of a file — stay refused, which is what the
original root-dir check was protecting against.

Claude-Session: https://claude.ai/code/session_01ArehnJt4QLb83xkEukqNFX
2026-08-23 02:20:23 +08:00
程序员阿江(Relakkes) ba8a5cb832 fix(desktop): show generated files once per turn 2026-08-23 02:18:50 +08:00