Commit Graph

1874 Commits

Author SHA1 Message Date
程序员阿江(Relakkes) aa237f2c22 feat(desktop): expose conversation-only rewind (#1273) 2026-09-01 21:33:55 +08:00
程序员阿江(Relakkes) ece7c54c45 fix(proxy): preserve DeepSeek reasoning content by model (#1269) 2026-09-01 21:33:18 +08:00
程序员阿江(Relakkes) 5eecf5cbc5 fix(api): tolerate progressing tool input streams (#1271) 2026-09-01 21:32:42 +08:00
程序员阿江(Relakkes) c0f3318d9f fix(telegram): preserve complete streamed replies (#1264) 2026-09-01 21:29:09 +08:00
程序员阿江(Relakkes) b1ecce7812 fix(desktop): preserve inline image placement (#1240) 2026-09-01 21:21:24 +08:00
程序员阿江(Relakkes) ee1efaff04 fix(desktop): clarify long rate-limit retries (#1266) 2026-09-01 21:13:25 +08:00
程序员阿江(Relakkes) 0fd6904a95 merge: integrate Computer Use into local main 2026-09-01 20:14:24 +08:00
程序员阿江(Relakkes) 8a37865536 fix(computer-use): harden native macOS automation runtime 2026-09-01 20:06:25 +08:00
程序员阿江(Relakkes) ad7321f7cc fix(computer-use): align long-running macOS automation with Codex 2026-09-01 17:31:17 +08:00
程序员阿江(Relakkes) 8751e558e5 fix(computer-use): preserve background input across focus changes
Track real and synthetic focus with process-bound monitor receipts.
Bind keyboard activation to the snapshot window and align targeted
pointer and keyboard bursts with the reference event sequence.
Wait for pasteboard data delivery before restoring the clipboard,
and add focus/stream diagnostics with transition regression coverage.

Validated with check:native and six covered NetEase focus cycles.
2026-08-31 23:03:54 +08:00
程序员阿江(Relakkes) 217867fff1 fix(computer-use): capture fresh aligned frames for covered windows 2026-08-31 12:30:58 +08:00
程序员阿江(Relakkes) 42d7cb1c93 fix(computer-use): maintain long-lived window capture streams 2026-08-31 11:25:38 +08:00
程序员阿江(Relakkes) 6ace7f865b fix(computer-use): restore background macOS automation 2026-08-26 20:38:47 +08:00
程序员阿江(Relakkes) 66206cd8a3 feat(desktop): surface the system prompt and injected context in trace
A traced session already recorded everything needed to explain a model
request, but the presentation kept it out of reach. The system prompt lived
inside each individual call, so reading it meant picking a call first. The
context the harness assembles — CLAUDE.md, system-reminder blocks,
deferred-tool rosters — was indistinguishable from what the person typed,
because the provider receives all of it as user-role text.

- The session overview names the system prompt and tool catalog, read once
  from the first model call rather than hunted down per request. This is the
  opening header, not a session-wide invariant: late tool registration and a
  mid-session model change rewrite it for later requests, which keep their own
  header on their own detail.
- A model call's detail separates injected context from the exchange, one row
  per injection labelled by its own content instead of its wrapper tag.
- Tool rows carry their input summary and model rows their token counts, so
  scanning the tree distinguishes one call from the next.
- An assistant turn that only reasoned says so, including when the provider
  withheld the reasoning, instead of rendering as a bare label.

Recognition of injected context is by wrapper tag against a closed list.
Position cannot stand in for it: hoistToolResults moves every tool result to
the front of a merged user message, so an attachment the harness appended and
an instruction typed after interrupting a tool arrive in the same shape. The
ambiguity is resolved toward the conversation — unrecognized text stays the
person's message, because looking for what you said and not finding it is the
worse failure.

requestParse also learns the OpenAI Responses wire format, which about a
quarter of traced sessions use and which previously yielded an empty message
list and no system prompt. Responses splits a tool round trip into sibling
function_call / function_call_output entries; across 150 real trace files
those outnumber plain messages 1199 to 878, so they are mapped onto the
tool_use / tool_result vocabulary rather than skipped. Its flat tool
`parameters` and `instructions` spellings are read too, the latter with `||`
so an empty `system` array cannot shadow it.

All new surfaces use design tokens, so the six themes need no per-theme work.

Claude-Session: https://claude.ai/code/session_0119s8U5VzUvVpNgWA7B3pSg
2026-08-26 19:17:24 +08:00
程序员阿江(Relakkes) a834434284 feat(desktop): add a sidebar task view grouped by run state and day
The sidebar could only group sessions by workspace. With many workspaces
open that is orthogonal to the question people actually ask — "which of
the tasks I just started are done?" — so a finished session had to be
hunted for across a dozen collapsed project groups.

A bell in the sidebar header now switches the list to a flat task view:
running sessions first, then sessions bucketed by calendar day (today,
yesterday, previous 7/30 days, earlier). Each row carries its title and
its owning directory, which is exactly the pairing the project grouping
folds away.

The bell shares one state with the existing "organize sidebar -> by
time" menu item, which promised this and still grouped by workspace.
That reuses the `projectOrganization` preference already persisted to
localStorage and to the desktop UI preferences service, so no new
storage key and no migration.

A session stopped on a permission prompt shows a warning dot rather than
the running spinner: it counts as unfinished, but it is waiting on the
user, not working.

Buckets use calendar-day boundaries, not elapsed hours — a task finished
at 23:50 belongs to "yesterday" when read at 00:10, and a mutation to
elapsed-hours arithmetic turns the guard test red.

Verified: check:desktop green (318 files, 4617 tests, lint + tsc +
build); the flat view, bell round trip and preference flip exercised in
a real browser against a sandboxed server with seeded sessions;
group/title/workspace contrast measured >= 4.83:1 across all six themes;
no horizontal overflow at the 240px minimum sidebar width.

Claude-Session: https://claude.ai/code/session_018mbioW3EEr7KbrWVgBvUvJ
2026-08-26 19:17:24 +08:00
程序员阿江(Relakkes) 7858c46bd2 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-26 19:17:24 +08:00
Relakkes Yang e5d385dac2 fix(computer-use): harden Windows desktop automation 2026-08-24 02:32:23 +08:00
程序员阿江-Relakkes a78b4e73ce feat(release): sign Windows artifacts with SignPath (#1265)
feat(release): sign Windows artifacts with SignPath
2026-08-23 19:57:27 +08:00
Relakkes Yang a8cb0c509d test(policy): budget the full dead-import sweep 2026-08-23 19:55:11 +08:00
Relakkes Yang 5fafabfca1 fix(release): isolate draft validation from published assets 2026-08-23 19:36:21 +08:00
Relakkes Yang 417cfdfd69 fix(release): preserve Windows updater config before signing 2026-08-23 19:17:29 +08:00
Relakkes Yang c7a90f2cd1 fix(release): separate SignPath test and release policies 2026-08-23 18:55:10 +08:00
程序员阿江(Relakkes) dd4edc0efe merge: bring main into the computer-use worktree
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.

Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.

`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.

`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.

`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.

Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:39:47 +08:00
程序员阿江(Relakkes) 96e31ede14 fix(computer-use): make the Windows helper refuse what it cannot deliver
Windows has no equivalent of `CGEvent.postToPid`. `pyautogui` bottoms out in
`SendInput`, which injects into the one system-wide input stream and warps the
one real cursor — so on Windows the agent shares the mouse and keyboard with
the user, and `SendInput` reports success unconditionally whether or not
anything acted on the events.

That combination produced the same lie the macOS engine was just fixed for:
click a point behind another window and it lands on that window; click one
off-screen and it lands nowhere; either way the helper answered "Action
completed".

So the helper now refuses instead of guessing:

  * `ForegroundLease` samples `GetLastInputInfo` around every mutating command.
    Interference before the action is `user_interference` — nothing ran, retry
    is safe. Interference during it is `user_interference_result_unknown`,
    because injection already went out and a retry could double-apply it. On a
    play/pause toggle those two differ by exactly one wrong outcome.
  * `ensure_point_on_screen` and `ensure_target_window_reachable` reject
    coordinates outside every display and targets whose windows are minimized
    or hidden, before anything is sent.

`GetLastInputInfo` is the signal because it needs no privileges and is not
advanced by `SendInput`, so the agent cannot trip its own detector. Both guards
fail open on an unreadable reading: a safety layer that turns the feature off
is not safety.

The guard set lives in one place rather than in each dispatcher branch — an
eleventh verb wired like the ten before it would otherwise be silently
unguarded, which is the bug class this pass removes. All ten mutating branches
now return through `_finish`, so no branch can write its own success response
and skip the post-action check.

Also adds a Windows cursor badge. It deliberately does NOT mirror the macOS
virtual cursor: there the real pointer never moves, so the drawn one is the
only cursor and replaces it. Here the real pointer does move, and a second
fake pointer would just be two cursors with one of them lying about where the
click lands. The badge annotates instead — it answers "is this me or the
agent?", which matters because grabbing the mouse mid-action is what makes the
two input streams interleave.

Retires `runtime/mac_helper.py` and its pyobjc requirements: macOS routes every
command to the signed native daemon and `helperBridge` refuses to fall back, so
both were unreachable. Renames the `callPythonHelper` alias to `callHelper`,
which is what it has actually imported since the native engine landed.

Verified with mutation testing — 13 injected regressions across both languages
(dropped lease, bypassed `_finish`, collapsed interference codes, guard reusing
the filtered `list_windows`, badge losing click-through), all caught.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:25:21 +08:00
程序员阿江(Relakkes) f3a6e02903 fix(computer-use): refuse to act through a window that cannot receive input
Three failures shared one root cause: the engine reported success for input it
had no way to deliver or verify.

A fully covered Chromium/CEF window stops drawing, so every screenshot returns
the last frame it painted while each action still answers "Action completed".
One session read a frozen image for three and a half minutes, pressed play four
times, and reported a song playing that the final capture showed paused.
Mutating commands now fail closed with `window_occluded` instead.

The physical-input monitor counted `.mouseMoved`, so moving the cursor anywhere
on screen aborted background automation — the one thing the feature exists to
allow. Movement is not an interaction with any app; presses, drags, keys and
scroll still are.

Focus notifications went out on NSEvent type 13 alone. The target accepted them,
ignored them, and reported nothing: 24 actions, 1 effective. Key-focus subtypes
travel on type 21, and `keyFocusReturned` (0x8000) needs a sign-preserving
truncation or it collapses to subtype 0 — a notification the target accepts and
discards.

The stale-capture notice claimed input does not depend on visibility. Measured
behaviour contradicts it, so it now says actions are refused while covered, and
says it without asking the model to consult the user about window management —
an earlier wording did, and the model handed back a three-step task after one
action.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:24:47 +08:00
Relakkes Yang af4454f38a feat(release): sign Windows artifacts with SignPath 2026-08-23 18:14:52 +08:00
程序员阿江(Relakkes) 27c31bdc15 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-23 17:21:53 +08:00
程序员阿江(Relakkes) ae6e11eeac chore(release): prepare v0.5.5 v0.5.5 2026-08-23 02:33:24 +08:00
程序员阿江(Relakkes) ba859a28ae fix(imagegen): accept user-attached images as edit inputs
ImageEdit trusted only two per-session directories: the bridge upload dir
and its own generated-images dir. Neither is where a user's image actually
lands. An @-mentioned file keeps its original path anywhere on disk, and a
pasted or dropped image goes to ~/.claude/image-cache/<sessionId>/, so every
image a user supplied was refused.

The tool description made it worse by pointing the model at
[Image source: ...] paths, which are exactly the image-cache paths the check
then rejected. The model followed the description, got refused, and asked
users to re-attach a file they had just attached.

Record images the user names explicitly in a session-scoped registry, and
trust the paste directory by location. Paths the model found on its own — a
Glob hit, a path read out of a file — stay refused, which is what the
original root-dir check was protecting against.

Claude-Session: https://claude.ai/code/session_01ArehnJt4QLb83xkEukqNFX
2026-08-23 02:20:23 +08:00
程序员阿江(Relakkes) ba8a5cb832 fix(desktop): show generated files once per turn 2026-08-23 02:18:50 +08:00
程序员阿江(Relakkes) 4575887630 fix(desktop): make the open-with menu a short, named list
The menu poured out everything LaunchServices returns. Measured on the dev
machine: 26 applications for a `.csv`, 24 for `.md`, 16 for `.pdf` — 34, 32 and
24 menu rows once the system default, six IDEs, the copy entries and Finder were
added. Most of it was not installed software but browser cores staged in caches:
three copies of Chrome for Testing (playwright, agent-browser), two of
BitBrowser, Warp's autoupdate directory, a LibreOffice shipped inside a runtime.

`discoverNativeApplications` deduped by `appPath` while the target id came from
`bundleId || appPath`, so the two disagreed. Copies of one bundle each survived
the filter and then collapsed onto a single id: duplicate React keys, and
`openTarget`'s `find` always returning the first record — clicking the second
copy launched the first. The keys are one function now, so they cannot drift
again.

What replaces the dump is a ranking, not a whitelist. A whitelist of install
directories was the obvious fix and it is wrong: Safari really lives in
`/System/Volumes/Preboot/Cryptexes/App/System/Applications`, so it would have
silently dropped the default browser for HTML. Location became one sort key
among five — `isDefault`, an extension-to-bundle-id table (`pdf` → Preview,
`docx` → Word/Pages), location tier, Spotlight's `kMDItemUseCount`, name — and
the list is cut to five. Cache copies never accumulate a launch count, because
nothing execs them through LaunchServices; every tier still ships, so a
misjudged location sorts lower instead of disappearing.

Icons never rendered, and not only the discovered ones. `TargetIcon` pointed an
`<img src>` at `/api/open-targets/icons/…`, which is a cross-origin no-cors
subresource: no Authorization header, so the server's fetch-metadata policy
answers 401 and Chrome drops it as ERR_BLOCKED_BY_ORB. The same URL returns 200
from a terminal, which is what hid this. Icons now come through the credential
path as blob URLs, with hits and misses both cached — a miss costs a `sips` run
per row otherwise. Same failure and same fix as d14866f6d, which never reached
this branch.

The rest of the menu:

  - The system-default row is named after the application it will actually open,
    but still launches through the system-default target. Only that path carries
    the guard that refuses to hand an executable to the shell.
  - A file type nothing can edit no longer offers an editor. Reusing the
    workspace preview gate rather than writing a third extension table.
  - One editor is guaranteed a slot and the rest fill what the applications
    leave. On Windows and Linux there is no application list, so a hard cap of
    one would have dropped installed editors for nothing.
  - The file manager is named per platform instead of interpolating the server's
    label, which produced "在 Explorer 中显示" — an English name inside a Chinese
    sentence, and not what Windows calls it. Same shape as #1236.
  - `max-height` and a scrollbar. That needed `useDismissable` to stop treating a
    scroll inside the overlay as a viewport change: the listener is on capture,
    so the menu closed the instant the user reached for its own scrollbar.

Deliverables written by a shell command reach the transcript again. The
checkpoint records only the file-editing tools, so `Write plan.md` plus
`python make_report.py` lost the report: the mention did not match a changed
file and was dropped. Document formats now survive that lookup — not Markdown or
images, which this product reads as much as it writes. A bare filename is
anchored to the directory the turn wrote into, since the prose gives the
directory once and then lists basenames; without that the card rendered and
could not be opened.

And when a path really is gone, the click says so. Those failures landed in
floating promises, so a stale reference did nothing at all and read as a broken
button.

Claude-Session: https://claude.ai/code/session_019nFuRvsozQcErtGjVkzZCN
2026-08-22 21:08:33 +08:00
Relakkes Yang fef789f5a8 fix(quality-gate): stabilize Windows validation 2026-08-22 16:06:17 +08:00
程序员阿江(Relakkes) 6fb8105439 fix: resolve reported provider and composer issues 2026-08-21 17:51:27 +08:00
程序员阿江(Relakkes) 9a602e0e17 fix(api): stop truncated tool calls from hanging sessions #1237 2026-08-18 19:04:22 +08:00
程序员阿江(Relakkes) b09e252fcc fix(desktop): stabilize chat image rendering
- preserve readable single-image previews while keeping multi-image layouts compact
- give assistant output images one sandboxed rendering owner
- reject unsafe Markdown images before the browser can request them
2026-08-18 17:52:01 +08:00
程序员阿江(Relakkes) a0e63ccae7 fix(desktop): adapt file manager labels by platform #1236 2026-08-18 17:21:49 +08:00
程序员阿江(Relakkes) 3b5e63e42b fix(agents): degrade worktree isolation outside a git repo
Requesting worktree isolation in a workspace with no git repository and
no WorktreeCreate hook failed every subagent it was asked for, so a
whole workflow run came back with nothing. Isolation is best-effort, not
a precondition for running: the agent now runs in the workspace
directory instead, and the skip is reported rather than hidden — once
per run in the workflow log, and in the Agent tool's result.

Explicit entry points that ask for a worktree (EnterWorktree, --worktree,
bridge worktree mode) still fail loudly. Deliberate divergence from
upstream, which hard-fails this path in both the Agent tool and the
workflow harness.
2026-08-18 00:37:50 +08:00
程序员阿江(Relakkes) 783f539b84 feat(desktop): add model-aware default reasoning effort 2026-08-16 16:35:20 +08:00
程序员阿江(Relakkes) 1d221f9f8d fix(desktop): open generated files with native apps #1233 2026-08-16 15:33:33 +08:00
程序员阿江(Relakkes) 2792406815 fix(provider): make Tool Search opt-in #1229 2026-08-16 15:32:27 +08:00
程序员阿江(Relakkes) b5d60a77ae fix(desktop): stabilize virtual row measurement #1223 2026-08-16 01:42:46 +08:00
程序员阿江(Relakkes) e2e302e1b3 fix(desktop): reconcile stale H5 task state #1214 2026-08-16 01:40:56 +08:00
程序员阿江(Relakkes) 7eb406243e fix(desktop): warn on proxy-only settings takeover #1219 2026-08-16 00:33:34 +08:00
程序员阿江(Relakkes) 5d7e1dcada fix(local-index): coalesce pending reconciliation work 2026-08-16 00:25:04 +08:00
程序员阿江(Relakkes) 5001502882 fix(desktop): recover safely from sidecar crashes #1227 2026-08-16 00:18:09 +08:00
程序员阿江(Relakkes) e72c55264f fix(rewind): prevent long-session checkpoint stalls 2026-08-15 23:59:18 +08:00
程序员阿江(Relakkes) d52bbec707 chore(release): prepare v0.5.4 v0.5.4 2026-08-14 18:11:23 +08:00
程序员阿江(Relakkes) 9f5c1951a2 fix(desktop): reuse model picker for agent settings 2026-08-14 16:43:44 +08:00
程序员阿江(Relakkes) 49cfd179fa docs: add Agent Teams and dynamic Workflow to READMEs 2026-08-14 16:01:04 +08:00
程序员阿江(Relakkes) a8c900b591 fix(auth): align Claude OAuth model and recovery 2026-08-14 14:59:44 +08:00