Route native-only changes through macOS checks and verify relocated cursor resources in final packages. Reject resources that escape the app and exercise visible click feedback against a disposable native receiver.
Connect the regressions to required checks and capture shell fixture output through temporary files.
Add a persistent isolated JavaScript worker for native app actions and batch
known operations without a model round trip between each input. Preserve
per-cell context, native errors, screenshot coordinates, and image types.
Align macOS gesture, key, inventory, capture, scroll, and clipboard behavior;
include a signed native receiver fixture and compiled sidecar regression tests.
Keep the Windows pixel route and fix cancellation with session-owned mouse
cleanup, lock revalidation, and portable signing-fixture tests.
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.
Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.
`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.
`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.
`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.
Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.
Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
A workflow is a JS script the model writes in the moment and hands to the
Workflow tool, which runs it in a locked-down `node:vm` and orchestrates
subagents through `agent()`/`parallel()`/`pipeline()`/`phase()`. Saving one
as a `/name` command is the secondary path; the inline script is the point.
Runtime: cross-realm value marshalling so a script cannot reach the host
`Function`, determinism guards on `Date.now()`/`Math.random()` (they would
make a resume replay diverge), a FIFO concurrency gate, and a journal that
lets an interrupted run resume from its longest unchanged prefix.
Desktop: the run shows up as a `workflow` section in the existing activity
panel — phases as headings, their agents beneath. A workflow agent is an
ordinary subagent run by the same runner, so its row opens the existing
subagent page rather than a parallel viewer of its own; that needed a
`by-agent` lookup, because these agents have no parent `Agent` tool call to
key off. Finished runs are rebuilt from the per-agent sidecars when a
session is reopened, since the live progress stream does not outlive the
process.
Rebuilt against main so the branch carries the Computer Use work and no other
divergence. Three unrelated efforts had been sitting uncommitted in this
worktree and were swept into an earlier commit; they are preserved on
cu-worktree-full-backup and belong on their own branches — adapter control
credentials, Electron asar sealing, and the sidecar code-loading audit. Every
file outside Computer Use now matches main exactly.
The engine
A Swift helper drives apps through the accessibility tree, with coordinate
actuation for the Chromium and Electron apps whose tree is a bare window
frame. Ten primitives matching the shape Codex uses, so an app's guidance and
the model's habits transfer.
Coordinate actions resolve their target window once and refuse when none can
be named. The unbound event they used to fall back to is discarded by custom
renderers, so a minimized target produced a whole session of "Action
completed" with nothing behind it.
Input acceptance is established for typing and key presses as well as clicks:
each MCP call is seconds apart, so the keyboard cannot inherit the focus a
click established. The synthetic focus notification is gated on the target
not already being active — sent unconditionally it names window 0 at an app
that already owns a key window, and nine window-bound clicks were discarded
with the traffic lights fully lit.
State the model can trust
An off-screen target says so, and says which tools still reach it: element
actions need no on-screen geometry, so an app with a real tree can still be
driven from the Dock. A fully covered window is recovered once, then left
alone — burying it again is the user wanting their screen back. A repeated
capture is reported with the cause that actually applies rather than both,
because coverage is something we compute.
Signing
The helper is signed under a stable identity before electron-builder sees it,
and excluded from re-signing: macOS ties Accessibility and Screen Recording
grants to the signing identity, so rotating it drops both on every update.
Discoverability
The desktop slash menu falls back to a directory scan while a session's CLI
has not started, which is when the menu is first opened. Built-ins and
bundled skills live in the binary, so /computer-use was absent until after
the first message.
Nothing checked src/ for unreferenced imports. desktop/tsconfig.json sets
noUnusedLocals and eslint covers desktop/ only; the root tsconfig.json sets
no such option and nothing installs typescript or bun-types at the root, so
no tool reads it at all. Splitting handler.ts left nine imports whose
symbols had moved out, and a human found them by reading the diff.
Measured before choosing. Under a temporary tsconfig extending the root
one, tsc reports 3225 errors over src/ and scripts/ before noUnusedLocals
and 3871 after — 646 net, on a baseline that already fails. That option
would land disabled, so this adds a narrow check instead: imports only,
one directory.
The analysis is lexical like module-graph.ts, but it blanks comments and
literal text first, so a symbol kept alive only by a comment still reports
dead — the exact shape the handler.ts split left behind. Template
substitutions, `//` inside a URL string and quotes inside a regex literal
must survive that blanking or a live import reads as dead;
src/utils/terminalShellEnvironment.ts is the last one and mis-lexing it
blanked 130 lines. Where blanking desyncs anyway the result no longer
parses, so the file is reported as degraded rather than mis-analysed: 6 of
2149 files repo-wide, none under src/server/ws.
src/server/ws/ also joins policyPrefixes. The check reads its files rather
than importing them, so the import graph cannot route a ws-only diff to
this lane, and without the prefix the check would never run on the diffs it
was written for.
Verified by mutation, each reverted from an explicit backup:
- planted `import { randomUUID }` into src/server/ws/events.ts →
check:policy 216 pass / 1 fail, reporting
"src/server/ws/events.ts:1 randomUUID"
- disabled line-comment blanking → 1 fail; block-comment blanking → 2 fail;
regex-literal detection → 2 fail (the extra one is the desync guard)
- pointed DEAD_IMPORT_ROOTS at a missing directory → 1 fail, so the check
cannot silently scan nothing
- dropped src/server/ws/ from policyPrefixes → 1 fail, and change-policy
reports policy=false for a ws-only diff
check:policy is 219 pass / 0 fail; src/server/__tests__/websocket-handler
.test.ts is 87 pass / 0 fail.
`check:agent-flow` proves the protocol with the mock CLI, which is what makes it
CI-safe and lets any contributor run it with no credentials. It cannot prove the
thing this product actually is: a desktop agent talking to a real model. That can
only run where the credentials are, so this lane is local and manual by
construction — registered in no quality-gate mode, referenced by no workflow, and
live.test.ts fails if either changes.
Six scenarios, sharing the existing harness rather than a second copy of it:
first turn, permission allow, permission deny, interrupt, reconnect, and history
recovery. Prompts induce the behaviour instead of dictating it, and assertions
only look at protocol shape and side effects on disk — never at generated text —
so the lane passes on any provider, including a local one. The three flows left
out (api-error, tool-error, runtime-select) each carry a written reason, because
a silently missing flow reads as a covered one.
Spending someone's quota is the failure mode worth engineering against, so the
runner refuses to guess: no implicit fallback to the active provider, an ambiguous
selector is an error rather than a pick, and without --yes it prints the provider,
model and config path it would use and exits without sending anything. User state
is copied into a throwaway config dir and the real ~/.claude is fingerprinted
before and after — a run that writes to it fails loudly instead of being cleaned
up quietly.
Not yet run end to end: the local LM Studio endpoint answers 502 here, so the six
runners have only been verified for structure. Target resolution, the
confirmation gate, and lane placement are covered by 14 tests that need no
provider at all.
Avoid Bun filter-mode repository scans that exhaust macOS file descriptors and corrupt subprocess test evidence. Apply rooted filters across server, contract, coverage, persistence, policy, desktop native, and adapter test entrypoints.
Confidence: high
Scope-risk: narrow
Tested: bun run check:policy; bun run check:server; bun run check:chat-contract
Run required server and contract suites in credential-free sandboxes, fail closed on incomplete coverage or test output, and preserve the desktop active-turn permission guard across stale tab interactions.\n\nTested: bun run check:policy (115 pass); bun run check:server (1605 pass before final runner evidence check); bun run check:desktop; bun run check:provider-contract; bun run check:chat-contract\nConfidence: high\nScope-risk: broad
Route required checks by changed surface, add offline provider and chat contracts, and keep fork PRs independent of live credentials. Layer agent guidance by subtree and enforce a compact instruction budget.
Tested: bun run check:policy
Confidence: high
Scope-risk: broad
Introduce the Electron desktop shell alongside the existing React renderer and local Bun server boundary. The migration keeps the DesktopHost contract explicit across Tauri, Electron, and browser runtimes while adding Electron main/preload services for dialogs, shell, notifications, updates, tray/window lifecycle, terminal, preview WebContentsView, app mode, and release/package validation.
The commit also carries the latest local main desktop command updates, including agent slash entries and hidden-by-default markdown thinking details, so the packaged Electron build matches the current main UX surface.
Constraint: React renderer, local Bun server, REST/WebSocket, and sidecar boundaries must remain reusable during the migration
Constraint: macOS dev packages are ad-hoc signed and cannot prove Developer ID notarization or Gatekeeper release launch
Rejected: Browser-only smoke validation | it cannot exercise native dialogs, keychain prompts, notification behavior, or packaged app startup
Confidence: medium
Scope-risk: broad
Directive: Do not remove Tauri host support until signed Electron release artifacts pass native OS smoke on macOS, Windows, and Linux
Tested: bun run check:desktop
Tested: cd desktop && bun run check:electron
Tested: CSC_IDENTITY_AUTO_DISCOVERY=false bun run electron:package:dir
Tested: bun run test:package-smoke --platform macos --package-kind dir --artifacts-dir desktop/build-artifacts/electron
Tested: Computer Use read packaged Electron app window at desktop/build-artifacts/electron/mac-arm64/Claude Code Haha.app
Not-tested: Developer ID signed/notarized Gatekeeper launch
Not-tested: Real OS notification click-to-session action
Not-tested: Windows and Linux packaged app smoke on real hosts
Daily pushes should catch policy and path-aware local failures without making every contributor wait for full coverage. The hook now runs a new quality:push entrypoint that reuses the PR quality gate while skipping coverage; verify, quality:pr, and CI still retain the full coverage gate for PR readiness.
Constraint: Forks and local contributors need a faster default push path
Rejected: Remove coverage from quality:pr | PR readiness and CI still need the ratcheted coverage signal
Confidence: high
Scope-risk: narrow
Directive: Keep pre-push on quality:push; use verify or quality:pr when coverage evidence is required
Tested: bun test scripts/pr/quality-contract.test.ts scripts/git-hooks/install.test.ts
Tested: bun run check:policy
Tested: bun run quality:push
Not-tested: Live provider smoke; intentionally remains opt-in
Desktop users can carry provider indexes, managed settings, localStorage state, and native update state from builds that no longer match current readers. This adds startup migrations and recovery paths before server and React state are consumed, plus a persistence upgrade gate so future storage protocol changes ship with old-format fixtures.
Constraint: Existing installs may contain malformed or legacy JSON/localStorage that must not block startup.
Constraint: Local verify should evaluate the current worktree diff rather than unrelated detached-worktree history.
Rejected: Treat invalid persisted state as fatal | reproduces white-screen and startup failure behavior for existing users.
Rejected: Bypass PR policy locally | hides real gate behavior and does not fix detached-worktree false positives.
Confidence: high
Scope-risk: moderate
Directive: Any local JSON, localStorage, or app config shape change must add a migration fixture and keep `bun run check:persistence-upgrade` green.
Tested: bun run check:persistence-upgrade; bun run check:policy; bun run check:desktop; bun run check:server; bun run check:native; bun run verify (9 passed, 1 coverage baseline failure)
Not-tested: Live provider baseline; existing user configs beyond covered fixtures
Contributors and coding agents need one local command that both reports and enforces the quality contract. This change turns the PR gate into the shared verification entrypoint, adds path-selected local lanes, tightens coverage accounting around changed lines, and documents the repair loop in contributor and agent-facing guidance.
Constraint: Ordinary PR verification must stay non-live and runnable without provider credentials
Constraint: Coverage policy updates in this commit require maintainer approval before push/merge
Rejected: Keep quality guidance only in docs | agents need executable scripts and AGENTS.md instructions to follow the loop consistently
Confidence: high
Scope-risk: broad
Directive: Do not bypass `bun run verify` for production changes; fix failed lanes and coverage reports instead of lowering thresholds
Tested: bun run check:policy
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run verify
Not-tested: live provider baseline; no provider credentials were required for this non-live PR gate
The repository now has a measurable PR quality path instead of a loose set of
manual checks. Coverage, quarantine governance, provider smoke, desktop smoke,
and workflow wiring all produce durable reports that contributors and maintainers
can inspect without reconstructing terminal output.
This also fixes the desktop smoke current-runtime path so browser-driven smoke
runs use the desktop default active provider instead of forcing the official
current model, and records that runtime decision as an artifact.
Constraint: Default PR gates must remain non-live and contributor-safe while live model checks stay explicit.
Constraint: Release packaging is still GitHub Actions based, so release preflight must run before the build matrix.
Rejected: Make live provider or desktop smoke mandatory on every PR | secrets, quotas, and model availability are maintainer-controlled.
Rejected: Let PRs lower coverage baselines in the same change | base-branch ratchet comparison must remain authoritative.
Confidence: high
Scope-risk: moderate
Directive: Do not relax coverage or quarantine policy without a maintainer approval label and a fresh quality report.
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run quality:gate --mode pr
Tested: bun run quality:gate --mode baseline --allow-live --only provider-smoke:* --provider-model nvidia-custom:main:nvidia-custom-main --artifacts-dir /tmp/quality-gate-live-smoke
Tested: bun run quality:gate --mode baseline --allow-live --only desktop-smoke:* --provider-model current:current:current-runtime --artifacts-dir /tmp/quality-gate-desktop-smoke-fixed
Tested: git diff --check
Not-tested: Full live release mode with multiple providers in hosted CI; provider credentials and quota remain maintainer-controlled.
This captures the pending worktree fixes before applying them to the
current local main. The changes tighten IM adapter path and credential
handling, preserve retry behavior for failed desktop notifications, and
make Azure/OpenAI provider auth and stop reasons reflect actual runtime
state.
Constraint: Worktree was detached from an older local main with pending uncommitted fixes
Rejected: Merge the detached HEAD directly | would also replay unrelated stale history
Rejected: Leave notification dedupe as fire-and-forget | failed sends consumed retry keys
Confidence: high
Scope-risk: broad
Directive: Keep adapter absolute-path matching constrained to configured work roots
Tested: git diff --check
Not-tested: full quality gate before local main integration
The PR triage workflow mentioned Dosu inside inline-code formatting and before the final footer, which did not reliably wake the bot. Move the handoff so the last non-empty comment line is a plain-text @dosubot request, matching the working PR template pattern.
Constraint: GitHub bot mentions can be sensitive to markdown formatting and comment placement.
Rejected: Leave the mention as a copy-paste hint | it does not satisfy the maintainer need for automatic bot triggering.
Confidence: high
Scope-risk: narrow
Directive: Keep the generated triage comment's final non-empty line as a plain-text @dosubot request.
Tested: bun run check:policy
Tested: git diff --check
Tested: bun run check:pr
The live baseline previously accepted provider UUIDs, which made the gate hard to run on another contributor's machine. Add a local provider listing command and resolve quality-gate targets from stable provider-name selectors while keeping UUIDs and current runtime support.
Constraint: Provider configuration is local machine state under CLAUDE_CONFIG_DIR and must not expose API keys.
Rejected: Require contributors to inspect providers.json manually | too error-prone and leaks implementation detail into the workflow
Confidence: high
Scope-risk: narrow
Directive: Keep live-provider baseline selection copyable from quality:providers before adding more live test lanes.
Tested: bun test scripts/quality-gate/providerTargets.test.ts
Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts
Tested: bun run quality:providers
Tested: bun run quality:gate --mode baseline --dry-run --provider-model codingplan:main --provider-model minimax:main
Tested: bun run quality:gate --mode baseline --dry-run --provider-model custom:haiku
Tested: bun run quality:gate --mode pr --dry-run
Tested: bun run check:server
The desktop product needs a repeatable local gate that can prove the core Coding Agent loop still works after changes, not only that unit tests pass. This adds a quality-gate runner with PR, baseline, and release modes, structured reports, explicit quarantine metadata, and fixture-based live baseline cases that can run across provider/model targets.
Constraint: Existing check:pr and CI policy behavior must remain usable while the stronger baseline grows around it
Constraint: Default PR gates must not require real model credentials or provider quota
Rejected: Build a standalone QA platform first | too heavy before the baseline task shape is proven
Rejected: Keep unstable server exclusions hardcoded in run-server-tests | hides quarantine policy from reports and future review
Confidence: medium
Scope-risk: moderate
Directive: Expand baseline cases by adding focused fixtures and verifiers; do not make normal PR checks depend on live providers
Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts
Tested: bun run check:server
Tested: bun run quality:gate --mode baseline --allow-live --provider-model 2944f963-ce75-45b7-bac1-6e4f57df0970:kimi-k2.6:volc-kimi-k2.6 --provider-model 9c78d3df-7fb5-44c7-8436-3a41c3a59231:MiniMax-M2.7-highspeed:minimax-m2.7
Not-tested: desktop UI browser smoke and native release mode in this commit
Pull requests need a deterministic way to show changed areas, required checks, missing-test signals, and CLI-core risk before review. This adds a path-based policy gate, local impact reporting, PR triage labels/comments, and reusable check scripts so reviewers can evaluate blast radius without trusting contributor claims.
Constraint: CLI core should remain effectively frozen unless a maintainer explicitly overrides it.
Constraint: Default PR checks must be safe for external forks and avoid live model/provider calls.
Rejected: Run live provider tests on every PR | secrets, cost, network, and vendor instability would make the gate noisy and unsafe.
Rejected: Use Dosu as the merge gate | AI review is useful for risk explanation, but deterministic Actions must own blocking checks.
Confidence: high
Scope-risk: moderate
Directive: Keep real model/provider smoke tests in maintainer-controlled workflows; do not make them required for untrusted PRs.
Tested: bun run check:impact
Tested: bun run check:policy
Tested: ruby YAML parse for PR workflows
Tested: git diff --check
Tested: bun run check:native
Tested: npm run docs:build
Tested: bun run scripts/pr/run-server-tests.ts
Tested: bun run check:adapters
Tested: bun run check:desktop
Not-tested: GitHub-hosted pull_request_target label/comment execution before opening this PR
Desktop sessions were failing before the actual fetch because Anthropic's domain preflight can be unreachable on restricted networks, and the next runtime path was missing turndown for HTML-to-Markdown conversion.
This change defaults desktop sessions to skip the preflight unless the user explicitly overrides it, exposes that behavior as a desktop General setting, seeds new settings JSON with the desktop-safe default, and adds regression coverage for both the runtime default and the UI toggle. It also adds the missing turndown dependency so successful fetches can continue through HTML reduction instead of failing at module resolution.
Constraint: Desktop must keep an escape hatch for users who want upstream preflight restored explicitly
Rejected: Force skipWebFetchPreflight globally for every session | would silently change CLI and non-desktop behavior
Rejected: UI-only toggle without runtime default | existing desktop users would still fail until they manually opened settings
Confidence: high
Scope-risk: moderate
Reversibility: clean
Directive: Keep desktop-specific WebFetch behavior scoped to desktop session detection and explicit user settings; do not broaden it to general CLI flows without separate validation
Tested: bun test src/tools/WebFetchTool/utils.test.ts; cd desktop && bun run lint; cd desktop && bunx vitest run src/__tests__/generalSettings.test.tsx; runtime import verification for turndown via node
Not-tested: End-to-end desktop packaging smoke test against a freshly built DMG/app bundle