Add mobile provider and General settings while preserving desktop behavior.
Harden remote credential handling, ngrok session ownership and consent upgrades.
Add a persistent isolated JavaScript worker for native app actions and batch
known operations without a model round trip between each input. Preserve
per-cell context, native errors, screenshot coordinates, and image types.
Align macOS gesture, key, inventory, capture, scroll, and clipboard behavior;
include a signed native receiver fixture and compiled sidecar regression tests.
Keep the Windows pixel route and fix cancellation with session-owned mouse
cleanup, lock revalidation, and portable signing-fixture tests.
Add shared project and session selection across IM adapters, preserve bindings on
failed restoration, and synchronize permissions with desktop clients.
Fixes#1286
Connecting Feishu meant creating a bot by hand on the open platform and
pasting an App ID and App Secret back. Feishu also exposes an RFC 8628
device-authorization flow, so the desktop can now render a QR code, and
confirming it in the app creates the bot and stores its credentials
directly. `adapters/feishu/registration.ts` implements that protocol
rather than importing `registerApp` from `@larksuiteoapi/node-sdk@1.73`:
the repository pins 1.60 for the chat client, and the SDK runs the whole
poll inside one un-cancellable promise where the desktop needs the
stateless begin/poll pair the DingTalk registration already uses. The
scan is create-only, so it can never rewrite the configuration of a bot
the user already runs, and it pre-fills exactly the scopes, events and
callbacks this adapter calls. International tenants finish on Lark's
domain, which is now persisted and honoured by the client.
WeCom, QQ and Slack join the same session model. WeCom and QQ bind by
scanning; Slack has no scan flow, so it uses an app manifest that
pre-fills the scopes and Socket Mode. All three run over long
connections, so no public callback URL is needed, and all three accept
private chats only — pairing authorizes one person, and answering in a
group would extend that authorization to everyone else in the room.
They are built on a new `adapters/common/chat-runtime.ts` instead of a
fourth copy of the loop the five existing adapters each carry. A platform
supplies a `ChatPort` — how to say something, how to open a streaming
reply, optionally how to send an image — and the runtime owns pairing,
command routing, session restore, permission bookkeeping and the
translation of the server's stream. The existing five are deliberately
left on their own copies; migrating them is a separate change with its
own regression surface.
Attachments are downloaded through a deferred loader that runs after the
pairing gate and inside the per-chat queue. Resolving them eagerly would
let an unpaired stranger make the adapter fetch bytes and write them
under ~/.claude/im-downloads — on Slack with the bot token attached —
and would let a slow attachment overtake a text message sent after it.
The sidecar launcher's per-adapter branches become one table. It is
declared above the mode dispatch on purpose: `runAdapters` is hoisted and
runs at module top level, so a table declared below it is still in its
temporal dead zone when the adapters mode reads it — which type checks,
lints and unit tests all miss, and only the compiled binary reveals.
Verified with the checks `check:impact` selects: adapters, server,
desktop, electron, policy, chat-contract, agent-flow, docs, native
(sidecar compile, packaging and an adapters-mode smoke against the real
binary) and coverage. The scan flows themselves are not verified against
live platforms — that needs real WeCom, QQ and Slack accounts and would
create real bots.
Claude-Session: https://claude.ai/code/session_01CCGoP316AK7wdQG3Ms6Uwq
The instruction files told every coding agent to reach for agent-browser
whenever a change needed browser-level evidence: copilot-instructions
listed "E2E or agent-browser smoke" as the remedy for cross-boundary
flows, and both contributing guides repeated it. That wording outlived
the tool. With the agent-browser skill uninstalled, agents still parsed
those lines as a recommendation and went looking for the binary instead
of using the browser skill that is actually installed.
Deleting the references would have made the docs wrong. agent-browser is
still a real dependency: check:desktop-ui-smoke spawns it on Linux CI,
and seven maintainer-run e2e scripts under desktop/scripts drive it
directly. It cannot be swapped for ego-browser either — ego lite is a
macOS-only GUI app with no headless mode and a one-time interactive
onboarding, so it cannot run on ubuntu-latest at all.
So the lanes keep the binary and the prose loses the recommendation.
agent-browser is now described as an implementation detail of those two
call sites, and ad-hoc browser work — manual verification, screenshots,
exploratory UI checks — is pointed at the ego-browser skill.
The quality contract asserted the old string, so it would have failed
closed on the reworded line. It now pins the replacement plus the new
routing rule; flipping either sentence turns the test red.
Built-in agents pin their own models — Explore and claude-code-guide run
on Haiku, statusline-setup on Sonnet — and there was no way to change
that. The only override mechanism was a same-named user agent, which
replaces the definition wholesale: AGENT_SLUG_PATTERN is lowercase-only
so `Explore` and `Plan` cannot even be created, and a replacement loses
the runtime getSystemPrompt and the built-in tool privileges keyed off
`source`. CLAUDE_CODE_SUBAGENT_MODEL is the only other lever and forces
every subagent onto one model.
Add `builtInAgentOverrides` to settings.json, carrying model and effort
per agentType. It is applied in getBuiltInAgents(), the single choke
point every consumer goes through, so the effective value reaches
spawning, `/agents` and the desktop list without any of them knowing an
override exists — and `source` stays `built-in`, preserving the prompt
and tool privileges.
Three constraints are load-bearing rather than stylistic:
- `model: "inherit"` is a real value, not a reset. Built-in defaults
differ per agent and per build, so clearing means deleting the field;
the entry and then the key are removed once empty.
- getSystemPrompt is never wrapped. serializeActiveAgent branches on
`.length === 0`, and claude-code-guide declares one parameter while
the others declare none, so a wrapper makes it destructure undefined.
- The settings write is a read-modify-write inside the file lock.
updateUserSettings is a top-level shallow merge and would replace the
whole record, losing an update when two agents are changed in quick
succession.
strictPluginOnlyCustomization is enforced when resolving, not only when
writing, since settings.json is user-editable by definition.
Also surface edit and delete on the agent list rows. Both already
existed but were reachable only after opening an agent's detail page.
Rows become a div with the primary button and the actions as siblings
rather than nested buttons, keep focus-within so the controls are not
tabbable while invisible, and stay visible on touch where hover never
fires. Built-in rows get the override entry point and no delete.
The store's create/update/override paths now share runAgentMutation
instead of a third hand-copy of the out-of-order guards; the existing
store tests pass unchanged.
restoreAvailable conflated two questions: whether the files this checkpoint
reports can be put back, and whether the checkpoint saw every file the turn
touched. Any tool outside an 18-name allowlist — Bash, PowerShell, TaskCreate,
every MCP tool — forced the second to false, and one such call anywhere from
the target turn onward disabled undo for the whole range. In practice that is
every real turn, so v0.5.3 shipped with undo effectively dead.
Split the two. restoreAvailable now answers only the first question.
The second becomes unverifiedChangeSources: tool names whose file effects the
checkpoint could not capture. Undo stays available and restores exactly the
files it lists, and the card, the confirmation, and the completion toast each
name what it is leaving behind. An unrecognized tool now costs a warning
instead of the feature. A transcript that cannot be read still blocks, because
then even the reported file list may be wrong.
Bash calls are classified against the existing read-only allowlist via a new
recordedCommandIsReadOnly, which drops the sandbox/cwd checks that describe the
live process rather than the replayed session. `git status` no longer warns at
all, so the warning means something when it appears.
Rewind also takes a mode. `conversation` skips the file restore entirely, so a
turn whose files cannot be restored no longer costs the user the ability to
back out of the prompt — matching how upstream keeps Restore conversation
independent of Restore code.
Files written by shell commands were never recoverable here; the checkpoint
only ever covered the structured file tools. Reporting that is honest, and
docs now say so.
One conflict, same structural cause as the last merge: git offers the pre-split
Settings.tsx against the 183-line shell. All four of main's hunks belong to
ProviderFormModal, which now lives in settings/ProviderSettings.tsx — the useId
import, the addToast on a successful model fetch, and the base-URL field growing
an explicit label plus a help tooltip. Ported there; tsc caught the one import
(useUIStore) the move needed.
Everything else auto-merged, including the two files both sides changed:
chatStore.ts keeps main's pushAssistantHistoryThinking alongside the removal of
the content-equality replay guard — different functions, history mapping versus
the live path — and the five locales land at 2521 keys each.
Note: six generalSettings tests fail after this merge and they fail identically
on main. 62f648cb2 added an IconButton labelled 'Base URL help' next to the base
URL input without updating generalSettings.test.tsx, so getByLabelText(/Base
URL/i) now matches both. Confirmed against a temporary worktree at main: 6
failed / 104 passed there too. Not introduced here, and fixed separately.
4f9fec876 added this workflow with `cron: '0 18 * * *'`. That was the wrong call
to make unilaterally: the repository had no scheduled workflow at all before it,
so this was not one more cron among several but the introduction of recurring CI
spend — about ninety minutes per run — on a schedule nobody asked for.
The reasoning for the sweep still holds: a per-PR gate only covers what the diff
reaches, so it is blind to checks no recent PR selected and to failures that only
appear when the whole suite runs together. Keeping the workflow on
`workflow_dispatch` keeps that one click away without deciding for the maintainer
when to spend the time.
pr-quality-workflow.test.ts now asserts the absence of `schedule:` and `cron:`
rather than their presence, so a schedule cannot drift back in unnoticed —
verified by adding the cron back and watching the test go red. The docs' four-tier
table renames the tier accordingly; calling it "Nightly" when nothing runs nightly
is exactly the kind of comment that outlives its code.
- Add both gateways as featured presets pointing at their Anthropic-compatible
roots (https://api.fenno.ai, https://api.qnaigc.com) with auth_token strategy
- Ship no default model ids: the available catalog depends on the plan the user
bought, so they fetch the live list and pick one instead
- Keep modelContextWindows, which stays useful per picked model id and covers
ids the built-in table cannot resolve (claude-opus-5, namespaced 七牛云 ids)
- Add both as sponsors in the English and Chinese READMEs, and list them in the
preset docs
The site had drifted from the product. Every screenshot predated the
v0.5.0 UI redesign, the reading experience shipped no search and no
syntax highlighting, and a third of the pages were internal process
artefacts — migration task lists addressed to agentic workers, a
release runbook, a proposal marked "historical".
Reorganise around the only two people who read this: someone getting
the desktop app running for the first time, and someone reading the
source. Five sections replace nine — start / desktop / im / cli /
internals — and the pages that served neither reader are gone.
Site rewrite:
- Palette lifted from the desktop app's 「纸·墨·印」 themes, so the
site and the product read as one thing. Light mirrors 纯白, dark
mirrors 墨夜, and dark mode exists at all now.
- Fonts are self-hosted. The old @import from Google Fonts is
unreachable from mainland China, which left every heading in a
fallback serif; it also only requested weight 600 while the CSS
asked for 900, so Latin and CJK in the same heading disagreed.
- Docs were shipped as one 968KB manifest downloaded on every page
view. Split into a 32KB index plus one lazily imported chunk per
page; the entry bundle is now 101KB gzipped.
- Add search, syntax highlighting, per-route meta with canonical and
hreflang, a sitemap, and an error boundary. Replace the 44vh
mobile sidebar with a drawer.
- Image dimensions are read at build time and written into the tag,
so lazy images reserve their space instead of collapsing.
Screenshots are recaptured from a real v0.5.0 build against a clean
demo project, with tokens, QR codes and paired accounts redacted.
The previous set is deleted rather than kept alongside.
Routes follow file paths, so the restructure would have broken every
inbound link; 37 old paths redirect, in both languages. The PR policy
gate and CODEOWNERS also hardcoded docs/guide/contributing.md.
Verified: check:docs 78 pages / 323 links / 0 problems, check:policy
127 pass. Walked every route at 1440 and 390 in both themes for
overflow, contrast, keyboard reachability and focus management.
Importing an animated pet required a file that was exactly 1536x2288, laid
out as 88 seamless cells, with the last two rows holding sixteen distinct
gaze angles. No image model emits that. Whatever a user got back from Jimeng
or ChatGPT was some fixed size like 1024x1536, so the path ended at "the
animation atlas must be exactly 1536x2288 pixels" every time. The third card
was worse: "AI-generate full animation" was hardcoded `disabled`, so the one
entry point named after what people actually wanted to do was dead.
The fix was already in the tree. `scripts/assemble-generated-pet-atlas.py`
landed in the same commit as the four built-in pets, which is to say the
built-ins were produced this way — it takes an action sheet at any size,
slices it on an 8x9 grid, fits each cell to 192x208, mirrors the run row to
make run-left, and reuses rows to reach eleven. That capability was never
wired to anything a user could reach.
`petAtlasNormalize.ts` reimplements it on a canvas in the renderer, so an
author draws nine rows and the app derives the rest. Verified against the
reference assembler by reversing dada-code's atlas into a nine-row sheet and
re-normalizing it: every difference lands on semi-transparent antialiased
edges (2314 pixels, max channel delta 14/255) and opaque regions are
identical. That residue is canvas premultiplied-alpha round-tripping, not a
slicing bug.
Three contract details worth stating. Row frame counts are now derived from
`PET_ANIMATION_DEFINITIONS` rather than typed out a fourth time; they come
out equal to the assembler's `(6,8,8,4,5,8,6,6,6,8,8)`. A sheet already at
1536x2288 passes through byte-for-byte instead of being resliced, because
resampling finished artwork buys nothing. And since the validator never
inspects the alpha channel, a flattened white background used to import
happily and render as a rectangle on the desktop — the renderer now rejects
sheets whose atlas is under 5% transparent (the built-ins sit near 78%) with
a message that names the actual problem.
The copy stops describing the implementation. "Animate one image" and
"Import professional animation atlas / exact 1536x2288 v2 PNG" become "use a
picture you already have" and "I already have an action sheet"; the dead AI
card becomes a three-step walkthrough carrying a copyable prompt, a labelled
8x9 reference grid that can be saved locally, and the checks that catch the
common failures. Reference images are generated by a script rather than hand-
placed, in both languages. All five locales move together.
Caught while reviewing the real dialog in Electron: after finishing the
walkthrough the form heading fell through to the atlas branch and announced
"I already have an action sheet" to someone who had just been walked through
drawing one. Covered by a test now.
Not done: docs/images/desktop_ui/15_pet_create_methods.png still shows the
old dialog and needs a fresh capture from a running app to match the styling
of the shots around it.
Skills installed by Codex, Cursor, Gemini CLI, or opencode live under
`.agents/skills/` -- the convention documented at agentskills.io -- so users
had to keep a duplicate copy in `.claude/skills/` before cc-haha could see
them. The spec only constrains SKILL.md itself, which we already parse, so
this is purely a directory-discovery gap.
Scan both conventions at every scope: user (`~/.agents/skills`), project
(every level up to the repo root), `--add-dir` roots, on-demand monorepo
discovery, and the hot-reload watcher. skillRoots.ts is the single source of
truth for those paths; the REST API now reports which convention a skill came
from and the desktop list badges the `.agents` ones.
`.claude` is always ordered first, so first-wins dedup keeps a user's own
skills authoritative rather than letting an externally installed same-named
skill shadow them. Name collisions are collapsed per scope, which covers
copy-based sync -- symlink-based sync was already handled by the existing
realpath dedup. Collisions across different scopes, such as a user skill and
a project skill sharing a name, keep their existing behavior, and the shadowed
file's identity is recorded so overlapping roots (CLAUDE_CONFIG_DIR pointing
inside the project) cannot reintroduce it under the other scope.
getProjectDirsUpToHome is untouched -- it is shared by commands, agents,
output-styles, and workflows. Only its directory walk is extracted so the
skill resolver reuses the same submodule/worktree stop boundary. Writes stay
where they were: market installs and /skillify still target
`~/.claude/skills`.
`.agents` is anchored at $HOME rather than CLAUDE_CONFIG_DIR, since it is a
space shared with other tools that read $HOME. Note that Bun snapshots
os.homedir() at startup while Node re-reads $HOME per call, so the resolver
reads the env var explicitly.
Opt out with `disableAgentSkillsDirectory` in settings.json or
CLAUDE_CODE_DISABLE_AGENT_SKILLS_DIR=1; no `.agents` root is produced at all
then, so there is not even a failed stat.
(cherry picked from commit ba9338f85d6beaba237753e42e823c87b59633d4)
Route required checks by changed surface, add offline provider and chat contracts, and keep fork PRs independent of live credentials. Layer agent guidance by subtree and enforce a compact instruction budget.
Tested: bun run check:policy
Confidence: high
Scope-risk: broad
Replace the local pre-push quality gate with a reminder-only hook, keep manual quality commands documented, and update the contract test to prevent reintroducing a blocking push gate.
Tested: bash .git/hooks/pre-push </dev/null
Tested: bun test scripts/git-hooks/install.test.ts scripts/pr/quality-contract.test.ts
Tested: bun run check:policy
Tested: bun run check:docs
Tested: git diff --check
Confidence: high
Scope-risk: narrow
MiniMax's official M3 model page and release notes describe M3 as supporting up to a 1M-token context window with a guaranteed minimum of 512K. Keep the built-in MiniMax preset aligned with that official default instead of inheriting the older 204.8K MiniMax generation limit.
Constraint: Official MiniMax M3 docs state the M3 API supports up to 1M tokens context window.
Rejected: Keep 204800 and rely on manual provider overrides | that makes the built-in MiniMax preset wrong for new users.
Confidence: high
Scope-risk: narrow
Directive: Do not copy MiniMax-M2.7's 204.8K window onto MiniMax-M3 without re-checking official M3 docs.
Related: https://www.minimax.io/models/text/m3
Related: https://www.minimax.io/blog/minimax-m3
Tested: git diff --check
Tested: bun test src/server/__tests__/provider-presets.test.ts src/server/__tests__/providers-real.test.ts src/server/__tests__/provider-runtime-env.test.ts src/server/__tests__/providers.test.ts src/utils/__tests__/context.test.ts
Not-tested: bun run check:server blocked before execution by expired quarantine entries: server:cron-scheduler, server:providers-real, server:tasks, server:e2e:business-flow, server:e2e:full-flow
Remote main includes PR #694, which updates the MiniMax preset to MiniMax-M3 and removes deprecated built-in MiniMax context-window entries. Merge that provider baseline before applying the transcript context-window estimate fix so the local line keeps upstream provider defaults.
Constraint: Remote main advanced independently with PR #694 while local main carried unreleased commits.
Rejected: Reset local main to origin/main | would discard local-only main commits.
Confidence: high
Scope-risk: moderate
Directive: Preserve MiniMax-M3 provider preset changes when touching provider context-window logic.
Tested: bun test src/server/__tests__/conversations.test.ts src/server/__tests__/conversation-service.test.ts src/server/__tests__/providers.test.ts src/server/__tests__/provider-presets.test.ts src/server/__tests__/provider-runtime-env.test.ts src/utils/__tests__/context.test.ts src/utils/__tests__/contextBudget.test.ts
Not-tested: bun run check:server blocked by expired quarantine entries: server:cron-scheduler, server:providers-real, server:tasks, server:e2e:business-flow, server:e2e:full-flow
Add MiniMax-M3 as the new default for the minimax provider preset and
remove deprecated M2.5/M2.1/M2 entries. MiniMax-M2.7 and the highspeed
variant are kept for users who pinned the previous generation.
Updated places:
- providerPresets.json: default models -> M3, modelContextWindows trimmed
- modelContextWindows.ts: built-in window list mirrors the new set
- provider-presets / providers-real / thinking tests updated to assert
the new default
- .env.example and third-party docs (zh + en) recommend M3
- scripts/repro-issue-247-real.ts default model bumped to M3
Co-Authored-By: Octopus <liyuan851277048@icloud.com>
Introduce the Electron desktop shell alongside the existing React renderer and local Bun server boundary. The migration keeps the DesktopHost contract explicit across Tauri, Electron, and browser runtimes while adding Electron main/preload services for dialogs, shell, notifications, updates, tray/window lifecycle, terminal, preview WebContentsView, app mode, and release/package validation.
The commit also carries the latest local main desktop command updates, including agent slash entries and hidden-by-default markdown thinking details, so the packaged Electron build matches the current main UX surface.
Constraint: React renderer, local Bun server, REST/WebSocket, and sidecar boundaries must remain reusable during the migration
Constraint: macOS dev packages are ad-hoc signed and cannot prove Developer ID notarization or Gatekeeper release launch
Rejected: Browser-only smoke validation | it cannot exercise native dialogs, keychain prompts, notification behavior, or packaged app startup
Confidence: medium
Scope-risk: broad
Directive: Do not remove Tauri host support until signed Electron release artifacts pass native OS smoke on macOS, Windows, and Linux
Tested: bun run check:desktop
Tested: cd desktop && bun run check:electron
Tested: CSC_IDENTITY_AUTO_DISCOVERY=false bun run electron:package:dir
Tested: bun run test:package-smoke --platform macos --package-kind dir --artifacts-dir desktop/build-artifacts/electron
Tested: Computer Use read packaged Electron app window at desktop/build-artifacts/electron/mac-arm64/Claude Code Haha.app
Not-tested: Developer ID signed/notarized Gatekeeper launch
Not-tested: Real OS notification click-to-session action
Not-tested: Windows and Linux packaged app smoke on real hosts
Daily pushes should catch policy and path-aware local failures without making every contributor wait for full coverage. The hook now runs a new quality:push entrypoint that reuses the PR quality gate while skipping coverage; verify, quality:pr, and CI still retain the full coverage gate for PR readiness.
Constraint: Forks and local contributors need a faster default push path
Rejected: Remove coverage from quality:pr | PR readiness and CI still need the ratcheted coverage signal
Confidence: high
Scope-risk: narrow
Directive: Keep pre-push on quality:push; use verify or quality:pr when coverage evidence is required
Tested: bun test scripts/pr/quality-contract.test.ts scripts/git-hooks/install.test.ts
Tested: bun run check:policy
Tested: bun run quality:push
Not-tested: Live provider smoke; intentionally remains opt-in
The desktop app and documentation need to present the new logo consistently, so the canonical 1024px source icon now feeds the Tauri icon family, desktop public assets, README images, and VitePress brand imagery.
Constraint: Tauri bundle icons are generated assets and must be refreshed from the source icon rather than relying on README or public image replacement alone
Rejected: Rename logo paths | existing README, desktop, and docs references can keep stable paths and pick up the new files directly
Confidence: high
Scope-risk: moderate
Directive: Regenerate desktop/src-tauri/icons whenever desktop/src-tauri/app-icon.png changes
Tested: cd desktop && bun run build; npm run --loglevel=error docs:build; git diff --check; iconutil expanded icon.icns size ladder
Not-tested: full local quality gate per request
Contributors and coding agents need one local command that both reports and enforces the quality contract. This change turns the PR gate into the shared verification entrypoint, adds path-selected local lanes, tightens coverage accounting around changed lines, and documents the repair loop in contributor and agent-facing guidance.
Constraint: Ordinary PR verification must stay non-live and runnable without provider credentials
Constraint: Coverage policy updates in this commit require maintainer approval before push/merge
Rejected: Keep quality guidance only in docs | agents need executable scripts and AGENTS.md instructions to follow the loop consistently
Confidence: high
Scope-risk: broad
Directive: Do not bypass `bun run verify` for production changes; fix failed lanes and coverage reports instead of lowering thresholds
Tested: bun run check:policy
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run verify
Not-tested: live provider baseline; no provider credentials were required for this non-live PR gate
The repository now has a measurable PR quality path instead of a loose set of
manual checks. Coverage, quarantine governance, provider smoke, desktop smoke,
and workflow wiring all produce durable reports that contributors and maintainers
can inspect without reconstructing terminal output.
This also fixes the desktop smoke current-runtime path so browser-driven smoke
runs use the desktop default active provider instead of forcing the official
current model, and records that runtime decision as an artifact.
Constraint: Default PR gates must remain non-live and contributor-safe while live model checks stay explicit.
Constraint: Release packaging is still GitHub Actions based, so release preflight must run before the build matrix.
Rejected: Make live provider or desktop smoke mandatory on every PR | secrets, quotas, and model availability are maintainer-controlled.
Rejected: Let PRs lower coverage baselines in the same change | base-branch ratchet comparison must remain authoritative.
Confidence: high
Scope-risk: moderate
Directive: Do not relax coverage or quarantine policy without a maintainer approval label and a fresh quality report.
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run quality:gate --mode pr
Tested: bun run quality:gate --mode baseline --allow-live --only provider-smoke:* --provider-model nvidia-custom:main:nvidia-custom-main --artifacts-dir /tmp/quality-gate-live-smoke
Tested: bun run quality:gate --mode baseline --allow-live --only desktop-smoke:* --provider-model current:current:current-runtime --artifacts-dir /tmp/quality-gate-desktop-smoke-fixed
Tested: git diff --check
Not-tested: Full live release mode with multiple providers in hosted CI; provider credentials and quota remain maintainer-controlled.
Computer Use is useful when explicitly needed, but exposing its MCP tools by default creates unnecessary desktop-control surface for users who want coding-only sessions. This adds a shared disable path for CLI flags, environment, and desktop settings while keeping preauthorized app state in one config file.
The same change also preserves Windows and WSL shell startup behavior by applying the MSYS argument-conversion guard only on WSL-bound launches.
Constraint: Computer Use MCP must not be exposed to the Coding Agent when disabled
Constraint: Desktop settings and CLI sessions need to read the same persisted Computer Use config
Rejected: Environment-only disable switch | desktop users need a persistent Settings control
Rejected: Remove Computer Use setup entirely | enabled sessions still need the existing built-in MCP path
Confidence: high
Scope-risk: moderate
Directive: Keep every new Computer Use entrypoint wired through loadStoredComputerUseConfig or the CLI disable flag before adding MCP tools
Tested: bun test src/utils/computerUse/gates.test.ts src/utils/computerUse/preauthorizedConfig.test.ts src/server/__tests__/computer-use-api.test.ts src/utils/shell/wslInterop.test.ts desktop/src/pages/ComputerUseSettings.test.tsx
Tested: bun run check:server; bun run check:desktop; bun run check:docs; bun run check:policy; bun run check:native; git diff --check
Tested: SKIP_INSTALL=1 ./desktop/scripts/build-macos-arm64.sh; codesign verify; hdiutil verify; built CLI Computer Use E2E exposure and disable checks
Not-tested: Full screenshot/control action after granting macOS Screen Recording permission on this machine
Contributors need a visible path for local verification instead of relying on design notes or maintainer memory. Add a GitHub-facing CONTRIBUTING entrypoint, bilingual docs pages, README links, and VitePress sidebar navigation for PR gates, live model baselines, provider selection, reports, and release checks.
Constraint: Live baseline providers are local machine state; docs must tell contributors how to list and choose their own providers without maintainer UUIDs.
Rejected: Keep the instructions only in the quality-gate design doc | too hidden for clone-and-contribute workflows
Confidence: high
Scope-risk: narrow
Directive: Keep quality-gate commands documented wherever contributor onboarding links are exposed.
Tested: bun run check:docs
Tested: bun run quality:pr
Not-tested: live baseline after the docs-only wording change
Desktop sessions can switch provider and model while a CLI subprocess is already alive, so the server now serializes runtime restarts and marks provider-managed launches to prevent stale settings env from overriding the selected provider. Provider settings also write API key env consistently and clear stale managed keys before syncing.
This includes the related desktop/docs brand asset refresh and keeps the desktop locale default in Chinese, with tests updated to match the current provider semantics.
Constraint: Session-scoped model selection must win over cc-haha/settings.json and inherited ANTHROPIC_* values.
Rejected: Store the selected model as a global provider activeModel | chat runtime selection is per session.
Confidence: high
Scope-risk: moderate
Directive: Do not remove CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST without validating Desktop provider switching against stale settings env.
Tested: bun test src/server/__tests__/conversation-service.test.ts src/server/__tests__/conversations.test.ts src/server/__tests__/providers.test.ts src/server/__tests__/providers-real.test.ts
Tested: cd desktop && bun run test src/stores/settingsStore.test.ts
Tested: cd desktop && bun run lint
Tested: git diff --check
Not-tested: Full desktop production package/signing.
- Add MiniMax-M2.7 as the recommended default chat model
- Add MiniMax-M2.7-highspeed as a faster alternative
- Update base URL to api.minimax.io (international) with note about api.minimaxi.com for China
- Add model comparison table in documentation (both CN and EN)
- Add API key registration link to MiniMax Open Platform
Add GitHub-compatible slugify (github-slugger algorithm) to VitePress
config so heading anchor IDs are consistent across both platforms.
Update all inline anchor links in 22 doc files accordingly.
- Add Computer Use architecture deep dive documentation (CN/EN)
with 4 generated diagrams (architecture, security gates,
Python Bridge, patch environment comparison)
- Move Computer Use section below Skills in sidebar and README
- Fix 153 broken anchor links across 18 documentation files:
- CN docs: add missing `、` after Chinese numeral prefixes
- EN docs: add missing `_` prefix for number-starting heading IDs
- Fix em-dash and special character encoding in anchor hrefs