The instruction files told every coding agent to reach for agent-browser
whenever a change needed browser-level evidence: copilot-instructions
listed "E2E or agent-browser smoke" as the remedy for cross-boundary
flows, and both contributing guides repeated it. That wording outlived
the tool. With the agent-browser skill uninstalled, agents still parsed
those lines as a recommendation and went looking for the binary instead
of using the browser skill that is actually installed.
Deleting the references would have made the docs wrong. agent-browser is
still a real dependency: check:desktop-ui-smoke spawns it on Linux CI,
and seven maintainer-run e2e scripts under desktop/scripts drive it
directly. It cannot be swapped for ego-browser either — ego lite is a
macOS-only GUI app with no headless mode and a one-time interactive
onboarding, so it cannot run on ubuntu-latest at all.
So the lanes keep the binary and the prose loses the recommendation.
agent-browser is now described as an implementation detail of those two
call sites, and ad-hoc browser work — manual verification, screenshots,
exploratory UI checks — is pointed at the ego-browser skill.
The quality contract asserted the old string, so it would have failed
closed on the reworded line. It now pins the replacement plus the new
routing rule; flipping either sentence turns the test red.
The smoke waited on `textarea`, but the composer became a ProseMirror
contenteditable (MentionComposer) some time ago, so the selector could
never match. The lane hung for its full 30s timeout and failed on every
run — the screenshot it captured on the way out showed the app rendered
and idle, which is what makes this easy to misread as a UI break.
Target `[data-composer-editor]`, the attribute the editor puts on its
editable node and the same hook composerTestUtils already drives.
The check landed scoped to src/server/ws because blanking desynced on 6 of
2149 files and a desync reports a live import as dead. All four causes are
fixed, so the scope is now every source root no compiler checks — src,
scripts and adapters — and 0 of 2541 files desync.
The bugs, each with a regression test that places the import's only
reference after the construct so a desync makes it go dead:
- The token before a slash was read back out of the raw source, so the last
word of a preceding comment decided whether `/` opened a regex. In
useIssueFlagBanner.ts `// …correction tone` made the next line's regex lex
as a division, and its apostrophe opened a string that ate the line. A
comment is whitespace to the grammar; it now contributes nothing.
- That token accumulated across whitespace, so `return false` became one
token named `returnfalse` and the following `return /re/` no longer looked
like a keyword. This is what broke markdownImages.ts and dead-imports.ts
itself.
- `input! / 10` divides but `!/re/.test(x)` negates, and both put `!` before
the slash. What precedes the `!` settles it.
- Character classes and quotes inside regular expressions, fixed earlier.
desktop/ stays out of scope: its tsconfig already sets noUnusedLocals, and
scanning it anyway finds nothing — the cross-check that this agrees with a
real compiler. `blankingIsSound` still refuses to analyse a file whose
blanked form no longer parses, so a future desync is reported, not acted on.
Routing follows the scope: policyPrefixes now names adapters/, scripts/ and
src/ instead of the three scripts/ subdirectories and src/server/ws/. The
check reads these files rather than importing them, so the import graph
cannot select the lane on its own. `does not widen docs, policy, or coverage
lanes` split in two — its fixture selects the policy lane through its own
files now, so the dependent-must-not-widen invariant moved to a desktop
fixture that still shows it.
Verified by mutation, each reverted from an explicit backup: planting
`plantedProbeSymbol` into one file per root reported all four and failed
check:policy; reverting each of the three lexer fixes failed exactly its own
regression test; removing 'src/' from policyPrefixes failed the routing test
and made change-policy report policy=false for a src-only diff.
check:policy is 223 pass / 0 fail in 8.6s, up from 2s — the scan is 3.8s
over 2150 files and the planted-import test covers every one of them.
Nothing checked src/ for unreferenced imports. desktop/tsconfig.json sets
noUnusedLocals and eslint covers desktop/ only; the root tsconfig.json sets
no such option and nothing installs typescript or bun-types at the root, so
no tool reads it at all. Splitting handler.ts left nine imports whose
symbols had moved out, and a human found them by reading the diff.
Measured before choosing. Under a temporary tsconfig extending the root
one, tsc reports 3225 errors over src/ and scripts/ before noUnusedLocals
and 3871 after — 646 net, on a baseline that already fails. That option
would land disabled, so this adds a narrow check instead: imports only,
one directory.
The analysis is lexical like module-graph.ts, but it blanks comments and
literal text first, so a symbol kept alive only by a comment still reports
dead — the exact shape the handler.ts split left behind. Template
substitutions, `//` inside a URL string and quotes inside a regex literal
must survive that blanking or a live import reads as dead;
src/utils/terminalShellEnvironment.ts is the last one and mis-lexing it
blanked 130 lines. Where blanking desyncs anyway the result no longer
parses, so the file is reported as degraded rather than mis-analysed: 6 of
2149 files repo-wide, none under src/server/ws.
src/server/ws/ also joins policyPrefixes. The check reads its files rather
than importing them, so the import graph cannot route a ws-only diff to
this lane, and without the prefix the check would never run on the diffs it
was written for.
Verified by mutation, each reverted from an explicit backup:
- planted `import { randomUUID }` into src/server/ws/events.ts →
check:policy 216 pass / 1 fail, reporting
"src/server/ws/events.ts:1 randomUUID"
- disabled line-comment blanking → 1 fail; block-comment blanking → 2 fail;
regex-literal detection → 2 fail (the extra one is the desync guard)
- pointed DEAD_IMPORT_ROOTS at a missing directory → 1 fail, so the check
cannot silently scan nothing
- dropped src/server/ws/ from policyPrefixes → 1 fail, and change-policy
reports policy=false for a ws-only diff
check:policy is 219 pass / 0 fail; src/server/__tests__/websocket-handler
.test.ts is 87 pass / 0 fail.
`check:agent-flow` proves the protocol with the mock CLI, which is what makes it
CI-safe and lets any contributor run it with no credentials. It cannot prove the
thing this product actually is: a desktop agent talking to a real model. That can
only run where the credentials are, so this lane is local and manual by
construction — registered in no quality-gate mode, referenced by no workflow, and
live.test.ts fails if either changes.
Six scenarios, sharing the existing harness rather than a second copy of it:
first turn, permission allow, permission deny, interrupt, reconnect, and history
recovery. Prompts induce the behaviour instead of dictating it, and assertions
only look at protocol shape and side effects on disk — never at generated text —
so the lane passes on any provider, including a local one. The three flows left
out (api-error, tool-error, runtime-select) each carry a written reason, because
a silently missing flow reads as a covered one.
Spending someone's quota is the failure mode worth engineering against, so the
runner refuses to guess: no implicit fallback to the active provider, an ambiguous
selector is an error rather than a pick, and without --yes it prints the provider,
model and config path it would use and exits without sending anything. User state
is copied into a throwaway config dir and the real ~/.claude is fingerprinted
before and after — a run that writes to it fails loudly instead of being cleaned
up quietly.
Not yet run end to end: the local LM Studio endpoint answers 502 here, so the six
runners have only been verified for structure. Target resolution, the
confirmation gate, and lane placement are covered by 14 tests that need no
provider at all.
4f9fec876 added this workflow with `cron: '0 18 * * *'`. That was the wrong call
to make unilaterally: the repository had no scheduled workflow at all before it,
so this was not one more cron among several but the introduction of recurring CI
spend — about ninety minutes per run — on a schedule nobody asked for.
The reasoning for the sweep still holds: a per-PR gate only covers what the diff
reaches, so it is blind to checks no recent PR selected and to failures that only
appear when the whole suite runs together. Keeping the workflow on
`workflow_dispatch` keeps that one click away without deciding for the maintainer
when to spend the time.
pr-quality-workflow.test.ts now asserts the absence of `schedule:` and `cron:`
rather than their presence, so a schedule cannot drift back in unnoticed —
verified by adding the cron back and watching the test go red. The docs' four-tier
table renames the tier accordingly; calling it "Nightly" when nothing runs nightly
is exactly the kind of comment that outlives its code.
Coverage answers "is this tested", never "should this exist", and the difference
cost real work. BackgroundTasksBar, SessionTaskBar and TeamStatusBar lost their
last import in 56a4be3d1 when SessionActivityPanel replaced them. Nothing noticed:
598b968ee then wrote tests for two of them — "cover the components left at zero" —
purely to lift changed-lines coverage past the gate, and c712f5285 restyled all
three during the UI redesign. 577 component lines plus 348 test lines were kept
alive for code no user could reach, along with nine translation keys carried in
five locales.
Delete all of it, and add the check that would have caught it: every .tsx under
src must be reachable by static import from a script tag in index.html or
gallery.html. Reading the entries out of the HTML rather than hardcoding them
means a new entry brings its whole subtree with it. The allowlist is empty and
should stay that way.
Scoped to .tsx deliberately. The .ts side has entry points a static graph cannot
see — workspaceDiffHighlight.worker.ts is a `new Worker(new URL(...))` target and
src/preview-agent/** is built into its own bundle — so covering it needs an
allowlist, which is where this kind of check goes to die.
Two mutations: putting BackgroundTasksBar.tsx back names it exactly, and breaking
ENTRY_HTML trips the entry-point guard first so the reader is not sent hunting
through 160 falsely-unreachable components.
Also drops the now-dangling desktop/src/mocks/ coverage exclusion, deleted in
33df50b9c.
Two additions to the provider settings page, both modelled on cc-switch.
One-click import from cc-switch:
- Reads the local cc-switch installation and offers its Claude Code
providers for bulk import. SQLite (cc-switch v3.8.0+) is the primary
store, with the legacy v2 config.json as a fallback used only when no
database exists — once cc-switch migrates it archives that file, so
reading it alongside a database would surface pre-migration data.
- Supports cc-switch v3.1.0 and newer. Older installs wrote a config
format cc-switch itself dropped in v3.6.0; those are refused explicitly
rather than reported as an empty scan.
- Degrades honestly when cc-switch's storage moves: a structure we cannot
read reports why, distinguishing "cc-switch too old" from "cc-haha does
not recognise this layout" from "the file could not be read at all".
Only id/app_type/settings_config are required; other columns are
optional and unknown ones are ignored.
- Full credentials are resolved server-side during import and never
appear in the scan payload.
Fetch model lists:
- Probes the provider's Base URL for an OpenAI-compatible /models
endpoint, walking cc-switch's candidate ladder (version segments and
nine vendor compat suffixes) and falling through on 404/405.
- Only http(s) endpoints are fetched; a 2xx that carries no model list is
reported as a failure with the upstream's own message rather than as an
empty catalog, so a key rejected behind a 200 is not read as "this
provider has no models".
The English README becomes README.md so the GitHub landing page reads in
English; the Chinese one moves to README.zh-CN.md. Both cross-link at the
top. Drops the source-origin framing from the intro and removes the
Disclaimer section, and points the license badge at MIT.
change-policy and pr-triage hardcoded README.en.md, which would have
dropped README.zh-CN.md out of the docs lane.
The site had drifted from the product. Every screenshot predated the
v0.5.0 UI redesign, the reading experience shipped no search and no
syntax highlighting, and a third of the pages were internal process
artefacts — migration task lists addressed to agentic workers, a
release runbook, a proposal marked "historical".
Reorganise around the only two people who read this: someone getting
the desktop app running for the first time, and someone reading the
source. Five sections replace nine — start / desktop / im / cli /
internals — and the pages that served neither reader are gone.
Site rewrite:
- Palette lifted from the desktop app's 「纸·墨·印」 themes, so the
site and the product read as one thing. Light mirrors 纯白, dark
mirrors 墨夜, and dark mode exists at all now.
- Fonts are self-hosted. The old @import from Google Fonts is
unreachable from mainland China, which left every heading in a
fallback serif; it also only requested weight 600 while the CSS
asked for 900, so Latin and CJK in the same heading disagreed.
- Docs were shipped as one 968KB manifest downloaded on every page
view. Split into a 32KB index plus one lazily imported chunk per
page; the entry bundle is now 101KB gzipped.
- Add search, syntax highlighting, per-route meta with canonical and
hreflang, a sitemap, and an error boundary. Replace the 44vh
mobile sidebar with a drawer.
- Image dimensions are read at build time and written into the tag,
so lazy images reserve their space instead of collapsing.
Screenshots are recaptured from a real v0.5.0 build against a clean
demo project, with tokens, QR codes and paired accounts redacted.
The previous set is deleted rather than kept alongside.
Routes follow file paths, so the restructure would have broken every
inbound link; 37 old paths redirect, in both languages. The PR policy
gate and CODEOWNERS also hardcoded docs/guide/contributing.md.
Verified: check:docs 78 pages / 323 links / 0 problems, check:policy
127 pass. Walked every route at 1440 and 390 in both themes for
overflow, contrast, keyboard reachability and focus management.
The Windows x64 build job went red on `Verify compiled Windows sidecar
startup`, asserting that a loopback request without
`CC_HAHA_LOCAL_ACCESS_TOKEN` returns 403 while it returned 200. Nothing
about x64 is involved — that step carries
`if: matrix.smoke_platform == 'windows' && matrix.arch == 'x64'`, so it is
the only job in the whole matrix that runs the smoke at all. Any regression
in this area can surface nowhere else.
The 200 is correct. `7d2a8a3cd keep loopback trusted without the desktop
process token` deliberately made the token additive again: gating every
local request behind it turned the Grok OAuth success page, `/preview-fs`
links and plain `curl` into 401s, because none of that traffic can ever
carry the token. Loopback is trusted on its own; the token is demanded only
on the `/api/h5-access` control plane, where another program on the same
box must not be able to publish the user's sessions to the network. The
assertion, written before that change, was still guarding the path that had
been intentionally opened.
So the probe moves to the boundary that is actually enforced, and gains a
positive assertion — loopback without a token must be 200 — so the additive
model is pinned down rather than merely no longer contradicted. Reverting to
the pre-`7d2a8a3cd` behaviour now fails the smoke instead of passing it.
Three copies of the stale assertion existed; all three are updated. Only the
compiled-sidecar smoke runs in CI, but `local-index-benchmark.ts` and its
corpus test were already failing the same way for anyone running them
locally. The benchmark also cancels the probe response bodies now: an unread
body holds its connection open, and that would land in the event-loop delay
and RSS samples taken immediately after.
Verified with the CI parameters — `bun run build:sidecars` then
`CC_HAHA_COMPILED_SIDECAR_SMOKE_STARTS=20 bun run test:compiled-sidecar-smoke`,
8/8 — and by running the benchmark directly, which now reports
`loopbackAuth` as 200/403/403/200 with validation intact.
Independent QA on 3264db10 flagged changed-lines coverage at 86.17%
(5539/6428), under the 90% gate. Two causes, handled differently.
`desktop/src/dev/` joins `mocks/` and `types/` in the coverage exclusions.
It holds the component gallery — 260 of the 889 uncovered lines, and by
far the largest single contributor. Vite never bundles it (the build
input is `index.html` alone), and unit-testing a page whose whole job is
rendering every primitive would assert that the primitives render, which
their own tests already do. Excluding it is a scope correction, not a
threshold adjustment.
The rest are three components this branch touched that had no test file
at all. They now have one each, covering the behavior that changed:
- `BackgroundTasksBar` — drawer open/close including Escape, the running
vs finished split, dismissed-key filtering, and that clearing reports
every finished key while keeping the drawer open if work continues.
- `TeamStatusBar` — the progress bar's `aria-valuenow`, lead exclusion
from both list and count, and that it greens on "nothing running"
rather than on 100%: one completed plus one errored is done at 50%,
which is why `tone="auto"` would have been wrong here.
- `MarketSkillDetail` — skeleton semantics, retry, install/uninstall by
`installState`, and the disabled+spinner state mid-install.
Changed-lines coverage: 91.06% (5610/6161).
The QA report's second finding, `check:impact` blocking on a missing
`allow-cli-core-change` label, is an artifact of the branch trailing
main. `check:impact` diffs against `main`, so main's own newer commits —
9 files under `src/` — are counted as this branch's. Against the merge
base the same evaluator returns `areas: desktop, blocked: false`. No code
change here; the branch needs a rebase before it can pass that lane.
Avoid Bun filter-mode repository scans that exhaust macOS file descriptors and corrupt subprocess test evidence. Apply rooted filters across server, contract, coverage, persistence, policy, desktop native, and adapter test entrypoints.
Confidence: high
Scope-risk: narrow
Tested: bun run check:policy; bun run check:server; bun run check:chat-contract
Run required server and contract suites in credential-free sandboxes, fail closed on incomplete coverage or test output, and preserve the desktop active-turn permission guard across stale tab interactions.\n\nTested: bun run check:policy (115 pass); bun run check:server (1605 pass before final runner evidence check); bun run check:desktop; bun run check:provider-contract; bun run check:chat-contract\nConfidence: high\nScope-risk: broad
Fail PR and release quality runs when the impact policy blocks the diff, and accept root-runtime regression coverage across service and utility seams.
Tested: bun run check:policy
Confidence: high
Scope-risk: narrow
Route required checks by changed surface, add offline provider and chat contracts, and keep fork PRs independent of live credentials. Layer agent guidance by subtree and enforce a compact instruction budget.
Tested: bun run check:policy
Confidence: high
Scope-risk: broad
Prevent repeated update checks from clearing the pending Electron update, recover the UI when relaunch does not exit, and keep public releases draft until all updater assets are uploaded.
Tested: bun test scripts/release-update-metadata.test.ts scripts/pr/release-workflow.test.ts
Tested: cd desktop && bun run test -- src/stores/updateStore.test.ts --run
Tested: cd desktop && bun run test -- electron/services/updater.test.ts --run
Tested: git diff --check
Tested: bun run check:policy
Tested: bun run check:desktop
Not-tested: bun run check:native
Not-tested: bun run check:coverage
Not-tested: bun run verify
Confidence: medium
Scope-risk: moderate
Reset session messageCount when /clear is confirmed so active headers and sidebars do not keep stale message totals after the transcript is cleared.
Also wait for the restored project chip in the live desktop smoke lane before filling the prompt, preventing the test from racing against EmptySession and creating a default-home session.
Tested: cd desktop && bun run test -- --run src/stores/chatStore.test.ts src/stores/sessionStore.test.ts
Tested: bun test scripts/quality-gate/desktop-smoke/execute.test.ts
Tested: bun run check:desktop
Tested: SKIP_INSTALL=1 MAC_TARGETS=zip desktop/scripts/build-macos-arm64.sh
Tested: bun run quality:smoke --provider-model codingplan:main:codingplan-main
Not-tested: bun run verify
Confidence: high
Scope-risk: narrow
Fix cross-issue regressions found during post-0.4.4 merge review:\n\n- preserve permission mode across clear and empty-session replacement flows\n- keep provider effort passthrough and context-window estimates aligned with runtime metadata\n- invalidate recent project caches and trace message signatures when sessions change\n- recognize Windows ARM64 unpacked package-smoke output\n\nTested: bun test scripts/quality-gate/package-smoke/index.test.ts scripts/quality-gate/runner.test.ts\nTested: bun run check:desktop\nTested: bun run check:server\nConfidence: high\nScope-risk: moderate
Add Windows ARM64 desktop release packaging, verify architecture-specific sidecar/native files in package smoke, and give the Electron sidecar more startup time plus early diagnostics for slow Windows ARM launches.
Tested: bun test electron/services/sidecarManager.test.ts
Tested: bun test scripts/quality-gate/package-smoke/index.test.ts scripts/release-update-metadata.test.ts scripts/pr/release-workflow.test.ts
Tested: bun run check:native
Tested: bun run check:policy
Not-tested: full bun run verify / coverage; this was a local issue-fix handoff, not PR-ready validation.
Confidence: high
Scope-risk: moderate
Constraint: providers-real remains a live MiniMax connectivity check, so it stays quarantined from non-live gates.
Tested: bun run check:quarantine
Tested: bun run check:policy
Confidence: high
Scope-risk: narrow
Make workflow_dispatch draft handling explicit and add a post-publish guard that edits the release back to draft when manual draft runs upload assets to an existing release.\n\nTested: bun test scripts/pr/release-workflow.test.ts\nTested: git diff --check\nScope-risk: narrow\nConfidence: high
Replace electron-builder's internal macOS notarization wait with an explicit notarytool flow: build a signed app with update config, submit it with a bounded notarytool timeout, staple and validate it, then rebuild dmg/zip from the notarized app via --prepackaged. Keep draft-only signed/no-notary builds available for fast artifact checks.\n\nTested: bun test scripts/pr/release-workflow.test.ts\nTested: git diff --check\nTested: bun run scripts/release.ts 0.4.3 --dry\nTested: local arm64 signed/no-notary build followed by --prepackaged dmg/zip rebuild\nTested: bun run test:package-smoke --platform macos --package-kind release --artifacts-dir desktop/build-artifacts/electron\nTested: codesign --verify --deep --strict --verbose=2 desktop/build-artifacts/electron/mac-arm64/Claude\ Code\ Haha.app\nConfidence: medium\nScope-risk: moderate
Add a workflow_dispatch-only notarize_macos switch so draft release runs can build Developer ID signed macOS artifacts without waiting on Apple notarization when GitHub runner networking is failing. Tag push releases still default to notarization and Gatekeeper release checks.\n\nTested: bun test scripts/pr/release-workflow.test.ts\nTested: git diff --check\nTested: bun run scripts/release.ts 0.4.3 --dry\nConfidence: medium\nScope-risk: moderate
Capture the signed macOS electron-builder exit code before the retry branch so watchdog timeouts and notarization failures fail the signed step directly instead of falling through to package-smoke.\n\nTested: bun test scripts/pr/release-workflow.test.ts\nTested: git diff --check\nTested: bun run scripts/release.ts 0.4.3 --dry\nTested: local bash watchdog timeout returns status 124\nConfidence: high\nScope-risk: narrow
Add an outer watchdog around the signed macOS electron-builder step so hung notarytool waits return control to the retry loop instead of idling until the GitHub step timeout.\n\nTested: bun test scripts/pr/release-workflow.test.ts\nTested: git diff --check\nTested: bun run scripts/release.ts 0.4.3 --dry\nConfidence: medium\nScope-risk: narrow