Commit Graph

50 Commits

Author SHA1 Message Date
程序员阿江(Relakkes) 7c7ce7d6b9 feat(desktop): add independent session collaboration 2026-09-21 01:16:27 +08:00
程序员阿江(Relakkes) 27f660fe5c fix(perf): bound large session history and tracing resource usage
Page transcript and trace reads, bound UI caches and retained task records,
and replace full-file background polling with incremental projections.
Preserve recovery and ownership semantics across pages and cancel stale work.
2026-09-19 20:29:20 +08:00
程序员阿江(Relakkes) f7fcbc2d8a merge: integrate local main with skills and connectors 2026-09-14 10:14:37 +08:00
程序员阿江(Relakkes) c4a4178f18 feat(desktop): unify skills and connectors with managed installation and mentions 2026-09-14 10:14:15 +08:00
程序员阿江(Relakkes) 9fce822df3 feat: add secure ngrok remote access with device pairing 2026-09-13 23:28:30 +08:00
程序员阿江(Relakkes) 0bcbfe6921 fix(computer-use): enforce cursor regression and package gates
Route native-only changes through macOS checks and verify relocated cursor resources in final packages. Reject resources that escape the app and exercise visible click feedback against a disposable native receiver.

Connect the regressions to required checks and capture shell fixture output through temporary files.
2026-09-11 15:53:22 +08:00
程序员阿江(Relakkes) 74f87a27e4 revert(session): remove protocol-based session restrictions
Revert efe5cae19 so existing sessions remain usable and models can be
switched without protocol admission checks. Retain the additive index
schema for already-upgraded caches and rebuild their summaries.
2026-09-10 14:00:50 +08:00
程序员阿江(Relakkes) 56c6a9aa8e feat(computer-use): align native app automation with Codex
Add a persistent isolated JavaScript worker for native app actions and batch
known operations without a model round trip between each input. Preserve
per-cell context, native errors, screenshot coordinates, and image types.

Align macOS gesture, key, inventory, capture, scroll, and clipboard behavior;
include a signed native receiver fixture and compiled sidecar regression tests.
Keep the Windows pixel route and fix cancellation with session-owned mouse
cleanup, lock revalidation, and portable signing-fixture tests.
2026-09-10 03:34:39 +08:00
程序员阿江(Relakkes) efe5cae19e fix(session): lock model switching to the session API protocol 2026-09-09 23:05:33 +08:00
程序员阿江(Relakkes) 5f7c719c96 fix(ci): preserve coverage logs across reporter cleanup 2026-09-07 23:30:57 +08:00
程序员阿江(Relakkes) c6146501cc fix(ci): capture coverage reports through regular file descriptors 2026-09-07 23:20:34 +08:00
程序员阿江(Relakkes) 8a37865536 fix(computer-use): harden native macOS automation runtime 2026-09-01 20:06:25 +08:00
程序员阿江(Relakkes) dd4edc0efe merge: bring main into the computer-use worktree
main is 87 commits ahead and carries a large amount of fixed behaviour this
branch should not be re-deciding. The rule applied throughout: this worktree
owns Computer Use, main owns everything else.

Only 12 files were touched on both sides, and Git merged all of them without
reporting a conflict — but two of those silent merges were wrong, and neither
was visible until the checks ran.

`desktop/src/api/client.ts` ended up with two `apiGetBlob` implementations.
Both sides had independently hit the same problem (an `<img src>` pointed at an
API endpoint is a cross-origin subresource, so it carries no Authorization
header and the server's fetch-metadata policy refuses it) and both had written
the same fix. Git saw two additions in different places and kept both, which
does not even compile. main's version survives: it builds its headers through
the shared `buildHeaders()` rather than assembling them inline, so it inherits
whatever main adds there later.

`src/server/api/computer-use.ts` still imported `runtime/mac_helper.py` and
`runtime/requirements.txt` as compile-time text, both deleted on this branch.
Nothing at runtime referenced them, which is why the deletion looked clean; the
bundler resolves those imports when the server module is loaded, so the failure
surfaced only when the tests actually imported it. That path is now Windows-only
in the same sense the rest of the Python bridge is, and it also ships
`win_cursor_badge.py`, which the badge needs because it runs as its own process.

`computer-use-requirements.test.ts` drops its darwin half for the same reason —
the pins it guards still matter, but only one requirements file is left.

Verified: server 3869 tests / 331 files, desktop 4612 tests / 319 files
(lint + tsc + build), Swift 272 XCTest + 14 Swift Testing, Python 25.

Claude-Session: https://claude.ai/code/session_015j1yxxaoonyAS2iZ7qGnTS
2026-08-23 18:39:47 +08:00
程序员阿江(Relakkes) 27c31bdc15 fix(models): preserve GPT relay reasoning effort (#1238) 2026-08-23 17:21:53 +08:00
Relakkes Yang fef789f5a8 fix(quality-gate): stabilize Windows validation 2026-08-22 16:06:17 +08:00
程序员阿江(Relakkes) a0afcc7efa feat(desktop): add custom project display names 2026-08-11 02:38:55 +08:00
程序员阿江(Relakkes) 4314345b6c fix(quality-gate): drive the ProseMirror composer in the desktop smoke
The smoke waited on `textarea`, but the composer became a ProseMirror
contenteditable (MentionComposer) some time ago, so the selector could
never match. The lane hung for its full 30s timeout and failed on every
run — the screenshot it captured on the way out showed the app rendered
and idle, which is what makes this easy to misread as a UI break.

Target `[data-composer-editor]`, the attribute the editor puts on its
editable node and the same hook composerTestUtils already drives.
2026-08-09 02:47:12 +08:00
程序员阿江(Relakkes) b8a90626ce feat(computer-use): native macOS engine, on top of main and nothing else
Rebuilt against main so the branch carries the Computer Use work and no other
divergence. Three unrelated efforts had been sitting uncommitted in this
worktree and were swept into an earlier commit; they are preserved on
cu-worktree-full-backup and belong on their own branches — adapter control
credentials, Electron asar sealing, and the sidecar code-loading audit. Every
file outside Computer Use now matches main exactly.

The engine
  A Swift helper drives apps through the accessibility tree, with coordinate
  actuation for the Chromium and Electron apps whose tree is a bare window
  frame. Ten primitives matching the shape Codex uses, so an app's guidance and
  the model's habits transfer.

  Coordinate actions resolve their target window once and refuse when none can
  be named. The unbound event they used to fall back to is discarded by custom
  renderers, so a minimized target produced a whole session of "Action
  completed" with nothing behind it.

  Input acceptance is established for typing and key presses as well as clicks:
  each MCP call is seconds apart, so the keyboard cannot inherit the focus a
  click established. The synthetic focus notification is gated on the target
  not already being active — sent unconditionally it names window 0 at an app
  that already owns a key window, and nine window-bound clicks were discarded
  with the traffic lights fully lit.

State the model can trust
  An off-screen target says so, and says which tools still reach it: element
  actions need no on-screen geometry, so an app with a real tree can still be
  driven from the Dock. A fully covered window is recovered once, then left
  alone — burying it again is the user wanting their screen back. A repeated
  capture is reported with the cause that actually applies rather than both,
  because coverage is something we compute.

Signing
  The helper is signed under a stable identity before electron-builder sees it,
  and excluded from re-signing: macOS ties Accessibility and Screen Recording
  grants to the signing identity, so rotating it drops both on every update.

Discoverability
  The desktop slash menu falls back to a directory scan while a session's CLI
  has not started, which is when the menu is first opened. Built-ins and
  bundled skills live in the binary, so /computer-use was absent until after
  the first message.
2026-08-05 22:06:23 +08:00
程序员阿江(Relakkes) 0cea7b5a61 feat(quality-gate): drive the agent flow with a provider the user configured
`check:agent-flow` proves the protocol with the mock CLI, which is what makes it
CI-safe and lets any contributor run it with no credentials. It cannot prove the
thing this product actually is: a desktop agent talking to a real model. That can
only run where the credentials are, so this lane is local and manual by
construction — registered in no quality-gate mode, referenced by no workflow, and
live.test.ts fails if either changes.

Six scenarios, sharing the existing harness rather than a second copy of it:
first turn, permission allow, permission deny, interrupt, reconnect, and history
recovery. Prompts induce the behaviour instead of dictating it, and assertions
only look at protocol shape and side effects on disk — never at generated text —
so the lane passes on any provider, including a local one. The three flows left
out (api-error, tool-error, runtime-select) each carry a written reason, because
a silently missing flow reads as a covered one.

Spending someone's quota is the failure mode worth engineering against, so the
runner refuses to guess: no implicit fallback to the active provider, an ambiguous
selector is an error rather than a pick, and without --yes it prints the provider,
model and config path it would use and exits without sending anything. User state
is copied into a throwaway config dir and the real ~/.claude is fingerprinted
before and after — a run that writes to it fails loudly instead of being cleaned
up quietly.

Not yet run end to end: the local LM Studio endpoint answers 502 here, so the six
runners have only been verified for structure. Target resolution, the
confirmation gate, and lane placement are covered by 14 tests that need no
provider at all.
2026-08-04 17:15:29 +08:00
程序员阿江(Relakkes) 537b48ba7b test(desktop): fail on unreachable components instead of covering them
Coverage answers "is this tested", never "should this exist", and the difference
cost real work. BackgroundTasksBar, SessionTaskBar and TeamStatusBar lost their
last import in 56a4be3d1 when SessionActivityPanel replaced them. Nothing noticed:
598b968ee then wrote tests for two of them — "cover the components left at zero" —
purely to lift changed-lines coverage past the gate, and c712f5285 restyled all
three during the UI redesign. 577 component lines plus 348 test lines were kept
alive for code no user could reach, along with nine translation keys carried in
five locales.

Delete all of it, and add the check that would have caught it: every .tsx under
src must be reachable by static import from a script tag in index.html or
gallery.html. Reading the entries out of the HTML rather than hardcoding them
means a new entry brings its whole subtree with it. The allowlist is empty and
should stay that way.

Scoped to .tsx deliberately. The .ts side has entry points a static graph cannot
see — workspaceDiffHighlight.worker.ts is a `new Worker(new URL(...))` target and
src/preview-agent/** is built into its own bundle — so covering it needs an
allowlist, which is where this kind of check goes to die.

Two mutations: putting BackgroundTasksBar.tsx back names it exactly, and breaking
ENTRY_HTML trips the entry-point guard first so the reader is not sent hunting
through 160 falsely-unreachable components.

Also drops the now-dangling desktop/src/mocks/ coverage exclusion, deleted in
33df50b9c.
2026-08-04 13:59:05 +08:00
程序员阿江(Relakkes) 4f9fec8760 ci(quality-gate): route checks by import graph and add offline agent QA 2026-08-03 21:46:54 +08:00
程序员阿江(Relakkes) 8bad67e590 fix(release): integrate RPM artifacts into desktop release 2026-07-28 00:51:10 +08:00
程序员阿江(Relakkes) 598b968eec test(desktop): cover the components left at zero and exclude dev tooling
Independent QA on 3264db10 flagged changed-lines coverage at 86.17%
(5539/6428), under the 90% gate. Two causes, handled differently.

`desktop/src/dev/` joins `mocks/` and `types/` in the coverage exclusions.
It holds the component gallery — 260 of the 889 uncovered lines, and by
far the largest single contributor. Vite never bundles it (the build
input is `index.html` alone), and unit-testing a page whose whole job is
rendering every primitive would assert that the primitives render, which
their own tests already do. Excluding it is a scope correction, not a
threshold adjustment.

The rest are three components this branch touched that had no test file
at all. They now have one each, covering the behavior that changed:

- `BackgroundTasksBar` — drawer open/close including Escape, the running
  vs finished split, dismissed-key filtering, and that clearing reports
  every finished key while keeping the drawer open if work continues.
- `TeamStatusBar` — the progress bar's `aria-valuenow`, lead exclusion
  from both list and count, and that it greens on "nothing running"
  rather than on 100%: one completed plus one errored is done at 50%,
  which is why `tone="auto"` would have been wrong here.
- `MarketSkillDetail` — skeleton semantics, retry, install/uninstall by
  `installState`, and the disabled+spinner state mid-install.

Changed-lines coverage: 91.06% (5610/6161).

The QA report's second finding, `check:impact` blocking on a missing
`allow-cli-core-change` label, is an artifact of the branch trailing
main. `check:impact` diffs against `main`, so main's own newer commits —
9 files under `src/` — are counted as this branch's. Against the merge
base the same evaluator returns `areas: desktop, blocked: false`. No code
change here; the branch needs a rebase before it can pass that lane.
2026-07-26 18:31:34 +08:00
程序员阿江(Relakkes) c2fd674662 docs: rebuild documentation site with React 2026-07-23 20:46:33 +08:00
程序员阿江(Relakkes) 7e998e0365 fix(search): bundle cross-platform ripgrep #923 2026-07-14 22:11:22 +08:00
程序员阿江(Relakkes) ad2a504bba feat(permission): add Auto mode #978 2026-07-11 20:06:36 +08:00
程序员阿江(Relakkes) 08dd64596a fix(ci): root explicit Bun test filters
Avoid Bun filter-mode repository scans that exhaust macOS file descriptors and corrupt subprocess test evidence. Apply rooted filters across server, contract, coverage, persistence, policy, desktop native, and adapter test entrypoints.

Confidence: high

Scope-risk: narrow

Tested: bun run check:policy; bun run check:server; bun run check:chat-contract
2026-07-10 23:49:46 +08:00
程序员阿江(Relakkes) 173d6a84fc fix(ci): isolate deterministic PR test evidence
Run required server and contract suites in credential-free sandboxes, fail closed on incomplete coverage or test output, and preserve the desktop active-turn permission guard across stale tab interactions.\n\nTested: bun run check:policy (115 pass); bun run check:server (1605 pass before final runner evidence check); bun run check:desktop; bun run check:provider-contract; bun run check:chat-contract\nConfidence: high\nScope-risk: broad
2026-07-10 22:59:02 +08:00
程序员阿江(Relakkes) 8cada04d10 fix(ci): align local verify with PR policy
Fail PR and release quality runs when the impact policy blocks the diff, and accept root-runtime regression coverage across service and utility seams.

Tested: bun run check:policy
Confidence: high
Scope-risk: narrow
2026-07-10 20:37:43 +08:00
程序员阿江(Relakkes) 435e4ccc9f ci: harden deterministic PR quality gates
Route required checks by changed surface, add offline provider and chat contracts, and keep fork PRs independent of live credentials. Layer agent guidance by subtree and enforce a compact instruction budget.

Tested: bun run check:policy
Confidence: high
Scope-risk: broad
2026-07-10 20:29:01 +08:00
程序员阿江(Relakkes) 782c32c1fc fix(desktop): keep cleared session metadata in sync
Reset session messageCount when /clear is confirmed so active headers and sidebars do not keep stale message totals after the transcript is cleared.

Also wait for the restored project chip in the live desktop smoke lane before filling the prompt, preventing the test from racing against EmptySession and creating a default-home session.

Tested: cd desktop && bun run test -- --run src/stores/chatStore.test.ts src/stores/sessionStore.test.ts

Tested: bun test scripts/quality-gate/desktop-smoke/execute.test.ts

Tested: bun run check:desktop

Tested: SKIP_INSTALL=1 MAC_TARGETS=zip desktop/scripts/build-macos-arm64.sh

Tested: bun run quality:smoke --provider-model codingplan:main:codingplan-main

Not-tested: bun run verify

Confidence: high

Scope-risk: narrow
2026-07-02 23:06:26 +08:00
程序员阿江(Relakkes) 35f43e8289 fix: resolve post-release semantic conflicts
Fix cross-issue regressions found during post-0.4.4 merge review:\n\n- preserve permission mode across clear and empty-session replacement flows\n- keep provider effort passthrough and context-window estimates aligned with runtime metadata\n- invalidate recent project caches and trace message signatures when sessions change\n- recognize Windows ARM64 unpacked package-smoke output\n\nTested: bun test scripts/quality-gate/package-smoke/index.test.ts scripts/quality-gate/runner.test.ts\nTested: bun run check:desktop\nTested: bun run check:server\nConfidence: high\nScope-risk: moderate
2026-07-02 22:11:43 +08:00
程序员阿江(Relakkes) 96a5842a01 fix(native): support Windows arm64 desktop startup (#954)
Add Windows ARM64 desktop release packaging, verify architecture-specific sidecar/native files in package smoke, and give the Electron sidecar more startup time plus early diagnostics for slow Windows ARM launches.

Tested: bun test electron/services/sidecarManager.test.ts
Tested: bun test scripts/quality-gate/package-smoke/index.test.ts scripts/release-update-metadata.test.ts scripts/pr/release-workflow.test.ts
Tested: bun run check:native
Tested: bun run check:policy
Not-tested: full bun run verify / coverage; this was a local issue-fix handoff, not PR-ready validation.
Confidence: high
Scope-risk: moderate
2026-07-02 18:27:55 +08:00
程序员阿江(Relakkes) c14a4ffe67 chore: refresh quarantine review window
Constraint: providers-real remains a live MiniMax connectivity check, so it stays quarantined from non-live gates.
Tested: bun run check:quarantine
Tested: bun run check:policy
Confidence: high
Scope-risk: narrow
2026-07-01 20:43:07 +08:00
程序员阿江(Relakkes) d7eef0ea51 fix(release): align Linux package smoke artifacts
Match Electron Builder Linux output by accepting linux-*-unpacked directories,
treating AppImage blockmaps as optional, and validating release asset names
against the generated x86_64/amd64 and arm64 artifacts.

Tested: bun test scripts/quality-gate/package-smoke/index.test.ts
Tested: bun test scripts/pr/release-workflow.test.ts scripts/quality-gate/package-smoke/index.test.ts
Tested: bun run check:policy
Tested: bun run check:native
Tested: bun run verify
Confidence: high
Scope-risk: moderate
2026-06-03 20:19:07 +08:00
程序员阿江(Relakkes) aed3e7d038 fix(release): restore quarantine quality gates
Separate quarantine review enforcement from server and coverage file selection so expired review dates fail only the governance lane.

Refresh stale server quarantine suites and keep only the live provider test quarantined for non-live PR gates.

Tested: bun run check:policy
Tested: bun run check:server
Tested: bun run check:coverage
Tested: bun run verify
Tested: git diff --check
Confidence: high
Scope-risk: moderate
2026-06-03 15:11:47 +08:00
程序员阿江(Relakkes) 232855fe77 fix(desktop): serve packaged H5 shell from Electron builds
Electron's sidecar runs outside app.asar, so H5 static files must be available as normal unpacked files. Point the sidecar at the unpacked renderer dist and keep a server fallback for stale app.asar-style paths.

Constraint: Packaged Bun sidecars cannot read app.asar paths with ordinary fs stat calls.
Rejected: Serve H5 from app.asar directly | the external sidecar is not Electron and does not get asar filesystem support.
Confidence: high
Scope-risk: narrow
Directive: Keep package-smoke checking app.asar.unpacked/dist/index.html before changing asarUnpack or H5 dist paths.
Tested: bun test src/server/__tests__/h5-access-auth.test.ts src/server/__tests__/h5-access-policy.test.ts
Tested: bun test desktop/electron/services/sidecarManager.test.ts scripts/quality-gate/package-smoke/index.test.ts
Tested: bun run check:server
Tested: cd desktop && bun run check:electron
Tested: git diff --check
Tested: SKIP_INSTALL=1 SIGN_BUILD=0 MAC_TARGETS=dmg desktop/scripts/build-macos-arm64.sh
Tested: packaged sidecar curl /?serverUrl=...&h5Token=... returned HTTP 200
Not-tested: Gatekeeper notarization for the local ad-hoc DMG
2026-06-01 23:32:56 +08:00
程序员阿江(Relakkes) 386a41e606 feat(desktop): enable Electron migration path
Introduce the Electron desktop shell alongside the existing React renderer and local Bun server boundary. The migration keeps the DesktopHost contract explicit across Tauri, Electron, and browser runtimes while adding Electron main/preload services for dialogs, shell, notifications, updates, tray/window lifecycle, terminal, preview WebContentsView, app mode, and release/package validation.

The commit also carries the latest local main desktop command updates, including agent slash entries and hidden-by-default markdown thinking details, so the packaged Electron build matches the current main UX surface.

Constraint: React renderer, local Bun server, REST/WebSocket, and sidecar boundaries must remain reusable during the migration
Constraint: macOS dev packages are ad-hoc signed and cannot prove Developer ID notarization or Gatekeeper release launch
Rejected: Browser-only smoke validation | it cannot exercise native dialogs, keychain prompts, notification behavior, or packaged app startup
Confidence: medium
Scope-risk: broad
Directive: Do not remove Tauri host support until signed Electron release artifacts pass native OS smoke on macOS, Windows, and Linux
Tested: bun run check:desktop
Tested: cd desktop && bun run check:electron
Tested: CSC_IDENTITY_AUTO_DISCOVERY=false bun run electron:package:dir
Tested: bun run test:package-smoke --platform macos --package-kind dir --artifacts-dir desktop/build-artifacts/electron
Tested: Computer Use read packaged Electron app window at desktop/build-artifacts/electron/mac-arm64/Claude Code Haha.app
Not-tested: Developer ID signed/notarized Gatekeeper launch
Not-tested: Real OS notification click-to-session action
Not-tested: Windows and Linux packaged app smoke on real hosts
2026-06-01 22:43:16 +08:00
程序员阿江(Relakkes) 4d6b1855bf release: prepare v0.3.2
Prepare the desktop release metadata and concise release notes while keeping
tagging and release publishing for a later step. The staged local build-script
updates keep desktop commands on checked-in local toolchain paths and avoid
rewriting preview-agent output when the built content is unchanged.

The persistence-upgrade gate now runs the focused desktop Vitest migration
suite in non-watch mode, matching the broader desktop quality lane and avoiding
pre-push termination during release preparation.

Constraint: Release publishing is intentionally deferred per request
Constraint: Desktop release metadata must keep package, Tauri config, Cargo metadata, and Cargo.lock aligned
Confidence: medium
Scope-risk: moderate
Directive: Do not tag v0.3.2 until release dry-run and final release verification are rerun on the release candidate
Tested: bun run scripts/release.ts 0.3.2 --dry
Tested: cd desktop && bun run lint
Tested: bun run check:persistence-upgrade
Tested: git diff --check
Not-tested: Full bun run verify
2026-06-01 02:30:01 +08:00
程序员阿江(Relakkes) e070c4a0d0 Make upgrade failures diagnosable without deleting user state
Desktop startup can fail before React mounts on older WebViews, so an HTML-level watchdog now renders startup diagnostics even when the module bundle never reaches the app code. Persistent provider migration now imports legacy root provider config into cc-haha-owned storage without deleting the old source file, and plugin marketplace cleanup refuses obvious corrupted cache roots or outside paths.

Constraint: User explicitly accepted the current reviewed state for landing despite remaining review concerns.

Constraint: Global ~/.claude state is user-owned and protected; automatic repair must avoid deleting shared config, transcripts, skills, MCP, plugins, OAuth, adapters, and teams.

Rejected: Tell users to delete ~/.claude or ~/.claude/cc-haha | unsafe because it can destroy user-owned Claude state and still may not fix WebView compatibility failures.

Confidence: medium

Scope-risk: moderate

Directive: Do not weaken protected-path checks; future deletion paths should validate real paths and symlink behavior before recursive rm.

Tested: bun test src/utils/plugins/installedPluginsManager.test.ts src/utils/plugins/marketplaceManager.test.ts

Tested: bun run check:server

Tested: cd desktop && bun run test -- --run src/main.test.tsx index-html.test.ts vite-config.test.ts src/theme/globals.test.ts

Tested: cd desktop && bun run build

Tested: bun run check:coverage

Not-tested: live macOS 12/Safari 15 WKWebView startup on an affected machine.

Not-tested: H5 diagnostic URL redaction and symlink-escape hardening are known follow-up risks from review.
2026-05-14 18:04:47 +08:00
程序员阿江(Relakkes) 90c9494c37 fix: remove pr-checks assertions and ratchet agent-utils coverage baseline
- Remove pr-checks lane assertions from runner.test.ts (lane was removed)
- Update agent-utils functions coverage baseline from 12.64% to 12.08% to match current measurement (total expanded from 4004 to 3989 functions)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 21:07:58 +08:00
程序员阿江(Relakkes) f379a9928a fix: remove pr-checks lane and fix test env isolation
- Remove pr-checks lane from quality gate (not needed locally)
- Fix launcherRouting tests to pass explicit null envAppRoot to avoid CLAUDE_APP_ROOT pollution
- Fix cron-scheduler-launcher test: CLAUDE_CODE_ENTRYPOINT is correctly set to sdk-cli for scheduled tasks, update assertion and restore env var in cleanup

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 21:00:06 +08:00
程序员阿江(Relakkes) fa47adc33a fix: prevent upgrade crashes from stale persistence
Desktop users can carry provider indexes, managed settings, localStorage state, and native update state from builds that no longer match current readers. This adds startup migrations and recovery paths before server and React state are consumed, plus a persistence upgrade gate so future storage protocol changes ship with old-format fixtures.

Constraint: Existing installs may contain malformed or legacy JSON/localStorage that must not block startup.
Constraint: Local verify should evaluate the current worktree diff rather than unrelated detached-worktree history.
Rejected: Treat invalid persisted state as fatal | reproduces white-screen and startup failure behavior for existing users.
Rejected: Bypass PR policy locally | hides real gate behavior and does not fix detached-worktree false positives.
Confidence: high
Scope-risk: moderate
Directive: Any local JSON, localStorage, or app config shape change must add a migration fixture and keep `bun run check:persistence-upgrade` green.
Tested: bun run check:persistence-upgrade; bun run check:policy; bun run check:desktop; bun run check:server; bun run check:native; bun run verify (9 passed, 1 coverage baseline failure)
Not-tested: Live provider baseline; existing user configs beyond covered fixtures
2026-05-06 23:20:21 +08:00
程序员阿江(Relakkes) b156be8d8d feat: make PR quality verification self-enforcing
Contributors and coding agents need one local command that both reports and enforces the quality contract. This change turns the PR gate into the shared verification entrypoint, adds path-selected local lanes, tightens coverage accounting around changed lines, and documents the repair loop in contributor and agent-facing guidance.

Constraint: Ordinary PR verification must stay non-live and runnable without provider credentials
Constraint: Coverage policy updates in this commit require maintainer approval before push/merge
Rejected: Keep quality guidance only in docs | agents need executable scripts and AGENTS.md instructions to follow the loop consistently
Confidence: high
Scope-risk: broad
Directive: Do not bypass `bun run verify` for production changes; fix failed lanes and coverage reports instead of lowering thresholds
Tested: bun run check:policy
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run verify
Not-tested: live provider baseline; no provider credentials were required for this non-live PR gate
2026-05-06 22:33:43 +08:00
程序员阿江(Relakkes) 9719726cd2 feat: make quality gates observable and enforceable
The repository now has a measurable PR quality path instead of a loose set of
manual checks. Coverage, quarantine governance, provider smoke, desktop smoke,
and workflow wiring all produce durable reports that contributors and maintainers
can inspect without reconstructing terminal output.

This also fixes the desktop smoke current-runtime path so browser-driven smoke
runs use the desktop default active provider instead of forcing the official
current model, and records that runtime decision as an artifact.

Constraint: Default PR gates must remain non-live and contributor-safe while live model checks stay explicit.
Constraint: Release packaging is still GitHub Actions based, so release preflight must run before the build matrix.
Rejected: Make live provider or desktop smoke mandatory on every PR | secrets, quotas, and model availability are maintainer-controlled.
Rejected: Let PRs lower coverage baselines in the same change | base-branch ratchet comparison must remain authoritative.
Confidence: high
Scope-risk: moderate
Directive: Do not relax coverage or quarantine policy without a maintainer approval label and a fresh quality report.
Tested: ALLOW_CLI_CORE_CHANGE=1 ALLOW_COVERAGE_BASELINE_CHANGE=1 bun run quality:gate --mode pr
Tested: bun run quality:gate --mode baseline --allow-live --only provider-smoke:* --provider-model nvidia-custom:main:nvidia-custom-main --artifacts-dir /tmp/quality-gate-live-smoke
Tested: bun run quality:gate --mode baseline --allow-live --only desktop-smoke:* --provider-model current:current:current-runtime --artifacts-dir /tmp/quality-gate-desktop-smoke-fixed
Tested: git diff --check
Not-tested: Full live release mode with multiple providers in hosted CI; provider credentials and quota remain maintainer-controlled.
2026-05-06 16:25:10 +08:00
程序员阿江(Relakkes) 6b83ffa4f5 Report the full quality gate impact
Release quality gates are used as maintainer evidence, so a live baseline failure must not hide later provider or desktop-smoke lanes. Run every lane and summarize all pass/fail statuses so reviewers can tell whether the whole matrix was exercised.

Constraint: Live model lanes can fail independently and may be slow, but release evidence needs complete coverage.

Rejected: Keep fail-fast behavior | it makes reports ambiguous after the first live failure.

Confidence: high

Scope-risk: narrow

Directive: Do not reintroduce fail-fast behavior for release gates without adding explicit unexecuted-lane reporting.

Tested: bun test scripts/quality-gate/runner.test.ts scripts/quality-gate/baseline/cases.test.ts scripts/quality-gate/providerTargets.test.ts

Tested: bun run quality:gate --mode release --allow-live --provider-model codingplan:main:codingplan-main --provider-model minimax:main:minimax-main

Not-tested: clean-worktree release rerun before this commit; run immediately after committing.
2026-05-02 16:08:28 +08:00
程序员阿江(Relakkes) 5f59c693c4 Let contributors choose quality-gate providers
The live baseline previously accepted provider UUIDs, which made the gate hard to run on another contributor's machine. Add a local provider listing command and resolve quality-gate targets from stable provider-name selectors while keeping UUIDs and current runtime support.

Constraint: Provider configuration is local machine state under CLAUDE_CONFIG_DIR and must not expose API keys.

Rejected: Require contributors to inspect providers.json manually | too error-prone and leaks implementation detail into the workflow

Confidence: high

Scope-risk: narrow

Directive: Keep live-provider baseline selection copyable from quality:providers before adding more live test lanes.

Tested: bun test scripts/quality-gate/providerTargets.test.ts

Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts

Tested: bun run quality:providers

Tested: bun run quality:gate --mode baseline --dry-run --provider-model codingplan:main --provider-model minimax:main

Tested: bun run quality:gate --mode baseline --dry-run --provider-model custom:haiku

Tested: bun run quality:gate --mode pr --dry-run

Tested: bun run check:server
2026-05-02 15:35:59 +08:00
程序员阿江(Relakkes) 5ed9903d2e Exercise desktop chat through a real browser baseline
The live baseline covered the server WebSocket path, but it still did not prove the desktop app can open a session, apply a selected provider/model, send a chat turn, and surface model/tool progress through the UI. This adds an agent-browser driven smoke lane that starts the local server and Vite desktop app, restores an isolated session tab with the requested runtime selection, submits a small coding task through the composer, and accepts the run only when the fixture diff and tests pass.

Constraint: Desktop smoke must stay behind --allow-live because it launches browsers and real models.

Constraint: The smoke temporarily enables bypassPermissions for the isolated run and restores the previous mode afterward.

Rejected: Wait for a final marker phrase | model reasoning and echoed prompts can contain the same marker before the work is actually done.

Rejected: Use only DOM text as success proof | the browser can show progress while the project files are still unchanged.

Confidence: high

Scope-risk: moderate

Directive: Future desktop smoke cases should verify project state and artifacts, not only UI copy.

Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts

Tested: bun run quality:gate --mode baseline --dry-run --provider-model minimax-m2.7

Tested: bun run quality:gate --mode baseline --allow-live --provider-model minimax-m2.7 (9 passed, 0 failed)

Tested: bun run check:server

Not-tested: Kimi desktop smoke completion because the provider returned AccountQuotaExceeded / API Error 429 until its reset window.
2026-05-02 14:07:06 +08:00
程序员阿江(Relakkes) d1219b9682 Broaden live agent baseline beyond toy edits
The first quality gate proved the execution path, but two tiny fixtures were not enough to reveal regressions in real agent behavior. This expands the maintained baseline with deeper failure recovery, workspace search, artifact creation, and cross-module refactor cases, and records per-case diffs so reviewers can see exactly what each model changed.

Constraint: Baseline runs must remain explicit maintainer-controlled live checks, not default PR checks.

Rejected: Require models to edit tests in every contract-change case | correct fixes can satisfy prewritten acceptance tests without modifying tests.

Rejected: Trust only command exit codes | passing tests alone does not show whether the agent edited forbidden or unrelated files.

Confidence: high

Scope-risk: moderate

Directive: Keep live baseline cases product-shaped and inspect both verification output and diff.patch before accepting future fixture changes.

Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts

Tested: bun run quality:gate --mode baseline --dry-run --provider-model volc-kimi-k2.6 --provider-model minimax-m2.7

Tested: bun run quality:gate --mode baseline --allow-live --provider-model volc-kimi-k2.6 --provider-model minimax-m2.7 (14 passed, 0 failed)

Not-tested: Desktop UI browser smoke and native Tauri release packaging remain outside this baseline expansion.
2026-05-02 13:22:55 +08:00
程序员阿江(Relakkes) f6511ab278 Protect release confidence with live agent baselines
The desktop product needs a repeatable local gate that can prove the core Coding Agent loop still works after changes, not only that unit tests pass. This adds a quality-gate runner with PR, baseline, and release modes, structured reports, explicit quarantine metadata, and fixture-based live baseline cases that can run across provider/model targets.

Constraint: Existing check:pr and CI policy behavior must remain usable while the stronger baseline grows around it
Constraint: Default PR gates must not require real model credentials or provider quota
Rejected: Build a standalone QA platform first | too heavy before the baseline task shape is proven
Rejected: Keep unstable server exclusions hardcoded in run-server-tests | hides quarantine policy from reports and future review
Confidence: medium
Scope-risk: moderate
Directive: Expand baseline cases by adding focused fixtures and verifiers; do not make normal PR checks depend on live providers
Tested: bun test scripts/quality-gate/*.test.ts scripts/quality-gate/baseline/*.test.ts
Tested: bun run check:server
Tested: bun run quality:gate --mode baseline --allow-live --provider-model 2944f963-ce75-45b7-bac1-6e4f57df0970:kimi-k2.6:volc-kimi-k2.6 --provider-model 9c78d3df-7fb5-44c7-8436-3a41c3a59231:MiniMax-M2.7-highspeed:minimax-m2.7
Not-tested: desktop UI browser smoke and native release mode in this commit
2026-05-02 13:00:55 +08:00