docs: streamline agent guidance for GPT-6 Astra

This commit is contained in:
程序员阿江(Relakkes)
2026-09-06 14:37:38 +08:00
parent 633320e9f2
commit c91f844712
14 changed files with 194 additions and 223 deletions
+6 -14
View File
@@ -2,18 +2,10 @@
Follow the root `AGENTS.md` and the nearest nested `AGENTS.md` for the files you edit.
For every feature, bugfix, refactor, or workflow change:
Use the root completion and authorization boundaries; consult `docs/internals/contributing.md` for task-specific testing and failure diagnosis. Continue scoped local implementation, deterministic checks, and fixes without asking at each step.
- Treat tool access as capability, not authorization. Do not commit, push, open/merge a PR, release, run live providers, or change repository settings unless explicitly requested.
- Inspect `git status --short`, identify the changed surface, define the intended behavior and failure signal, and inspect the nearest implementation and tests before editing.
- Identify the changed surface before coding: `desktop`, `server/runtime`, `adapter`, `native`, `docs`, `provider/runtime`, `agent-loop`, `persistence`, `policy/ci`, or `release`.
- Add same-area tests with the production change. Do not leave production behavior untested unless the PR explicitly carries the maintainer override `allow-missing-tests`.
- Preserve or improve the coverage ratchet. New or changed executable production lines must pass the changed-line coverage threshold in `scripts/quality-gate/coverage-thresholds.json`; do not edit coverage baselines or thresholds without maintainer approval via `allow-coverage-baseline-change`.
- Use unit tests for pure logic, API/request-shape tests for server/provider/runtime behavior, Testing Library/Vitest for desktop UI and stores, and E2E or desktop UI smoke for user-visible cross-boundary flows.
- Ad-hoc browser automation (manual verification, screenshots, exploratory UI checks) goes through the `ego-browser` skill. The `agent-browser` binary is reserved for the committed `check:desktop-ui-smoke` lane and `desktop/scripts/e2e-*-agent-browser.sh`; do not reach for it as a general browser tool.
- Provider/auth/runtime-env/model-window/proxy changes require offline `bun run check:provider-contract`; desktop chat/WebSocket/session-runtime changes require `bun run check:chat-contract`.
- Required PR evidence must be deterministic: use fake credentials, temporary config/home paths, mocked or loopback transports, explicit cleanup, and restored environment state. Never call a real provider or use saved machine credentials in required tests.
- For agent loop, tool execution, provider routing, model selection, file editing, permissions, session resume, and desktop chat changes, include mock/fixture/contract tests. Live smoke is trusted-maintainer evidence only and requires explicit authorization; finding local credentials is not authorization.
- Run the focused regression first, then `bun run check:impact` and every selected surface/contract check. Run `bun run verify` only before claiming PR-ready/push-ready or when full validation was requested.
- Do not present skipped, blocked, not-run, mock, build-only, or stale evidence as passed live/runtime verification.
- In the final handoff or PR description, include changed files, tests added, commands actually run with pass/fail counts, checks not run, coverage report path when generated, deterministic E2E evidence, live report path or explicit maintainer-only deferral, and known residual risk.
- Add same-area tests with the production change as defined by `scripts/pr/change-policy.ts`. Preserve or improve the coverage ratchet and meet the changed-line coverage threshold; maintainer overrides remain explicit decisions.
- Use E2E or desktop UI smoke when unit tests cannot prove a user-visible cross-boundary flow. Ad-hoc browser automation uses `ego-browser`; committed smoke scripts retain their own runner.
- Provider/auth/runtime-env/model-window/proxy changes require offline `bun run check:provider-contract` when selected by `check:impact`; desktop chat/WebSocket/session changes likewise use `check:chat-contract`.
- Live smoke is trusted-maintainer evidence only and requires explicit authorization. Required tests use isolated fixtures and no saved credentials.
- Follow the root verification policy once for the final diff. In the handoff, include changed files, tests added, commands actually run with pass/fail counts when available, checks not run, evidence paths when generated, and remaining risk.
+29 -86
View File
@@ -1,111 +1,54 @@
# Repository Instructions
This file is the entry point for coding agents. Keep it short: it should route an agent to the right code, tests, and deeper documentation rather than duplicate them.
This is a routing guide for coding agents. Keep shared instructions model-independent; put task-specific detail next to the code or in the linked guides.
Rules closer to the code take precedence. Before editing `.github/`, `src/`, `desktop/`, `adapters/`, or `docs/`, read the nested `AGENTS.md` in that directory.
Rules closer to the code take precedence. For the directory you are changing, read the nested `AGENTS.md` in that directory and any applicable ancestors. Load other documentation when the task needs it.
## Start Here
- Run `git status --short` before editing. Preserve all existing user changes and never revert, restage, reformat, or overwrite unrelated work.
- Identify the affected surface and inspect its production path, nearest tests, and existing implementation pattern before proposing a change. Check recent history when regression context matters.
- For bugs, reproduce the failure or add a regression test that fails for the intended reason. If reproduction is impossible, state the limitation instead of guessing.
- Define the smallest behavior change and the proof that will demonstrate it. Stop and re-scope if the diff crosses an unplanned surface, adds a dependency, or grows beyond the verified seam.
- For broad investigation, parallel read-only subagents are encouraged. Give editing agents non-overlapping file ownership; the primary agent owns integration and final verification.
- Tool access is capability, not authorization. Do not create/switch branches, commit, push, open or merge a PR, publish a release, change repository settings, or spend live-provider quota unless the user explicitly requests that operation.
- Run `git status --short` before editing and preserve existing user changes.
- Carry the requested outcome through implementation, relevant verification, and repair of failures caused by the change. Local edits and deterministic checks using disposable fixtures are authorized within that scope; do not stop for approval after the first implementation or each test run.
- Investigate the affected behavior and its callers. Follow a fix across boundaries when needed to complete the request; ask only when a material product decision or action outside the authorized scope is required.
- Use subagents for independent investigations or non-overlapping implementation work when useful. Give each a concrete question or file ownership and expected evidence; the primary agent owns integration and final verification.
- Tool access is capability, not authorization. Do not create/switch branches, commit, push, open or merge a PR, publish a release, change repository settings, or spend live-provider quota unless the user explicitly requests that operation. Authorization already given in the conversation remains valid.
## Repository Map
- `src/`: CLI, Ink UI, commands, services, tools, shared runtime utilities, and the local API/WebSocket server.
- `desktop/`: React desktop UI, Electron host, native/sidecar resources, and desktop build scripts.
- `adapters/`: Telegram, Feishu, WeChat, DingTalk, WhatsApp, WeCom, QQ, Slack, and shared IM adapter utilities.
- `site/`: React documentation site and build tooling. `docs/` and `docs/en/` are its Chinese and English Markdown content sources; keep counterparts aligned when both exist.
- `.github/workflows/`, `scripts/pr/`, and `scripts/quality-gate/`: CI routing and quality policy.
- `release-notes/`, `scripts/release.ts`, and `.github/workflows/release-desktop.yml`: desktop release automation.
| Surface | Entry point |
| --- | --- |
| CLI, tools, runtime, local API/WebSocket server | [src/AGENTS.md](src/AGENTS.md) |
| React desktop UI, Electron host, native/sidecars | [desktop/AGENTS.md](desktop/AGENTS.md) |
| IM platforms and shared chat runtime | [adapters/AGENTS.md](adapters/AGENTS.md) |
| Chinese and English product/source documentation | [docs/AGENTS.md](docs/AGENTS.md) |
| React documentation site and build tooling | [site/AGENTS.md](site/AGENTS.md) |
| CI and quality policy | [.github/AGENTS.md](.github/AGENTS.md), `scripts/pr/`, `scripts/quality-gate/` |
| Desktop releases and auto-update | `release-notes/`, `scripts/release.ts`, [release guide](docs/internals/contributing.md#发版与自动更新) |
## Implementation Rules
- Make narrow, owned diffs. Every changed line must trace to the request, a failing test, or a verified compatibility constraint.
- Prefer existing utilities, stores, services, and test harnesses. Do not add dependencies or speculative abstractions unless the task requires them.
- Production changes under `src/`, `desktop/src/`, or `adapters/` require a same-area regression test unless a maintainer explicitly approves an exception. A test that only covers the hop you just changed satisfies this rule and still lets the next change break — see "Writing a test that holds" below.
- Keep TypeScript ESM style: 2-space indentation, no semicolons, `PascalCase` components, and `camelCase` functions/hooks/stores.
- Use structured parsers and existing boundaries instead of ad hoc string manipulation. Add comments only for non-obvious control flow or external constraints.
- Keep changes tied to the requested behavior. Reuse existing utilities, stores, services, and test harnesses; add dependencies or abstractions only when the task needs them.
- Executable JS/TS production changes under `src/`, `desktop/src/`, or `adapters/` require a same-area regression test unless a maintainer explicitly approves an exception. For bugs, reproduce the failure or add a test that fails for the intended reason; report when reproduction is unavailable. Test the behavior and affected boundaries. See [test design](docs/internals/contributing.md#回归测试设计) for state transitions, replay, and coverage caveats.
- Keep TypeScript ESM style: 2-space indentation, no semicolons, `PascalCase` components, and `camelCase` functions/hooks/stores. Use structured parsers and existing boundaries for structured data.
- Do not commit generated output such as `artifacts/`, coverage reports, `node_modules/`, build directories, or Rust `target/` trees.
- When publishing is explicitly requested, use Conventional Commit subjects and normal product branch prefixes such as `fix/`, `feat/`, or `docs/`; do not create `codex/` branches in this repository.
## Writing a Test That Holds
Most regressions here are repairs of a recent repair: 21 of the last 70 `fix` commits
edit lines another `fix` wrote within 30 days. Coverage is not the missing signal —
`ContextUsageIndicator.tsx` sits at 87% branch coverage and was fixed three times in
ninety minutes. What those tests had in common is shape, so choose it deliberately.
- **Drive the transition; never hand-write the state it produces.** Component tests in
`desktop/src` call `setState` 744 times and a real store action 3 times. State you
assigned is self-consistent by construction and cannot expose "transition A did not
update B" — which is where these bugs live. Use `handleServerMessage`, store actions,
and real user events.
- **Assert the invariant, not today's output.** `2262973a4` shipped
`expect(getByText('deepseek-reasoner'))` at a moment when the screen showed another
model's number: it wrote the bug in as a passing assertion, and the next fix had to
invert that exact line. Ask what must be true after this step, not what it prints now.
- **Cover both directions of any rule that drops or merges something.** The replay guard
was tested for "a replay must be discarded" and never for "a genuine repeat must be
kept", so it shipped dropping real replies.
- **Test the join, not each end.** Server, store, and component each had a test for
`runtime_config_applied`; nothing crossed them, and deleting the term that joins them
(`ChatInput.tsx` `refreshNonce`) left 314 tests green.
- **Never retune an existing test's inputs to keep it green.** `128f75ab5` changed five
tests' props (`messageCount={0}` → `{1}`) instead of accepting that they described
states a real session cannot reach. If a test only passes after you edit its inputs,
the test was describing the implementation.
- **Do not mock the module under test.** A hand-written factory freezes an interface
snapshot: the store can be renamed or gutted and the test still passes.
- **If you are comparing content to decide identity, the identity exists upstream.**
Deduping by text cannot separate a replay from a legitimate repeat; forward the id
(`uuid`, `toolUseId`) instead of guessing.
Blind spots to check rather than trust:
- `desktop/electron/` is not instrumented at all (`vitest.config.ts` collects only
`desktop/src`), so main-process diffs score zero covered lines.
- Bun's LCOV emits no branch records, so `src/` and `adapters/` report **100% branch
coverage** for data that was never collected (`pct(0, 0) === 100`). Only `desktop/`
has real branch numbers.
- When publishing is explicitly requested, use Conventional Commit subjects and product branch prefixes such as `fix/`, `feat/`, or `docs/`; do not create `codex/` branches in this repository.
## Verification
1. Run the narrowest relevant test while iterating.
2. Run `bun run check:impact`; every command it selects is part of the minimum handoff for the current diff. Selection is import-aware: a change is routed to every surface that imports it, not only to its own directory. The report's `## Cross-surface impact` section names the importer that pulled in each extra check.
3. Run `bun run verify` only when full validation is requested or before claiming a code change is PR-ready or push-ready.
Additional invariants:
- Required PR checks must be deterministic and work on an untrusted fork: no real models, public network, repository secrets, saved providers, or real user home/config. Use fake credentials, fixtures, mocked/loopback transports, temporary directories, and explicit cleanup.
- `bun run check:agent-flow` is the deterministic end-to-end agent lane: it drives the real server and WebSocket through session creation, runtime selection, streaming, tool permission allow/deny, tool failure, API error, interrupt, reconnect replay, and session recovery using the repository's mock SDK CLI. It needs no provider, credentials, or network, so every contributor can run it.
- `bun run check:desktop-ui-smoke` drives the real desktop UI against that same mock runtime and answers the permission dialog by clicking the real button. It skips with a printed reason when `agent-browser` or desktop dependencies are missing.
- `agent-browser` is an implementation detail of that committed lane (which runs headless on Linux CI) and of the maintainer-run `desktop/scripts/e2e-*-agent-browser.sh` scripts. It is not the tool for ad-hoc browser work: manual verification, screenshots, and exploratory UI checks go through the `ego-browser` skill instead.
- Quality-gate lanes that boot the real server must run in a sandbox config dir (`scripts/quality-gate/sandbox.ts`) and fail if they wrote to the developer's real `~/.claude`.
- Provider/auth/proxy/runtime changes may select `bun run check:provider-contract`; desktop chat/WebSocket/session changes may select `bun run check:chat-contract`. These contracts are offline and do not replace their selected surface checks.
- Any persisted JSON, `localStorage`, or app-config shape change requires a forward migration, an old-fixture regression test, and `bun run check:persistence-upgrade`.
- User-visible desktop or cross-process behavior needs an actual browser/desktop smoke path when unit tests cannot prove the workflow.
- Live model checks are separate maintainer evidence. Run them only after deterministic checks pass and a maintainer explicitly authorizes quota use; finding credentials on the machine is not authorization.
- `bun run check:docs` runs `npm ci`; run it sequentially with checks that rely on root `node_modules`.
- `bun run check:impact` selects the required checks using paths and imports. Run the selected checks for the final diff; use focused tests while fixing failures. `package.json` and `scripts/pr/change-policy.ts` are the command and routing sources of truth.
- Use `bun run verify` when full validation is requested or before claiming a code change is PR-ready or push-ready. It runs the selected PR lanes, so there is no need to run every lane separately first. Reuse passing results for unchanged code; rerun or broaden checks when subsequent edits, failures, or unresolved risks warrant it.
- Required PR checks must be deterministic: no real models, public network, repository secrets, saved providers, or real user home/config. Use fake credentials, fixtures, mocked/loopback transports, temporary directories, and cleanup. Server-booting quality lanes use `scripts/quality-gate/sandbox.ts` and must fail on real user-state writes.
- For user-visible desktop or cross-process changes that unit tests cannot prove, exercise a browser/desktop smoke path. Ad-hoc browser work uses the `ego-browser` skill; `agent-browser` is reserved for the committed smoke lane and `desktop/scripts/e2e-*-agent-browser.sh`. See the [deterministic agent/UI lanes](docs/internals/contributing.md#无模型的端到端-agent-门禁).
- Live model checks are separate maintainer evidence after deterministic checks pass and quota use is explicitly authorized; finding credentials on the machine is not authorization.
## User-State Safety
- Never use or mutate the developer's real `~/.claude`, keychain, tokens, transcripts, providers, or project settings in tests. Redirect every relevant path to a temporary directory.
- Treat `~/.claude/settings.json` as user-owned shared state: preserve unknown fields, merge additively, and never add a repository-owned global schema marker.
- Repair/Doctor flows are deny-by-default. They may automatically change only explicitly allowlisted, regenerable desktop UI state; protected user data requires a reviewed, backup-first manual flow.
- Any persisted JSON, `localStorage`, or app-config shape change requires a forward migration, an old-fixture regression test, and `bun run check:persistence-upgrade`.
- Repair/Doctor flows are deny-by-default. Automatic repair may change only explicitly allowlisted, regenerable desktop UI state; protected user data requires a reviewed, backup-first manual flow.
## Handoff
- Review `git diff --check`, `git diff`, and `git status --short` before reporting completion.
- Report only evidence from the current worktree: changed files, tests added, commands actually run and their observed results, checks not run, blockers, and remaining risk.
- `passed`, `failed`, `skipped`, `blocked`, and `not run` are different states. A build is not E2E, a mock is not live-provider evidence, and an older report becomes stale after relevant edits.
## Deeper Guides
- Contributor workflow and quality lanes: `CONTRIBUTING.md` and `docs/internals/contributing.md`
- Package scripts and path routing: `package.json` and `scripts/pr/change-policy.ts`
- PR evidence contract: `.github/pull_request_template.md`
- Desktop release and auto-update runbook: `docs/desktop/10-release-auto-update.md`
- Report changed files, tests added, commands actually run and their observed results, checks not run, blockers, and remaining risk. Distinguish `passed`, `failed`, `skipped`, `blocked`, and `not run`; build-only, mock, live, and stale evidence are not interchangeable.
- Contributor workflow, failure diagnosis, and instruction maintenance: [CONTRIBUTING.md](CONTRIBUTING.md) and [detailed guide](docs/internals/contributing.md). PR evidence: [.github/pull_request_template.md](.github/pull_request_template.md).
+5 -3
View File
@@ -8,7 +8,7 @@
bun run check:impact
```
`check:impact` 会列出这次改动选中的检查。先运行对应的窄命令;准备声明 PR-ready、改动风险较高,或维护者需要复现完整 CI 时,再运行统一入口:
`check:impact` 会列出这次改动选中的检查。普通任务运行这些检查;准备声明 PR-ready 或需要完整验证时,直接运行统一入口,不必先单独重复执行其全部检查:
```bash
bun run verify
@@ -16,6 +16,8 @@ bun run verify
`bun run verify` 会按改动路径执行被选中的 policy、产品、契约、持久化和 coverage lane,但不会调用真实模型。小范围贡献无需在本机重复所有无关模块;GitHub PR gate 会再次执行并严格核对 selected / skipped 状态。
修复期间重跑受影响的窄检查,交付时补齐最终 diff 的验证证据;没有后续改动或未解决风险时,已通过的检查无需重跑。
`git push` 不再自动运行本地质量门禁。需要质量检查时请手动运行 `bun run quality:push` 或 `bun run verify`;完整覆盖率仍以 `bun run verify` 为准。
只改了某个模块时可以用窄命令快速迭代:
@@ -42,8 +44,8 @@ artifacts/quality-runs/<timestamp>/logs/<lane>.log
改动涉及用户可见 UI、跨 WebSocket/进程流程、Electron host 或 native/packaging 时,除了自动门禁外,还应在真机上验证相关流程。纯样式、纯工具或已有组件单元测试能够完整证明的改动,不要求重复无关流程:
- 起本地服务 `SERVER_PORT=3456 bun run src/server/index.ts`
- 起桌面端 `cd desktop && bun run dev`
- 优先使用 `bun run check:desktop-ui-smoke` 的隔离 mock runtime 验证真实桌面 UI
- 需要手动启动服务和桌面端时,复用 `scripts/quality-gate/sandbox.ts` 的临时配置和测试环境,避免服务读写真实 `~/.claude`;临时浏览器操作使用 `ego-browser` skill
- 验证改动涉及的交互流程:页面渲染、按钮/表单行为、弹窗、快捷键、多窗口等
- 必要时打本地 macOS 包 `desktop/scripts/build-macos-arm64.sh` 做完整验证
+4 -8
View File
@@ -1,6 +1,6 @@
# 桌面端组件规范
编辑 `desktop/src/components/` 下任何文件前先读本文。它是可复用组件的权威索引、新组件的放置规则,以及样式 / i18n / 无障碍 / 测试的强制约定。
本文提供可复用组件索引、放置规则及样式 / i18n / 无障碍 / 测试约定。按改动涉及的组件查阅对应小节,无需在每次编辑前重读全文。
规则优先级:根 `AGENTS.md` < `desktop/AGENTS.md` < 本文件。冲突时以本文件为准。
@@ -138,7 +138,7 @@ style={{ zIndex: 'var(--z-dialog)' }}
1. 先确认这不是"这处本就该定制"。列表项、树节点、标签页项、拖拽把手、OS 窗口按钮这些本来就不该塞进通用组件。
2. 如果是通用需求(同样的形态在多处出现),**给组件加一个 variant/tone/size**,连同测试一起。
3. 改不了或拿不准,**保持原样并记下来** —— 一个诚实的"跳过 + 原因"比一个视觉走样的替换有价值得多。
3. 在当前任务范围内补齐需要的组件能力,并验证调用方、交互和主题效果。只有缺少必要的产品决定或超出授权范围时才请求确认;无法完成的部分说明具体 blocker。
`IconButton` 的 `secondary` tone、`2xs`–`2xl` 六档尺寸、`bordered`、`solid`、`hoverTone="danger"`、`pressed`、`surface="sidebar"`、`surface="terminal"`(墨色终端标题栏,纸主题下 `--color-text-tertiary` 在其上不可见),`Button` 的 `base`(h-8)、`inverse` 与 `tonal-outline`(陶土描边 hover 反色)变体,`Card` 的 `shadow`/`lift`/`container` 档与 rest 透传,`Badge` 的 `wrap`/`title`/rest 透传,`SearchField` 的 `clearLabel` 与 `xl`(44px) 档 —— 全部来自这个流程。多轮独立的替换工作各自撞到同一批缺口,然后一次补齐。
@@ -233,13 +233,9 @@ cd desktop && bun run dev
---
## 八、提交前
## 八、验证与交付
```bash
cd desktop && bun run lint && bun run test -- --run
```
然后 `bun run check:impact`,桌面改动通常会选中 `bun run check:desktop`(lint + test + build 三步)。
迭代时运行相关组件测试;最终按根规则运行 `bun run check:impact` 选中的检查。`bun run check:desktop` 已包含 lint、完整 test 和 build,不必先重复跑一遍 lint / test。需要 PR-ready 或完整验证时使用 `bun run verify`;没有后续改动或未解决风险时,已通过的检查无需重跑。
自查清单:
+5 -5
View File
@@ -9,15 +9,15 @@ order: 3
子 Agent 就是 Claude 派出去的分身:给它一个明确的小任务,它自己带着独立上下文去干,干完只把结论交回来。
好处是主对话不会被一堆中间过程撑爆。比如「在这个仓库里找出所有调用 `validateUser` 的地方」,如果主 Agent 亲自去搜,几十个文件的内容全会挤进上下文;派个子 Agent 去,回来的只有一份清单。
子 Agent 可以把独立调查的中间过程留在自己的上下文中,再把结论和证据交回主对话。例如,调查一个模块的认证流程可以单独委派;只查找 `validateUser` 的调用位置,通常一次定向搜索就能完成,无需创建子 Agent。
## 什么时候该派
- **要翻很多文件才能回答的问题** — 找用法、理清依赖、统计某种模式在哪些地方出现。
- **可以并行的独立工作** — 前端一个、后端一个、测试一个,同时开工。
- **需要多个角度看同一件事** — 比如让几个 Agent 各自审一遍同一段代码。
- **可以独立完成的调查** — 需要阅读多个文件,并能明确约定问题、范围和应返回的证据。
- **可以并行的独立工作** — 比如分别调查前后端;涉及编辑时划清文件归属,由主 Agent 整合和验证。
- **有具体关注点的独立复核** — 比如分别检查权限边界和会话恢复,而非无差别地重复整轮审查。
反过来,你已经知道文件在哪、改哪一行,那就直接说,不用绕这一圈。
委派也有启动、传递上下文和汇总的成本。简单搜索、已定位的小改动,或下一步必须等待其结果的短任务,通常由主 Agent 直接完成更合适。
派出去的子 Agent 会出现在活动面板的「SubAgent」区块,工具活动实时冒泡,点进去能看它完整的运行记录和最终结果。后台跑的也一样,不用等它结束才知道在干什么。
+5 -5
View File
@@ -9,15 +9,15 @@ order: 3
A subagent is a copy of Claude sent off with one clearly scoped job. It works in its own context and reports back only the conclusion.
The point is that your main conversation doesn't get flooded. Ask "find every call site of `validateUser` in this repo" and, if the main agent searches itself, dozens of files end up in the context window. Delegate it and all that comes back is the list.
A subagent can keep an independent investigation's intermediate work in its own context, then return conclusions and evidence to the main conversation. Investigating a module's authentication flow can be delegated; finding call sites of `validateUser` usually takes one targeted search and does not need a subagent.
## When to delegate
- **Questions that require reading a lot of files** — finding usages, untangling dependencies, counting where a pattern occurs.
- **Independent work that can run in parallel** — frontend, backend, and tests at the same time.
- **The same thing from several angles** — several agents each reviewing the same code.
- **An investigation that can stand on its own** — it requires multiple files and has a clear question, scope, and expected evidence.
- **Independent work that can run in parallel** — such as separate frontend and backend investigations. For edits, assign file ownership and let the main agent integrate and verify the result.
- **An independent review with a specific focus** — such as separate checks of permission boundaries and session recovery, instead of repeating an entire review without a distinct purpose.
Conversely: if you already know the file and the line, just say so. No need for the detour.
Delegation also costs startup time, context transfer, and integration. The main agent should usually handle simple searches, small edits with a known location, or short tasks whose result is needed before anything else can proceed.
Delegated agents appear under **SubAgents** in the Activity panel with their tool activity streaming live. Open one to read its full transcript and final result. Background agents work the same way — you don't have to wait for them to finish to see what they're doing.
+26 -25
View File
@@ -13,14 +13,14 @@ Deconstructing the architecture behind the world's most popular AI code editor
## What This Framework Solves
Watch Claude Code closely and a few behaviors need explaining:
The runtime provides these execution mechanisms for the model:
- It can modify dozens of files in a single conversation with extremely few errors
- It automatically recovers from edge cases (token overflow, API timeouts, tool failures)
- It can simultaneously manage multiple subagents collaborating on complex tasks
- Long conversations don't degrade — they actually become more precise over time
- Execute multiple file edits in one conversation
- Recover from some token overflows, API timeouts, and tool failures
- Manage multiple subagents collaborating on complex tasks
- Trim tool results or generate summaries as context pressure grows
None of that comes from the model alone; it is designed into the framework. The rest of this page walks the source in order.
Task quality depends on the model, prompts, available tools, and runtime together. These mechanisms help work continue, but cannot guarantee correct edits or lossless long conversations. The rest of this page walks the source in order.
## The Core Agent Loop
@@ -67,16 +67,16 @@ The entire `while (true)` loop (`src/query.ts:307-1728`) consists of five phases
#### Phase 1: Message Preparation & Smart Compression (lines 365-543)
Before calling the API, conversation history goes through four layers of compression:
Before calling the API, the loop checks the available context processing paths; it does not necessarily perform all four kinds of compression on every turn:
| Compression Strategy | Mechanism | Trigger |
|---------------------|-----------|---------|
| **Snip Compression** | Smart deletion of redundant tokens in old messages | Every turn |
| **Micro Compression** | In-place modification of cached message content | Every turn |
| **Context Collapse** | Staged summarization of historical messages | When context nears limit |
| **Auto Compact** | Full summary generation via Claude | When context is critically low |
| **Snip Compression** | Entry point for trimming history | Requires `HISTORY_SNIP`; the module in this repository is a placeholder stub |
| **Micro Compression** | Clear some old tool results, or use cache editing | When the time threshold is met, or feature, model, and main-thread conditions hold |
| **Context Collapse** | Entry point for projecting historical summaries | Requires `CONTEXT_COLLAPSE`; the module in this repository is a placeholder stub |
| **Auto Compact** | Ask the model to generate a summary | When auto-compaction is enabled and the model-specific threshold is reached |
This is the key to Claude Code handling **extremely long conversations** without degradation — it doesn't simply truncate history, but **intelligently compresses while preserving critical information**.
Compression frees context space but can lose detail. Read files or original evidence again when a detail needs verification. See “Context Management & Compression” below for the available paths.
#### Phase 2: Streaming API Call (lines 652-954)
@@ -302,27 +302,27 @@ The model dynamically retrieves full definitions via the `ToolSearch` tool when
## Context Management & Compression
### The Secret Behind Unlimited Conversations
### Continuing Within a Finite Window
Claude Code claims "conversations have no context limit." Behind this is a **four-level compression system**:
The model's context window remains finite. `src/query.ts` contains four kinds of processing entry points, whose availability depends on build features, model, and configuration. The diagram does not mean every path is implemented or enabled in the current build.
![Context Compression Strategy](./images/14-context-compression.png)
#### Level 1: Snip Compression
Smart trimming of processed messages — removes duplicate file content, overly long tool outputs, etc.
`HISTORY_SNIP` gates the history-trimming entry point. The current `src/services/compact/snipCompact.ts` is a placeholder stub for external builds, so its presence does not establish that the app trims history through this path.
#### Level 2: Micro Compression
Modifies cached message content without changing the cache key. An "in-place optimization" strategy.
`src/services/compact/microCompact.ts` checks the time threshold first and clears some old tool results when the conditions hold. Cache editing separately requires `CACHED_MICROCOMPACT`, model support, and a main-thread source. When conditions do not hold, it can return messages unchanged and leave context pressure to Auto Compact.
#### Level 3: Context Collapse
Staged summarization of historical messages. Not all-at-once summarization, but **progressive folding** — summarize the oldest messages first, keeping recent details intact.
`CONTEXT_COLLAPSE` gates the historical-summary projection entry point. The current `src/services/contextCollapse/index.ts` is also a placeholder stub, so this entry point cannot be described as an enabled progressive summarization capability.
#### Level 4: Auto Compact
When all local optimizations are insufficient, Claude itself generates a complete conversation summary that replaces all historical messages.
When auto-compaction is enabled and its threshold is reached, `src/services/compact/autoCompact.ts` attempts a summary and continues with the compacted messages. Thresholds use the model window resolved by the runtime; consecutive failures stop automatic retries. Summaries cannot guarantee retention of every historical detail.
### System Context Injection
@@ -557,7 +557,7 @@ When images or other media cause token overflow:
| **Tool Execution** | After complete model response | During streaming |
| **State Management** | External Memory objects | Built-in state assignment + loop |
| **Error Recovery** | Manual orchestration required | 6 built-in recovery strategies |
| **Context Compression** | Simple truncation or summary | Four-level progressive compression |
| **Context Compression** | Simple truncation or summary | Trimming or summary paths enabled by build, model, and configuration |
| **Multi-Agent** | Chain/Graph explicit orchestration | Unified tool interface + state machine |
| **Extension Mechanisms** | Python class inheritance | Skills + Plugins + Hooks + MCP |
| **Caching Strategy** | None | Global / session / per-turn three-level cache |
@@ -621,13 +621,13 @@ From source code analysis, we can distill these core design principles:
### Streaming First
The entire architecture is designed around `AsyncGenerator` — everything is streamed:
The core loop uses `AsyncGenerator` to report progress incrementally:
- Model responses are streamed
- Tools execute during streaming
- Progress updates in real-time
- Compression strategies are progressive
Users **never have to wait** — they see the model thinking, tools executing, and results emerging.
Users can see model and tool progress before the task finishes. Model responses, tool execution, and compaction still take time.
### Intelligent Caching
@@ -645,7 +645,8 @@ This dramatically reduces latency and cost for every API call.
### Graceful Degradation
Six recovery strategies ensure Claude Code **almost never interrupts the user's workflow due to technical issues**:
The runtime provides several recovery paths. Whether work can continue depends on the error type, configuration, and retry limits:
- Token overflow? Auto-compress
- API timeout? Auto-retry
- Model failure? Fall back to alternate model
@@ -672,14 +673,14 @@ This avoids the "framework tax" — the abstraction layer that frameworks like L
### Tool-Driven Agent
Claude Code's philosophy: **an agent's capability equals the capability of its tools**.
Tool interfaces give the model entry points for taking action:
- Spawn a subagent? That's a tool (`AgentTool`)
- Manage a team? That's a tool (`TeamCreate`/`SendMessage`)
- Edit a file? That's a tool (`FileEdit`)
- Execute a skill? That's a tool (`SkillTool`)
**All capabilities are exposed through the unified tool interface**, and the model uses natural language reasoning to decide which tool to use. No explicit orchestration logic needed — the model itself is the orchestrator.
The model selects tools based on the task and context; the runtime handles scheduling, permission checks, and state updates. Tools define the available actions, while the model and prompts affect how those actions are selected, combined, and verified.
### Deep Developer Experience Integration
+35 -20
View File
@@ -71,7 +71,7 @@ Neither needs a provider, credentials, or the public network. `check:agent-flow`
Every quality-gate lane that boots the real server runs against a sandbox config dir (`scripts/quality-gate/sandbox.ts`) and fails if it wrote to the developer's real `~/.claude`.
Run the selected focused commands while developing. Before claiming PR-ready, for a high-risk change, or when reproducing the full hosted CI locally, use the unified entrypoint:
Run the selected focused commands while developing. For PR-ready or full validation, use the unified entrypoint directly without first running all of its lanes separately:
```bash
bun run verify
@@ -96,24 +96,33 @@ The coverage gate does four things: measures source-only coverage, enforces the
## AI Coding Agent Fix Loop
When asking an AI coding agent to work in this repo, use this as the acceptance instruction:
Completion means implementing the intended behavior, running the checks required for the current diff, and fixing failures caused by the change. Scoped local edits, isolated fixture checks, and related repairs do not need approval at each step. Commits, pushes, releases, repository settings, and live-model quota still follow the root `AGENTS.md` authorization boundaries.
```text
Run `bun run check:impact`, then run the selected focused checks. If the task
requires PR-ready/full validation, run `bun run verify`. If it fails, read the latest
`artifacts/quality-runs/<timestamp>/report.md` and the relevant lane log,
fix the missing tests, coverage failures, type/lint/build errors, or docs/native
failures, then rerun `bun run verify` until it passes. Do not lower coverage
baselines or thresholds unless a maintainer explicitly requested it.
```
Use `bun run check:impact` to determine the check scope. Run the selected checks for ordinary tasks; use `bun run verify` directly for PR-ready/full validation without first running all of its lanes separately. During repairs, rerun affected focused checks, then complete the evidence needed for the final diff. Do not repeat passing checks without subsequent edits or unresolved risks. Report unrelated existing failures or environment blockers instead of expanding the change merely to make everything green.
Agents should handle failures in this order:
When a check fails, consult the evidence for that failure:
1. Start with the Summary and Result Matrix in `artifacts/quality-runs/<timestamp>/report.md` to identify the failing lane.
2. If `Path-aware PR checks` failed, check for missing same-area tests, CLI core changes, or coverage policy changes. Do not bypass normal feature PRs with maintainer overrides.
3. If `Coverage gate` failed, open `artifacts/coverage/<timestamp>/coverage-report.md` or `coverage-report.json`, then fix `changedLines.failures` and `failures` first. `targetGaps` are technical-debt signals; touched areas should still improve.
4. If desktop/server/adapters/native/docs failed, read `artifacts/quality-runs/<timestamp>/logs/<lane>.log`, add tests or fix the build, then rerun the narrow command.
5. After narrow checks pass, run `bun run verify` when claiming PR-ready/full validation. The agent may only make that claim when the final Summary has `failed=0`.
| Failure | Evidence and action |
| --- | --- |
| Failed lane | Summary / Result Matrix in `artifacts/quality-runs/<timestamp>/report.md` and `logs/<lane>.log` |
| Path-aware PR checks | Check same-area tests, CLI core, and coverage policy; maintainer overrides require an explicit decision |
| Coverage gate | `artifacts/coverage/<timestamp>/coverage-report.md` or `.json`; address `changedLines.failures` / `failures`, while `targetGaps` signal technical debt |
| Build, types, lint, docs, or native | Fix issues caused by the change identified in the relevant log and rerun affected checks |
Claim PR-ready/full validation only after `bun run verify` passes for the final diff. Do not lower coverage baselines/thresholds or rewrite test expectations to hide failures.
## Regression Test Design
A same-area test file is the gate's minimum signal; tests also need to prove behavior:
- **Drive state transitions.** When testing a transition, produce the state through `handleServerMessage`, real store actions, or user events instead of directly assigning the expected result with `setState`. Direct state setup is still appropriate for fixture initialization.
- **Assert behavioral invariants.** Check which session or model the displayed data belongs to, rather than copying today's screen text. Test inputs and expectations must follow the intended behavior contract; do not change them to hide failures.
- **Cover both dropping and keeping.** Test what a deduplication, merging, or filtering rule should discard and retain. Message deduplication in particular must reject replays and preserve legitimate repeats; forward upstream identities such as `uuid` / `toolUseId` instead of guessing identity from text.
- **Test the connections across boundaries.** Separate green server, store, and component tests do not prove that messages drive the UI. Exercise risky connections through real entry points and do not mock the module under test.
Coverage reports have limits: `desktop/vitest.config.ts` collects only `src/**`, excluding the Electron main process. The repository's current Bun coverage baseline has zero branch records, and `coverage.ts` displays `0/0` as 100%; that does not prove all branches were exercised. Inspect current configuration and reports instead of using historical coverage figures as evidence for a new change.
### Coverage References
External reference points:
@@ -121,14 +130,20 @@ External reference points:
- [Microsoft Visual Studio / Azure DevOps docs](https://learn.microsoft.com/en-us/visualstudio/test/using-code-coverage-to-determine-how-much-code-is-being-tested): teams typically target about 80%, typical project requirements can be 75%, and generated code may be relaxed.
- [ChromiumOS EC](https://chromium.googlesource.com/chromiumos/platform/ec/+/main/docs/code_coverage.md): new or changed lines require at least 80% coverage.
## Maintaining Agent Instructions
Keep project constraints and entry points in root `AGENTS.md`, specialized rules near the code, and explanations/examples in on-demand documentation. Shared guidance must work for contributors using different models. Revisit duplicated workflows and broad stopping conditions as capabilities change, while preserving current safety and CI contracts. This cleanup draws on Eric Provencher's [Rethinking skills and prompts for GPT-6 Astra](https://x.com/pvncher/status/2095991462416490862) (2026-09-04).
Repository skill descriptions should identify the applicable task and necessary distinctions; put operational detail in the body or referenced files. Use a short router for multiple workflows and avoid broadening triggers just to match more keywords. Model defaults, tool formats, and compaction behavior describe product implementation, so check the source before updating those docs.
## Feature Quality Contract
Every feature, bugfix, and behavior change must ship with verifiable evidence. This rule applies to human authors and AI coding agents:
- Name the changed surface first: `desktop`, `server`, `adapter`, `native`, `docs`, `provider/runtime`, `agent-loop`, or `release`.
- Production changes under `desktop/src`, `src/server`, `src/tools`, `src/utils`, or `adapters` must include same-area tests in the same PR unless a maintainer explicitly applies `allow-missing-tests`.
- Executable JS/TS production changes must include same-area tests in the same PR. `scripts/pr/change-policy.ts` checks four areas separately: `desktop/src/`, `src/server/`, the rest of `src/`, and `adapters/`, unless a maintainer explicitly applies `allow-missing-tests`. Non-executable files such as prose or CSS do not independently require new tests under this rule; all impact-selected checks still apply.
- Pure logic needs unit tests. Server/API/provider/runtime behavior needs API or request-shape tests. Desktop UI/store/API behavior needs Vitest or Testing Library coverage. Cross-boundary user flows through UI, WebSocket, provider proxying, native sidecars, or release packaging need E2E or desktop UI smoke.
- Agent loop, tool execution, provider routing, model selection, file editing, permissions, session resume, and desktop chat changes need mock/fixture tests in PR, plus live smoke or baseline evidence when provider access is available.
- Agent loop, tool execution, provider routing, model selection, file editing, permissions, session resume, and desktop chat changes need mock/fixture tests in PR. Run live smoke or baseline only after deterministic checks pass and a maintainer explicitly authorizes quota use. Finding a local provider is not authorization; report when live checks were not run.
- Coverage is part of the feature. This project follows a Google/Microsoft-style policy: generated/build output is not counted as product coverage, maintained product areas should move toward 75-80%+, and new or changed executable production lines must pass the changed-line coverage threshold in `coverage-thresholds.json`.
- Do not lower `coverage-baseline.json` or `coverage-thresholds.json` just to pass the gate; real baseline/threshold changes require `allow-coverage-baseline-change` and a reason. Legacy low-coverage areas are debt; new PRs must leave touched areas better than they found them.
- The PR description must record changed files, tests added, coverage report path, E2E/live report path or blocker, and remaining risk.
@@ -187,7 +202,7 @@ bun run check:coverage # Root, desktop, and adapter coverage reports plus rat
Focused tests are the normal development loop. Run `bun run verify` locally when claiming PR-ready/full validation; hosted CI still executes every selected required lane.
Production code changes must include matching tests. Changes under `desktop/src/**`, `src/server/**`, `src/tools/**`, `src/utils/**`, or `adapters/**` without a same-area test file are blocked unless a maintainer applies `allow-missing-tests`. Coverage baseline/threshold changes are also blocked unless a maintainer applies `allow-coverage-baseline-change`.
Executable JS/TS production changes must include matching tests. See the Feature Quality Contract above and `scripts/pr/change-policy.ts` for area boundaries; missing same-area tests block the change unless a maintainer applies `allow-missing-tests`. Coverage baseline/threshold changes are also blocked unless a maintainer applies `allow-coverage-baseline-change`.
## Live Model Baseline
@@ -269,7 +284,7 @@ Before a release, run release mode:
bun run quality:gate --mode release --allow-live --provider-model <selector>:main
```
Release mode composes PR checks, baseline catalog validation, live baseline cases, provider smoke, native checks, and current-platform canonical release `package-smoke --package-kind release`. Reports are written to `artifacts/quality-runs/<timestamp>/`. The hosted release workflow now runs `bun run verify` as a non-live preflight before the packaging matrix; maintainers still need to run the live release gate explicitly with an available provider.
Release mode composes PR checks, baseline catalog validation, live baseline cases, provider smoke, native checks, and current-platform canonical release `package-smoke --package-kind release`. Reports are written to `artifacts/quality-runs/<timestamp>/`. `release-desktop.yml` builds and publishes artifacts without running `bun run verify`. Pre-release quality evidence comes from PR gates and explicitly run maintainer full checks and release gates.
In release mode, live lanes are not allowed to be silently skipped. Missing providers, model quota, or external account access will fail the gate and must be recorded as a release blocker.
+6 -3
View File
@@ -349,12 +349,13 @@ export const MAX_LISTING_DESC_CHARS = 250 // Max characters per descrip
```
formatCommandsWithinBudget(commands, contextWindowTokens)
├─ Cap description + whenToUse at 250 characters for every entry, including Bundled
├─ Calculate total budget = contextWindowTokens × 4 × 1%
├─ Try full descriptions
├─ Try retaining these already capped descriptions
│ └─ Total chars ≤ budget → output all
│
├─ Partition: Bundled (never truncated) + rest
│ ├─ Bundled Skills always retain full descriptions
├─ Partition: Bundled (no further truncation for the total budget) + rest
│ ├─ Bundled Skills retain their descriptions after the 250-character cap
│ └─ Remaining Skills split the leftover budget evenly
│
├─ Truncate descriptions → maxDescLen characters
@@ -363,6 +364,8 @@ formatCommandsWithinBudget(commands, contextWindowTokens)
└─ Output format: "- skill-name: description..."
```
The listing supports discovery; the body loads when the skill is invoked. Start descriptions with specific trigger conditions and purpose. Put steps, examples, and references in the body, and read them as the task requires. Avoid putting the entire workflow in the description or requiring every reference to be read on every invocation.
### SkillTool Prompt
The tool prompt definition seen by the model:
+3 -3
View File
@@ -61,12 +61,12 @@ To spend nothing and stay offline, run a model server on your machine and point
**Ollama**: run `ollama serve`, choose the `Ollama` preset, and use base URL `http://localhost:11434`.
Two hard requirements:
Check two things when connecting:
1. **Do not append `/v1` to the base URL.** Both of these expose an Anthropic-compatible protocol and the app uses that path. Adding `/v1` gives you a straight 404.
2. **Raise the context window — at least 200K.** Claude Code's system prompt, tool definitions, and Skills consume a substantial amount of context before your first message. A default 4K or 8K window can't even hold the opening. Change this in LM Studio or Ollama's own model settings, not in this app.
2. **Configure the context window to match what the model actually supports.** 200K is not a universal minimum. The window needs to hold the system prompt, tool definitions, loaded Skills, and current task. A smaller window may not fit that content, and longer tasks trigger compaction more often. Set the server's actual window in LM Studio or Ollama's model settings; do not declare a value beyond what it supports.
Whether a local model can actually sustain an agent workflow comes down to its tool-calling ability. Small models often talk endlessly without ever calling a tool — that's the model, not your configuration.
Whether a local model can complete an agent workflow depends on its tool-calling ability, API compatibility, configuration, and task. If it outputs text without calling tools, inspect the requests, responses, and tool configuration rather than inferring the cause from model size alone.
## The Add Provider dialog, field by field
+26 -25
View File
@@ -13,14 +13,14 @@ order: 6
## 这套框架要解决什么
观察 Claude Code 的行为,有几件事值得解释:
这套运行时为模型提供以下执行机制:
- 它能在一次对话中修改几十个文件,且极少出错
- 它能自动恢复各种边界情况(token 溢出、API 超时、工具失败)
- 它能同时管理多个子 Agent 协作完成复杂任务
- 长对话不会退化,反而能越来越精准
- 在一次对话中执行多步文件修改
- 对部分 token 溢出、API 超时和工具失败进行恢复
- 管理多个子 Agent 协作完成复杂任务
- 在上下文压力增大时裁剪工具结果或生成摘要
这些能力都不是模型自带的,而是框架设计出来的。下面按源码顺序拆开看。
任务质量取决于模型、提示词、可用工具和运行时的共同作用。这些机制支持任务持续执行,但不能保证修改正确或长对话不丢失信息。下面按源码顺序拆开看。
## 核心 Agent 循环
@@ -67,16 +67,16 @@ type State = {
#### 阶段 1:消息准备与智能压缩(第 365-543 行)
在调用 API 之前,对话历史会经过四层压缩处理:
调用 API 之前,循环会检查可用的上下文处理路径;并非每轮都会实际执行四种压缩:
| 压缩策略 | 原理 | 触发时机 |
|----------|------|----------|
| **Snip 压缩** | 智能删除旧消息中的冗余 token | 每轮自动 |
| **Micro 压缩** | 修改已缓存消息的内容 | 每轮自动 |
| **上下文折叠** | 分阶段摘要历史消息 | 上下文接近限制时 |
| **Auto Compact** | 通过 Claude 生成完整摘要 | 上下文严重不足时 |
| **Snip 压缩** | 裁剪历史消息的入口 | 需要 `HISTORY_SNIP`;当前仓库对应模块是占位 stub |
| **Micro 压缩** | 清空部分旧工具结果,或使用缓存编辑路径 | 满足时间阈值,或特性、模型与主线程条件时 |
| **上下文折叠** | 投影历史摘要的入口 | 需要 `CONTEXT_COLLAPSE`;当前仓库对应模块是占位 stub |
| **Auto Compact** | 调用模型生成摘要 | 自动压缩启用且达到模型对应阈值时 |
这是 Claude Code 能处理**极长对话**而不退化的关键——它不会简单地截断历史,而是**智能地压缩和保留关键信息**。
压缩可以释放上下文空间,也可能丢失细节;需要核实的内容仍应回到文件或原始证据中读取。具体路径见下文「上下文管理与压缩」。
#### 阶段 2:流式 API 调用(第 652-954 行)
@@ -302,27 +302,27 @@ Claude Code 有 48+ 个内置工具。如果每次 API 调用都把所有工具
## 上下文管理与压缩
### 无限对话的秘密
### 有限窗口中的对话延续
Claude Code 宣称"对话没有上下文限制",这背后是一套**四级压缩系统**:
模型的上下文窗口仍然有限。`src/query.ts` 中保留了以下四类处理入口,其可用性取决于构建特性、模型和配置;图中的入口不代表当前构建全部实现或启用。
![上下文压缩策略](./images/14-context-compression.png)
#### 第 1 级:Snip 压缩
对已处理的消息进行智能裁剪——移除重复的文件内容、过长的工具输出等。
`HISTORY_SNIP` 控制历史裁剪入口。当前 `src/services/compact/snipCompact.ts` 是外部构建占位 stub,不能据此认为应用已经执行了历史裁剪。
#### 第 2 级:Micro 压缩
修改已缓存消息的内容,而不改变缓存键。这是一种"原地优化"策略。
`src/services/compact/microCompact.ts` 先检查时间阈值,满足条件时清空部分旧工具结果。缓存编辑路径另受 `CACHED_MICROCOMPACT`、模型支持和主线程来源限制;条件不满足时可以原样返回消息,并由 Auto Compact 处理上下文压力。
#### 第 3 级:上下文折叠(Context Collapse)
将历史消息分阶段摘要。不是一次性摘要全部,而是**渐进式折叠**——先摘要最旧的消息,保留最近的细节。
`CONTEXT_COLLAPSE` 控制历史摘要投影入口。当前 `src/services/contextCollapse/index.ts` 同样是占位 stub,不能把这一入口描述为已启用的渐进式摘要能力。
#### 第 4 级:Auto Compact
当所有局部优化都不够时,通过 Claude 自身生成一个完整的对话摘要,替换所有历史消息。
自动压缩启用且达到阈值时,`src/services/compact/autoCompact.ts` 尝试生成摘要,并用压缩后的消息继续对话。阈值使用运行时解析的模型窗口;连续失败会停止自动重试。摘要不能保证保留全部历史细节。
### 系统上下文注入
@@ -557,7 +557,7 @@ if (error.type === 'prompt_too_long') {
| **工具执行** | 等待模型完整响应后执行 | 流式传输中即时执行 |
| **状态管理** | 外部 Memory 对象 | 内置状态赋值 + 循环 |
| **错误恢复** | 需要手动编排 | 6 种内置恢复策略 |
| **上下文压缩** | 简单截断或摘要 | 四级渐进式压缩 |
| **上下文压缩** | 简单截断或摘要 | 按构建、模型和配置启用裁剪或摘要路径 |
| **多 Agent** | Chain/Graph 显式编排 | 统一工具接口 + 状态机 |
| **扩展机制** | Python 类继承 | 技能 + 插件 + 钩子 + MCP |
| **缓存策略** | 无 | 全局/会话/按轮三级缓存 |
@@ -621,13 +621,13 @@ Claude Code 的优势在于**简单性**——不需要定义图结构,一个
### 流式优先(Streaming First)
整个架构围绕 `AsyncGenerator` 设计,一切都是流式的:
核心循环通过 `AsyncGenerator` 逐步产出进度:
- 模型响应是流式的
- 工具在流式中执行
- 进度实时更新
- 压缩策略是渐进式的
这意味着用户**永远不需要等待**——看到模型在思考、工具在执行、结果在产出。
用户可以在任务完成前看到模型和工具的进度;模型响应、工具执行和压缩仍需要时间。
### 智能缓存(Intelligent Caching)
@@ -645,7 +645,8 @@ Section Cache(轮级) ← systemPromptSection 记忆化
### 优雅降级(Graceful Degradation)
6 种恢复策略确保 Claude Code **几乎不会因为技术问题中断用户的工作流**:
运行时提供多种恢复路径,实际能否继续取决于错误类型、配置和重试限制:
- Token 超限?自动压缩
- API 超时?自动重试
- 模型失败?降级到备用模型
@@ -672,14 +673,14 @@ Claude Code 直接使用 Anthropic API 的原生能力:
### 工具驱动的 Agent(Tool-Driven Agent)
Claude Code 的哲学是:**Agent 的能力等于其工具的能力**。
工具接口为模型提供执行操作的入口:
- 子 Agent 生成?是一个工具(`AgentTool`)
- 团队管理?是一个工具(`TeamCreate`/`SendMessage`)
- 文件编辑?是一个工具(`FileEdit`)
- 技能执行?是一个工具(`SkillTool`)
这意味着**所有能力都通过统一的工具接口暴露**,模型通过自然语言推理来决定使用哪个工具。不需要显式的编排逻辑——模型本身就是编排器。
模型根据任务和上下文选择工具;运行时负责工具调度、权限检查和状态更新。工具决定可执行的操作范围,模型与提示词影响如何选择、组合和验证这些操作。
### 深度集成的开发体验
+35 -20
View File
@@ -71,7 +71,7 @@ bun run check:desktop-ui-smoke # 真实桌面 UI + 真实权限对话框 + mock
所有会启动真实 server 的 quality-gate lane 都跑在沙箱配置目录里(`scripts/quality-gate/sandbox.ts`),并在结束时校验没有写过开发者真实的 `~/.claude`;写了就判定 lane 失败。
开发时运行 impact report 选中的窄命令即可。准备声明 PR-ready、改动风险较高,或需要完整复现托管 CI 时,再运行统一入口:
开发时运行 impact report 选中的窄命令即可。需要声明 PR-ready 或完整验证时,直接使用统一入口,无需先单独执行其全部 lane:
```bash
bun run verify
@@ -96,24 +96,33 @@ PR 描述里请贴出你实际运行的命令和 summary。`quality:pr` / `quali
## AI Coding Agent 修复循环
给 AI 写代码时,可以直接把这段作为验收指令:
任务的完成标准是实现目标行为、运行当前 diff 必需的验证,并修复由本次改动造成的失败。任务范围内的本地编辑、隔离 fixture 检查和相关失败修复无需逐步请求批准;提交、推送、发布、仓库设置及真实模型额度仍遵循根 `AGENTS.md` 的授权边界。
```text
Run `bun run check:impact`, then run the selected focused checks. If the task
requires PR-ready/full validation, run `bun run verify`. If it fails, read the latest
`artifacts/quality-runs/<timestamp>/report.md` and the relevant lane log,
fix the missing tests, coverage failures, type/lint/build errors, or docs/native
failures, then rerun `bun run verify` until it passes. Do not lower coverage
baselines or thresholds unless a maintainer explicitly requested it.
```
`bun run check:impact` 用于确定检查范围;普通任务运行选中的检查,需要 PR-ready/full validation 时直接使用 `bun run verify`,不必先单独重复执行其全部 lane。修复期间先重跑受影响的窄检查,交付时补齐最终 diff 的所需证据。没有后续改动或未解决风险时,不要反复运行已经通过的检查。无关的现有失败或环境阻塞应准确报告,而不是为了让结果变绿擅自扩大修改范围。
Agent 应按这个顺序处理失败:
需要诊断失败时,按失败类型查阅相应证据:
1. 先看 `artifacts/quality-runs/<timestamp>/report.md` 的 Summary 和 Result Matrix,定位失败 lane。
2. 如果是 `Path-aware PR checks` 失败,优先看是否缺同区域测试、是否动了 CLI core、是否动了 coverage policy;不要用 override 绕过普通功能 PR。
3. 如果是 `Coverage gate` 失败,打开 `artifacts/coverage/<timestamp>/coverage-report.md` 或 `coverage-report.json`,优先修 `changedLines.failures` 和 `failures`;`targetGaps` 是技术债提示,新改动应让触达区域变好。
4. 如果是 desktop/server/adapters/native/docs 失败,读对应 `artifacts/quality-runs/<timestamp>/logs/<lane>.log`,补测试或修构建,再跑相关窄命令。
5. 窄命令通过后,如果要声明 PR-ready/full validation,再跑一次 `bun run verify`。只有最终 Summary 是 `failed=0`,才可以这样声明。
| 失败类型 | 证据与处理 |
| --- | --- |
| Lane 失败 | `artifacts/quality-runs/<timestamp>/report.md` 的 Summary / Result Matrix,以及 `logs/<lane>.log` |
| Path-aware PR checks | 核对同区域测试、CLI core 和 coverage policy;维护者 override 需要明确决定 |
| Coverage gate | `artifacts/coverage/<timestamp>/coverage-report.md` 或 `.json`;修复 `changedLines.failures` / `failures`,`targetGaps` 是技术债提示 |
| 构建、类型、lint、文档或 native | 修复对应日志指出的本次变更问题,重跑受影响检查 |
只有最终 diff 的 `bun run verify` 报告通过,才能声明 PR-ready/full validation。不要通过降低 coverage baseline/threshold 或改写测试预期掩盖失败。
## 回归测试设计
同区域测试文件是门禁的最低信号,测试还需要证明实际行为:
- **驱动状态迁移。** 需要验证迁移时,通过 `handleServerMessage`、真实 store action 或用户事件产生状态,避免直接 `setState` 写出本应由迁移生成的结果。初始化 fixture 仍可直接设置状态。
- **断言行为不变量。** 断言用户应看到哪个会话或模型的数据,而不是抄下当前屏幕的字符串;测试输入和预期应由预期行为契约支撑,不为掩盖失败而修改。
- **覆盖丢弃与保留两个方向。** 验证去重、合并和过滤规则应丢弃及应保留的情况。消息去重尤其要拦住 replay、保留真实重复;透传上游 `uuid` / `toolUseId` 等身份,避免用文本猜测身份。
- **跨边界测试连接点。** server、store、component 分别通过并不能证明消息真正驱动了 UI;通过真实入口验证有风险的连接,避免 mock 被测模块本身。
覆盖率报告也有边界:`desktop/vitest.config.ts` 只采集 `src/**`,不包含 Electron main process。仓库现有 Bun coverage baseline 中分支总数为 0,`coverage.ts` 把 `0/0` 显示为 100%;这不代表测到了全部分支。查看当前配置和报告,不把历史覆盖率数字当作新改动的证明。
### 覆盖率参考
外部参考口径:
@@ -121,14 +130,20 @@ Agent 应按这个顺序处理失败:
- [Microsoft Visual Studio / Azure DevOps 文档](https://learn.microsoft.com/en-us/visualstudio/test/using-code-coverage-to-determine-how-much-code-is-being-tested):团队通常以约 80% 为目标,典型项目要求可为 75%,生成代码可以放宽。
- [ChromiumOS EC](https://chromium.googlesource.com/chromiumos/platform/ec/+/main/docs/code_coverage.md):新增或变更行要求至少 80% 覆盖。
## 维护 Agent 指导
根 `AGENTS.md` 保留项目约束与入口,专项规则留在对应目录,解释和示例按需放到文档。共享指导应适用于贡献者使用的不同模型;能力升级后重新核对重复流程和宽泛停止条件,不能据此跳过现行安全或 CI 契约。这次整理参考了 Eric Provencher 的 [Rethinking skills and prompts for GPT-6 Astra](https://x.com/pvncher/status/2095991462416490862)(2026-09-04)。
仓库技能的描述只写适用任务和必要的区分信息,操作细节放正文或引用文件。多工作流技能用短入口路由;避免为了覆盖更多关键词而扩大触发范围。模型默认值、工具格式和压缩行为属于产品实现,更新相关文档前应先核对源码。
## Feature Quality Contract
所有新功能、bugfix 和行为变化都必须带着可验证证据交付。这条规则同时约束人和 AI Coding Agent:
- 先声明变更面:`desktop`、`server`、`adapter`、`native`、`docs`、`provider/runtime`、`agent-loop` 或 `release`。
- `desktop/src`、`src/server`、`src/tools`、`src/utils`、`adapters` 下的生产代码变更必须同 PR 带同区域测试;除非维护者显式加 `allow-missing-tests`。
- 可执行 JS/TS 生产代码变更必须同 PR 带同区域测试;`scripts/pr/change-policy.ts` 分别检查 `desktop/src/`、`src/server/`、其余 `src/` 和 `adapters/` 四个区域,除非维护者显式加 `allow-missing-tests`。文案或 CSS 等非可执行文件不因这一规则单独要求新增测试,仍需完成 impact 选中的检查。
- 纯逻辑写单元测试;server/API/provider/runtime 写 API 或 request-shape 测试;桌面 UI/store/API 写 Vitest/Testing Library;跨 UI、WebSocket、provider proxy、native sidecar、发布打包的用户流程要补 E2E 或桌面 UI smoke。
- agent loop、工具调用、provider 路由、模型选择、文件编辑、权限、会话恢复、桌面聊天改动,PR 内必须有 mock/fixture 测试;有 provider 条件时还要给 live smoke 或 baseline 证据。
- agent loop、工具调用、provider 路由、模型选择、文件编辑、权限、会话恢复、桌面聊天改动,PR 内必须有 mock/fixture 测试;live smoke 或 baseline 仅在确定性检查通过且维护者明确授权额度后运行。发现本机 provider 不代表获得授权,未运行时如实说明。
- 覆盖率是功能的一部分。本项目按 Google/Microsoft 风格执行:生成物/构建产物不计入产品覆盖率,维护中的产品区域要逐步达到 75-80%+,新增或变更的可执行生产代码行必须满足 `coverage-thresholds.json` 里的 changed-line coverage 门槛。
- 不要为了过门禁随便降低 `coverage-baseline.json` 或 `coverage-thresholds.json`;确实要改时必须有 `allow-coverage-baseline-change` 和原因。历史低覆盖区域是技术债,新 PR 至少要让触达区域更好。
- PR 描述必须写清楚:改了哪些文件、补了哪些测试、coverage 报告路径、E2E/live 报告路径或 blocker、剩余风险。
@@ -187,7 +202,7 @@ bun run check:coverage # root、desktop、adapters 覆盖率报告和 ratchet
如果只改了很窄的文件,先跑对应的定向测试即可;只有在声明 PR-ready/full validation 时才需要本地再跑 `bun run verify`,托管 CI 仍会执行所有被选中的必需 lane。
生产代码改动必须带对应测试文件:`desktop/src/**`、`src/server/**`、`src/tools/**`、`src/utils/**`、`adapters/**` 变更如果没有同区域测试,会触发阻断。只有维护者确认不适合自动化测试时,才能使用 `allow-missing-tests`。覆盖率 baseline/threshold 变更同样需要维护者确认并加 `allow-coverage-baseline-change`。
可执行 JS/TS 生产代码改动必须带对应测试文件;同区域划分见上文 Feature Quality Contract 和 `scripts/pr/change-policy.ts`,缺失时会触发阻断。只有维护者确认不适合自动化测试时,才能使用 `allow-missing-tests`。覆盖率 baseline/threshold 变更同样需要维护者确认并加 `allow-coverage-baseline-change`。
## 真实模型 Baseline
@@ -269,7 +284,7 @@ bun run quality:gate --mode baseline --allow-live
bun run quality:gate --mode release --allow-live --provider-model <selector>:main
```
release 模式会组合 PR checks、baseline catalog、live baseline、native checks,并用当前平台 canonical release artifact 跑 `package-smoke --package-kind release`。发版报告同样写入 `artifacts/quality-runs/<timestamp>/`。线上 release workflow 在打包矩阵前会先跑 `bun run verify` 作为非 live 预检;真实 live release gate 仍需要维护者用可用 provider 显式运行。
release 模式会组合 PR checks、baseline catalog、live baseline、native checks,并用当前平台 canonical release artifact 跑 `package-smoke --package-kind release`。发版报告同样写入 `artifacts/quality-runs/<timestamp>/`。`release-desktop.yml` 只负责构建与发布,不运行 `bun run verify`;发版前质量证据来自 PR 门禁及维护者显式运行的全量检查和 release gate。
release 模式下 live lane 不允许静默跳过。缺少 provider、真实模型额度或外部账号时,门禁会失败,并要求在发版记录里明确 blocker。
+6 -3
View File
@@ -349,12 +349,13 @@ export const MAX_LISTING_DESC_CHARS = 250 // 每条描述上限
```
formatCommandsWithinBudget(commands, contextWindowTokens)
├─ 所有条目的 description + whenToUse 先限制为 250 字符(含 Bundled)
├─ 计算总预算 = contextWindowTokens × 4 × 1%
├─ 尝试全量描述
├─ 尝试保留上述已限长描述
│ └─ 总字符 ≤ 预算 → 全部输出
│
├─ 分区: Bundled(不截断) + 其余
│ ├─ Bundled Skills 始终保留完整描述
├─ 分区: Bundled(不再按总预算截断) + 其余
│ ├─ Bundled Skills 保留经过 250 字符限制的描述
│ └─ 其余 Skills 平分剩余预算
│
├─ 截断描述 → maxDescLen 字符
@@ -363,6 +364,8 @@ formatCommandsWithinBudget(commands, contextWindowTokens)
└─ 输出格式: "- skill-name: description..."
```
技能列表用于发现,正文在调用时加载。编写描述时先写清具体触发条件和用途;将操作步骤、例子和参考文件放在正文,按任务需要读取。不要把整套工作流塞进描述,也不要要求每次调用都读完所有参考文件。
### SkillTool Prompt
模型看到的工具提示词定义:
+3 -3
View File
@@ -61,12 +61,12 @@ FennoAI 和七牛云 AI 能用哪些模型取决于你买的套餐,所以这
**Ollama**:`ollama serve` 起来之后,选 `Ollama` 预设,接口地址填 `http://localhost:11434`。
两条硬性要求:
连接时注意两点:
1. **接口地址后面不要加 `/v1`。** 这两家都提供 Anthropic 兼容协议,应用走的是那条路径,多加 `/v1` 会直接 404。
2. **把上下文窗口调大,建议至少 200K。** Claude Code 的系统提示词、工具定义和 Skills 本身就要吃掉不少上下文,默认的 4K/8K 窗口连开场都放不下。这个设置在 LM Studio 或 Ollama 自己的模型配置里改,不在本应用里。
2. **按模型实际支持的大小配置上下文窗口。** 200K 不是统一的最低门槛;窗口需要容纳系统提示词、工具定义、已加载的 Skills 和当前任务。较小窗口可能容纳不下这些内容,长任务也更容易触发压缩。服务端的实际窗口在 LM Studio 或 Ollama 自己的模型配置里设置,不要声明超出实际支持范围的数值。
本地模型能不能撑住完整的 Agent 工作流,取决于模型本身的工具调用能力。小参数量模型经常出现"一直说话但不动手"的情况——那是模型的问题,不是配置的问题。
本地模型能否完成 Agent 工作流,取决于模型的工具调用能力、接口兼容性、配置和任务。如果出现只输出文字却不调用工具,应结合请求、响应和工具配置排查,不能仅凭模型参数量判断原因。
## 「添加服务商」弹窗逐个字段