mirror of
https://github.com/NanmiCoder/claude-code-haha.git
synced 2026-10-10 11:53:10 +08:00
ffa2b59105
The instruction files told every coding agent to reach for agent-browser whenever a change needed browser-level evidence: copilot-instructions listed "E2E or agent-browser smoke" as the remedy for cross-boundary flows, and both contributing guides repeated it. That wording outlived the tool. With the agent-browser skill uninstalled, agents still parsed those lines as a recommendation and went looking for the binary instead of using the browser skill that is actually installed. Deleting the references would have made the docs wrong. agent-browser is still a real dependency: check:desktop-ui-smoke spawns it on Linux CI, and seven maintainer-run e2e scripts under desktop/scripts drive it directly. It cannot be swapped for ego-browser either — ego lite is a macOS-only GUI app with no headless mode and a one-time interactive onboarding, so it cannot run on ubuntu-latest at all. So the lanes keep the binary and the prose loses the recommendation. agent-browser is now described as an implementation detail of those two call sites, and ad-hoc browser work — manual verification, screenshots, exploratory UI checks — is pointed at the ego-browser skill. The quality contract asserted the old string, so it would have failed closed on the reworded line. It now pins the replacement plus the new routing rule; flipping either sentence turns the test red.
3.0 KiB
3.0 KiB
AI Coding Instructions
Follow the root AGENTS.md and the nearest nested AGENTS.md for the files you edit.
For every feature, bugfix, refactor, or workflow change:
- Treat tool access as capability, not authorization. Do not commit, push, open/merge a PR, release, run live providers, or change repository settings unless explicitly requested.
- Inspect
git status --short, identify the changed surface, define the intended behavior and failure signal, and inspect the nearest implementation and tests before editing. - Identify the changed surface before coding:
desktop,server/runtime,adapter,native,docs,provider/runtime,agent-loop,persistence,policy/ci, orrelease. - Add same-area tests with the production change. Do not leave production behavior untested unless the PR explicitly carries the maintainer override
allow-missing-tests. - Preserve or improve the coverage ratchet. New or changed executable production lines must pass the changed-line coverage threshold in
scripts/quality-gate/coverage-thresholds.json; do not edit coverage baselines or thresholds without maintainer approval viaallow-coverage-baseline-change. - Use unit tests for pure logic, API/request-shape tests for server/provider/runtime behavior, Testing Library/Vitest for desktop UI and stores, and E2E or desktop UI smoke for user-visible cross-boundary flows.
- Ad-hoc browser automation (manual verification, screenshots, exploratory UI checks) goes through the
ego-browserskill. Theagent-browserbinary is reserved for the committedcheck:desktop-ui-smokelane anddesktop/scripts/e2e-*-agent-browser.sh; do not reach for it as a general browser tool. - Provider/auth/runtime-env/model-window/proxy changes require offline
bun run check:provider-contract; desktop chat/WebSocket/session-runtime changes requirebun run check:chat-contract. - Required PR evidence must be deterministic: use fake credentials, temporary config/home paths, mocked or loopback transports, explicit cleanup, and restored environment state. Never call a real provider or use saved machine credentials in required tests.
- For agent loop, tool execution, provider routing, model selection, file editing, permissions, session resume, and desktop chat changes, include mock/fixture/contract tests. Live smoke is trusted-maintainer evidence only and requires explicit authorization; finding local credentials is not authorization.
- Run the focused regression first, then
bun run check:impactand every selected surface/contract check. Runbun run verifyonly before claiming PR-ready/push-ready or when full validation was requested. - Do not present skipped, blocked, not-run, mock, build-only, or stale evidence as passed live/runtime verification.
- In the final handoff or PR description, include changed files, tests added, commands actually run with pass/fail counts, checks not run, coverage report path when generated, deterministic E2E evidence, live report path or explicit maintainer-only deferral, and known residual risk.