The existing rule — production changes need a same-area regression test — is
satisfied by exactly the tests that failed to hold. ContextUsageIndicator.tsx was
fixed three times in ninety minutes; all three fixes shipped tests, and all three
tests covered only their own hop. 21 of the last 70 fix commits edit lines
another fix wrote within 30 days.
Coverage was never the missing signal. That file sits at 87% branch coverage.
What the tests had in common was shape, so this writes the shape down, with the
commit behind each rule:
- drive the transition instead of assigning the state it produces (component
tests call setState 744 times and a real store action 3 times, and assigned
state cannot expose "A did not update B")
- assert the invariant, not today's output (2262973a4 asserted the bug and the
next fix inverted that exact line)
- cover both directions of any drop/merge rule (the replay guard was only ever
tested for "discard the replay")
- test the join, not each end (deleting ChatInput's refreshNonce term left 314
tests green)
- never retune an existing test's inputs to keep it green (128f75ab5 changed five
tests' props rather than accept they described unreachable states)
- do not mock the module under test
- comparing content to decide identity means the identity exists upstream
Plus the two blind spots worth checking rather than trusting: desktop/electron is
not instrumented at all, and Bun's LCOV emits no branch records, so src/ and
adapters/ report 100% branch coverage for data that was never collected.
The desktop file gains the module-placement rules the Settings split and the
reachability guard produced.
`components/AGENTS.md` is the reference: what to use for a given need,
where a new component goes, and the style / i18n / a11y / test rules.
`desktop/AGENTS.md` now routes here — the line it replaces ("reuse the
existing desktop design system") named nothing to look up and so was not
an executable instruction for a person or a model.
`docs/component-library-plan.md` keeps the audit evidence and a record of
what actually shipped, including where the plan was wrong.
Two rules earned their own sections because the library broke them
itself:
- Overriding a component's utility with `className` does not reliably
win. Tailwind sorts same-utility arbitrary values by value and takes
the last, regardless of the order they were passed. `hoverTone="danger"`
was a no-op with exactly the two tones it was built for, and its test
passed because it only asserted the red class was *present*, never that
the neutral one was gone. Assert that a class prefix appears exactly
once, and prove it by reverting the fix.
- The library must not compose user-visible English. `SearchField` built
`Clear ${label}`, which made adopting it an i18n regression for every
caller that had already translated its clear button. The first repair —
falling back to `label` — was worse: the input and its clear button
then shared an accessible name and `getByLabelText` matched both.
Reuse went from 22% to 75% (524 component uses against 179 remaining
native buttons). The remainder are elements that should not be
components: `role="tab"`, `menuitem`, `option`, `treeitem`, `gridcell`,
whole-row and whole-card click targets, drag handles, and the OS
titlebar.
Route required checks by changed surface, add offline provider and chat contracts, and keep fork PRs independent of live credentials. Layer agent guidance by subtree and enforce a compact instruction budget.
Tested: bun run check:policy
Confidence: high
Scope-risk: broad