A session finished by another CLI run (official CLI, or one that has
exited) kept a background shell "running" forever: the sidebar spun, the
header said the session was active, and Stop failed with "CLI session is
not running".
The CLI records a task's end in several shapes depending on when the
notification was delivered, and history restore only read the user-turn
one:
- notifications drained mid-turn are queued_command attachments;
- notifications queued but never delivered survive only as a
queue-operation;
- TaskStop/KillShell leave no notification, only their own tool result.
Read all of them, and when no CLI process exists, converge a Stop of a
history-only task to stopped instead of failing.
Replace the drawer-based mobile layout with a single navigation stack. Home
is the session list grouped by status (waiting for you, working, then by
day) with the new-task composer docked under it; a session, Settings and
subagent or team-member pages push on top, and the system back gesture
moves up one level. Touch tablets keep the list beside the page.
- Composer: one row (+, model, send); permission mode, context usage and
the capability menu move into a + bottom sheet.
- Approvals: a waiting permission, question or plan takes the composer's
place; the transcript keeps a one-line marker. Denying can carry a reason.
- Messages: press and hold for copy, select text, edit and branch instead
of an always-visible icon row.
- Tool runs fold into one card per run and open the full timeline in a
sheet; the turn change card opens a file's diff full screen.
- Activity: a pill in the session bar opens the activity sheet instead of
auto-opening over the chat.
- Settings: one grouped list with the connection, provider, phone-editable
preferences and the computer-only settings shown read-only.
- Server: GET /api/sessions/live-status reports every running or waiting
session so the list can sort sessions the phone never opened.
An approved team handed every member its instructions at the same moment,
so members whose tasks depended on unfinished work started anyway: the
second and third layers of a plan ran before the first, a member could mark
a blocked task in progress, and results only ever went to the lead. The
official CLI states dependencies in tool prompts and nothing more, so the
plan the user approved was not what ran.
- A member whose tasks all wait on other tasks gets its instructions only
once one of them is ready. What the lead or a teammate sends it before
that waits in its inbox and arrives with the instructions; only the user
writing to it, or a shutdown request, reaches it earlier. This survives a
Stop and a server restart.
- TaskUpdate refuses a teammate that starts or completes a task while a
task it is blocked by is unfinished. The lead is not held to it.
- Each member is told who waits on its tasks and to send them its result
before completing, so a dependent member starts with that result in
hand. A member that starts without one is told whose is missing.
- A member that exits after approving a shutdown request is no longer
recorded as failed, and the lead gets no failure notice for it. The
shutdown-approval schema keeps leaving out the desktop backend on
purpose, now documented and tested.
- TeamPlan tells the lead that a dependent member's prompt is delivered
when its dependencies are done.
* fix(server): reap background shell tasks on Stop and after runtime exit
A session's background shell task (Bash `run_in_background`, Dream, workflow)
is tracked as a non-Agent task, and nothing on the server ever bounded or
reaped one. A user Stop only interrupted the foreground turn and the Agent
tasks, so the shell kept running after the user had asked everything to stop.
A client disconnect let the CLI live forever, because the disconnect watcher
waits for `hasActiveBackgroundTasks` to clear and a task that never emits a
terminal notification never clears it. A hard runtime death published no
terminal state at all, so a later reconnect re-hydrated a ghost "Running"
entry that could never reach a terminal state.
This closes all three from one file of server bookkeeping:
- `handleStopGeneration` now reaps the session's non-Agent background tasks
through the same `requestControl` stop path the panel's per-task Stop uses,
and the 3-second force-kill guard watches them too. Programmatic stops
(turn stop, runtime-config restart) keep the previous narrow behavior.
- A new 31-minute ceiling bounds a session that a disconnected client keeps
alive through background tasks alone; when it elapses the shared runtime is
stopped and terminal bookends are published. A reconnect clears the ceiling.
- `close()` and `closeStoppedAgentsAfterRuntimeExit` publish terminal bookends
whenever the runtime is already gone but task records remain, instead of
leaving them to be re-hydrated as "Running".
The new timer registry is released by `cleanupSessionRuntimeState` and
classified as `cleared` in the session-state cleanup guard, so a ceiling
scheduled for a session cannot outlive that session.
会话的后台 shell 任务(Bash `run_in_background`、Dream、workflow)登记为非 Agent 任务,
服务端一直没有收掉它们的路径:按停止键只断前台轮和 Agent 任务,shell 还在跑;客户端断开后
CLI 因为 `hasActiveBackgroundTasks` 恒为真而一直存活;运行时崩溃又不写终态,重连后又
变成幽灵 Running。这次在 handler 一处收掉这三种:
- 按停止键时一并停掉本会话的非 Agent 后台任务,走面板单任务 Stop 同一个 stop 通路,
三秒强杀闸也一并盯住它们;程序化停止(停轮、重启换配置)仍走原来的窄路径。
- 新增 31 分钟硬上限:没有客户端、只靠后台任务续命的会话,到期就停掉共享运行时并写终态;
重连则撤销该计时。
- 运行时已经没了但任务记录还在时,`close()` 与 `closeStoppedAgentsAfterRuntimeExit`
直接写终态,不再留着让重连把 Running 灌回来。
Refs #1132
* fix(runtime): converge background tasks and reap orphaned shells
* fix(desktop): restore stopped shell tasks after cold history load
---------
Co-authored-by: gugugaga <267102352+omazili-guga@users.noreply.github.com>
Co-authored-by: 程序员阿江(Relakkes) <relakkes@gmail.com>
The OpenAI Chat proxy decided image support from host and model names:
opencode.ai and DeepSeek hosts replaced every image with a text-only
placeholder unless the model id contained "vision" or was allowlisted.
Each new vision model on OpenCode Go (Kimi K3, Space Bunny, DeepSeek Flash)
therefore told users the endpoint could not read images.
Offer images to every model instead. When the upstream refuses a request
that carried images with a recognizable image-input rejection, resend it
once without images and, only if that succeeds, remember the endpoint and
model as text-only for 30 minutes and record an info diagnostic. Failures
about a particular image (format, MIME type, size, animation, image count)
surface unchanged so one bad image cannot disable images for the model.
The rejection matcher moves to a shared module used by the CLI error mapper
and the proxy, and now also recognizes vLLM and OpenAI wording.
Recording:
- While dictating, the composer's toolbar row becomes a recording bar:
cancel, a scrolling trace of what the microphone heard, the clock
(a countdown in the last 15 s), stop, and send. Recording is no longer
painted in the error color.
- Stop writes the text into the draft. Send writes it, then submits the
draft through the composer's own path once the render carries the text.
Text held back because the draft changed is never sent.
- Dictated text is tinted for a moment where it landed.
- The settings transcription test uses the same trace, clock and stop
disc. VoiceWave is removed.
Default:
- Voice input is on by default for preferences that never saved it; a
saved `false` is kept.
- In the desktop app the microphone shows before the model is downloaded
and opens Settings > Voice input. The browser (H5) never offers it.
* fix(desktop): show no error for images that are not there
A reply named the screenshots it had taken as
`/tmp/cc-haha-ui-review/0*.png`, and a red "Unable to load image" tile
with a retry button appeared under it. The inline gallery took the glob
for a file: its absolute-path pattern accepted any character but
whitespace and quotes. Wildcards and substitutions (`*`, `?`, `{name}`,
`${id}`, `$NAME`, `%03d`) now end a path there. Brackets stay, since real
directories use them.
The tile was the larger problem. In the local session history we
checked, 45 of the 115 pictures the gallery tried to show were red tiles
and only 18 existed: files cleaned out of /tmp, outputs deleted since,
web routes, example paths. A retry fixes none of them. Every failure
went red because an <img> error carries no status, but the
authenticated fetch that follows it does. A 400, 403, 404 or 413 from
the file routes now means the picture is not there to show: the inline
tile disappears, and a Markdown image falls back to its alt text. A
server fault, a dropped connection, a refused credential or bytes that
do not decode still show the retryable error, whose hint now names
those causes.
The guessed-name versus spelled-out-path split from #1429 goes: why the
load failed decides, not how the path was written. A server test pins
the 404 for a file missing from an allowed root, which the rule relies
on.
* fix(desktop): translate text typed into components
The header over the pictures in a reply read "1 IMAGE" in every locale.
The text was typed into JSX, and the type of the locale files only
proves that every key has a translation, not that components use one.
A scan of the 246 components found more: Mermaid's error, loading text
and preview hint, "Untitled" in the sidebar, Computer Use's "Failed to
check status.", token counts in /context, labels in the trajectory
view, and screen-reader labels such as the sidebar region and the diff
grid. They now go through t(), reusing a key where one existed and
adding 15 to all five locales. The English text is unchanged.
untranslatedText.test.ts keeps it that way. It reads every component
and fails on JSX text, text-bearing attributes and string literals
rendered as children that skip t(). Key caps, code, icon ligatures and
t() fallbacks are not prose. Product names, units and example values
are listed per file with the reason, and a listed text the source no
longer has fails too. Modal's "Close dialog" stays a known gap: a ui
primitive cannot read the locale, so the label has to come from its
callers.
When a provider closes a 200 stream before message_stop and every stream
retry fails, the error now names the upstream model provider as the cause,
states it is not a context-limit error, and records what the attempt
received: event count, last event, stop_reason, whether message_stop was
missing, open blocks, elapsed time and retries spent. That separates a
reply cut mid-block from one that only dropped message_stop.
The error carries a new upstream_stream_interrupted business code. The
desktop shows a localized explanation in all five locales and keeps the
raw evidence line under it.
* fix(agent-teams): keep a stopped team stopped and wake members after any failure
Stop killed every member, but the lead kept polling its mailbox: reports
sent before the Stop started hidden lead turns, and a single SendMessage
from one of them unpaused the whole team. Stop now sends team_plan_pause
so the lead holds teammate mail until the user's next message, and the
server only lets lead instructions written after that message resume the
team. A restart in flight when Stop arrives is stopped as well, and the
idle Stop button hides once every member is stopped.
Members that stopped on a usage limit, billing, network or crash failure
were only resumed if the lead happened to message them. The user's next
message now lists every stopped member with its failure and open tasks,
once per failure, whether or not Stop was pressed.
* feat(desktop): show where each team member stands on its tasks
A member that stopped or failed left its task reading "In progress" with
an animated bar, and its card said only "Stopped" or "Error". Tasks whose
owner stopped, failed or waits to retry now show that state without
animation, and the card says which task the member stopped on.
The member card chips tell done, current and next tasks apart, with the
subject on hover. The member drawer groups tasks into now, up next and
done, names the unfinished dependencies of blocked work, shows done as
3/5, and no longer prints +0:00 for a duration polling could not measure.
A reply that summarised commits quoted `开题报告2.docx` and
`D:/资料/测试文档1.docx` from their messages. The turn only ran `git log`,
yet three Word cards appeared under it and every name was a link, all
opening "file not found". Being mentioned is not being produced, and a
guessed path is not a file.
Output cards now need the turn to have produced the file. A changed file
in the turn checkpoint still counts. A turn whose checkpoint lists no
unverified write source (only read-only shell, no MCP writes) has no
unseen output, so other mentions are dropped without touching the disk.
Otherwise a mention is shown only once the disk reports it exists and
was modified since the turn began, compared on the server's clock via a
new checkpoint `startedAt`. Cards wait while the checkpoint loads.
File links guessed from code spans and prose render as plain code until
the workspace confirms the file exists, and stay plain when it does not.
Links the reply wrote in Markdown are kept as written. Clicking a file
that is gone shows the existing "could not open" toast instead of a tab.
Both checks share one stat channel: GET /api/sessions/:id/workspace/stat
(at most 20 paths, same boundary as the file route, the workspace root
resolved once per request). The client batches every message that mounts
in a tick into one request, dedupes in-flight paths and caches answers
for 10 s, so a remounted message renders its final state at once and a
link just verified opens without another round trip.
Sessions get a Chat / Trajectory switch in the header. The trajectory is a
dense one-line-per-event ledger (system prompt, user, injected context such
as skills and system reminders, assistant responses, tool calls) with a
three-lane minimap, turn folding, search, and a detail panel per row.
- CLI: record a deduplicated prompt snapshot sidecar for desktop sessions
(system prompt, tool catalog, user context) without touching the transcript.
- Server: project transcript records into trajectory rows with bounded,
cursor-paged reads, live appends, a turn index, row detail by byte range,
snapshot blobs, and a time-window lookup into the raw trace capture.
- Desktop: windowed ledger, minimap, detail panel with a raw request summary,
chat <-> trajectory navigation, subagent drill-in.
- Remove the Settings trace list, trace tabs and the standalone trace
window; migrate persisted trace tabs to session tabs.
The global effort default moves from max to low on both the server
(/api/effort fallback) and the desktop store. Users who already saved an
effort keep it.
ChatGPT Official and Grok Official are not in the saved provider list, so
their first draft selection carried no effort: the server ran the model
default while the selector showed the global value. Resolve the effort
against the model catalog for those providers too.
Every session route resolves its transcript first, and that lookup parsed
the whole file to rank candidates even when only one file matched. The
desktop pages a long session's history request by request, so reopening
it cost one full parse per page and grew quadratically with file size.
Skip the content check when a single file matches, and stop a multi-file
check at the first conversation record. Ranking and API responses are
unchanged.
Output cards scanned inline code spans as prose, which excludes CJK on
purpose, so `开题报告2.docx` became a `2.docx` card that opened a missing
file. Code spans are now parsed whole with the same Unicode-aware parser
the rendered chip uses.
An image the reply only named by a bare filename, with no changed file
corroborating its location, was resolved at the workdir root and showed a
red load error when it was not there. Such guesses now disappear on
failure; paths spelled out with a directory keep the error and retry.
Long Agent Teams runs lost members for good: a truncated provider stream
ended a member's turn with nobody to wake it, the desktop Stop button and
every lead restart killed all members and marked the plan interrupted,
mail sent to a stopped member landed in an inbox nothing read, and a lead
kept inside one long turn never saw member reports. Aligned with the
official CLI 2.1.284 and verified with DeepSeek Flash through a
fault-injecting proxy.
Stream recovery
- Re-send a stream that breaks before any tool ran (proxy truncation,
transport errors), with the existing retry budget and backoff; the
desktop drops the discarded attempt's tool cards and todo update.
Desktop team runtime (teamPlanRuntime)
- The server supervises members: a stopped member restarts from its own
transcript when messaged; transient failures continue automatically
(15s/45s/2m/5m/10m) and only exhausted retries reach the lead; ready
dependent tasks wake their owner; a crash-loop guard ignores user stops.
- Stop pauses the team instead of ending it; the lead's next user message
is followed by a notice listing the stopped members and their open
tasks. Lead restarts (model/permission switch, crash) keep members;
server restarts re-own the team. Teams end on /clear or session delete.
- Approving a plan no longer races a concurrent plan read into
"Launch ownership was lost".
Mailbox and messaging
- Atomic inbox writes, identity-based read marking, read history files,
idle notifications with result/failureReason, and write failures
reported instead of "Message sent". External builds keep the official
between-turn delivery to the lead.
- SendMessage resumes non-running in-process teammates from their
transcript, notes restarting desktop members, queues mail for members
of a plan awaiting approval, and rejects unknown names.
CLI in-process teammates
- Compaction uses the teammate's own controller and real history and no
longer kills it on error; failed turns are classified and continued;
the turn-end mailbox drains as one batch; one durable transcript per
teammate.
Lead behaviour
- An unmet /goal ends the lead turn while members work, so member reports
arrive; WaitSessions on own team members returns immediately.
Desktop UI
- Member states for stopped, auto-retrying and failed, with reason,
countdown and recovery hint in all five locales.
Tests and tooling
- Regression tests for every behaviour above; module mocks in four test
files are restored after use so the single-process coverage run is not
polluted; the desktop smoke asserts the new Stop semantics.
Open the skills market on a bundled catalog of 398 curated ClawHub and
SkillHub skills in 13 categories instead of querying both registries live.
Live search stays available as an explicit "search all markets" scope
whose results are marked as not curated.
- Server: catalog scope (default) with category filter, filtering before
pagination and batched install state; catalog metadata overlaid on live
detail; per-scanner ClawHub reports, changelog and page URL; manual
catalog refresh script.
- Pin ClawHub reads and installs to the card's owner so a same-slug copy
by another author is never shown or installed in its place.
- Desktop: category chips, curated cards with tags, locale-aware summaries,
detail page with stats, security report, changelog, capability panel and
triggers; install confirmation requires acknowledgement for unaudited or
flagged skills.
- Docs: rewrite the skills market section and refresh screenshots.
The downloader races Hugging Face against hf-mirror (and npmjs against
npmmirror) and keeps whichever answers first, so behind a rule-based
proxy the domestic mirror usually wins even though the configured
network proxy is in use. Add an auto/official/mirror preference that
pins one host without silent fallback, and show which network proxy
downloads go through with a link to General settings.
The @ mention list only loaded disk and plugin skills, so bundled skills
such as imagegen never appeared. imagegen is also gated on image provider
env that is injected into each session's CLI rather than the server, so
the mentions endpoint now accepts the session's providerId and evaluates
imagegen against that provider's runtime env.
The CLI proxy now rejects clients below 1.0.13. Advertise 1.0.46, use the official interactive grok-pager/grok-shell user agent, and send the client identifier and authenticate-response headers.
A single transcript the index could not project flipped the whole index to
degraded, and every session list request then fell back to a full JSONL
scan (about 1.8 GB here, 3.6-5 s of CPU on each cold start).
- Stream records over 8 MiB (Read-tool image results store their base64
twice) into a bounded skeleton instead of rejecting the transcript. The
selection logic is shared with the metadata reader via
boundedJsonProjection.
- Treat budget, changed-during-read and transient I/O failures as
source-scoped: list and sidebar reads keep serving the index and read
only those transcripts from disk. Index-wide failures still fall back.
- Persist budget failures in source_files.state so the next launch knows
them before its first read, and stop rereading them on appends.
- A targeted entry read that misses one transcript no longer cools down
index reads for every other session.
Adds a microphone button beside the composer. Click to record, click to stop; the audio is resampled to 16 kHz mono PCM16 WAV, posted to the local server and transcribed by a SenseVoice worker process, and the text lands in the draft without being sent. If the draft changed or an IME is composing, the result is held behind an insert button instead of overwriting the user's text.
Server: a small provider registry behind /api/voice/* (catalog, preferences, prepare/cancel/status/remove, transcribe). The engine and model are downloaded at runtime to <config>/cc-haha/voice with pinned sha256/sha512, HuggingFace plus hf-mirror and npm plus npmmirror, HTTP Range resume, automatic retry after interruptions, and a partial file kept across cancels. Recognition runs in a separate worker process (sidecar --voice-worker), started on demand and reclaimed when idle.
Desktop: an independent Voice input settings tab with enable switch, model download progress and resume, language, microphone selection and a transcription test with a live waveform. Uses the shared Dropdown/Card/Button components; adds a danger-ghost Button variant. Preferences live in desktop-ui.json (schemaVersion 6); the microphone device id stays in localStorage.
Electron: main-window media permission handler limited to app pages, main frame and audio only, plus the audio-input entitlement and NSMicrophoneUsageDescription.
Scope: Electron desktop only. Not verified on Windows or Linux, with a real microphone, or in a signed and notarized package.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397
A 413 from the API or a relay was reported as "Request too large (max 20MB)",
which is the PDF-only limit, and the only recovery stripped top-level media
from the single turn before the error. When the bytes sat in tool results,
@-mentioned images or older turns, nothing shrank and every later message
failed the same way.
The error now reports the size of what was sent and how much of it is images or
documents, names the provider or relay as the side that rejected it, and keeps
the upstream's own text in errorDetails. After the error, images and documents
in everything the failed request carried are replaced with placeholders,
including media nested in tool results and from @-mentioned attachments. The
rejection carries no sourceModel, so it still applies after a model switch.
Transcripts saved with the old wording keep working, and compaction shares the
placeholder logic instead of keeping its own copy.
Refs #1399
A UTF-8 BOM (PowerShell 5.x) or a zero-filled file (crash mid-write on NTFS)
made every read of scheduled_tasks.json and scheduled_tasks_log.json throw, so
listing and creating scheduled tasks failed with 500 until the file was fixed
by hand.
Both readers now strip the BOM and treat a blank file as empty. Content that
has data but does not parse still throws and is never overwritten.
Refs #1400
Everything the assistant did between a <task-notification> turn and the next
real user message was hidden, so the reply a model gave once a background
command finished (and any tool work after it) never reached the chat, live or
after a reload, even though it was written to the transcript.
Drop that suppression from every session history projection (full, paged,
recovery, sub-agent lookup) and from the desktop store (history mapping and the
live-stream flag). Only the injected notification prompt stays hidden; the
background task cards still come from its notification data. The suppression
was added for #886 to quiet replies to stale notifications; the CLI already
avoids those at the source when the model has read the task's result.
Refs #1389