A cold merge kept live thinking and reply rows beside their durable copies because thinking has no identity and live text only gets a transcript id when the cache was unchanged. Coalesce such rows with the durable row in the same gap between shared rows and the same number of user turns away. Also keep whitespace-only thinking deltas inside a streaming block so live thinking matches the transcript.
The default-permission menu in Settings > General grew rightwards from the
trigger's left edge. The trigger sits at the end of the row, so the 340px
menu ran past the content column and was clipped. Add a menuAlign option
and right-align the menu there; composer callers keep the left alignment.
Fixes#1474
* fix(desktop): label the custom provider preset in the active locale
The custom preset showed the catalog's bare English "Custom" in every
language and prefilled it as the config name. It now reads "Custom API"
(自定义模型 API, etc.) on the chip and in the prefilled name.
* fix(desktop): open the composer + submenu level with its row
The side panel was bottom-aligned to the root menu, so Skills and
Connectors both opened at the bottom. It now lines its title up with the
row that opened it and only moves to stay inside the window.
* fix(desktop): show pressed sidebar toggles in the theme accent
A pressed sidebar button used the hover fill, so the task-view bell looked
the same on and hovered. It now takes the brand accent. Pressed buttons
also stop emitting the tone's resting text color, which won over the
pressed color by class order.
* fix(desktop): place the browser edit bubble at the click and capture after it closes
- Picking an element that fills the view (such as <body>) squeezed the
bubble to its 160px minimum at the bottom-left. It now opens at full
height beside the click.
- The <webview> guest does not repaint before a native capture, so the
selection screenshot still showed the closed bubble and no annotation.
Hiding the page chrome now resolves after the next painted frame.
* feat(desktop): order the token usage page usage, activity, then models
The per-model table moves below the activity heatmap and insights under
its own "Usage by model" heading.
* fix(desktop): wrap the Agent Teams header in narrow panes
The header had a fixed 1180px minimum that scrolled sideways and let
flex squeeze button labels into vertical stacks. The controls now wrap
onto their own row, and buttons and timings never wrap.
A session finished by another CLI run (official CLI, or one that has
exited) kept a background shell "running" forever: the sidebar spun, the
header said the session was active, and Stop failed with "CLI session is
not running".
The CLI records a task's end in several shapes depending on when the
notification was delivered, and history restore only read the user-turn
one:
- notifications drained mid-turn are queued_command attachments;
- notifications queued but never delivered survive only as a
queue-operation;
- TaskStop/KillShell leave no notification, only their own tool result.
Read all of them, and when no CLI process exists, converge a Stop of a
history-only task to stopped instead of failing.
Connector rows in the + menu used a generic plug icon, while the extension
market shows each connector's logo from public/connectors/<id>.svg. The menu
now uses the same logo and falls back to the plug when it fails to load.
Market skill icons, which are absolute URLs, are no longer prefixed with
the app's base path.
- Group the root into "for this task" (skills, connectors, files) and
"how it runs" (Agent Teams and Computer Use switches); the rest moves
under More tools.
- Open categories in a side panel on desktop; narrow windows and the
phone sheet keep in-place navigation.
- Connectors list connected services, ones that need sign-in, and three
fixed suggestions, each opening its detail in the extension market.
- Drop the Plugins entry: plugin skills show under Skills, MCP plugins
under Connectors, the rest stay searchable.
- Skills show recently sent ones, all skills, and two curated market
skills that install in place.
- Agent Teams arms the next message with a team instruction and a
removable chip; the bubble and restored history show only the typed text.
- Turning Computer Use on from the menu goes through the same consent
dialog and grants as Settings.
The phone home is now the session list alone. The docked composer took a
quarter of the screen and could not send until a project was picked, so a
New task button floats over the list instead and opens a full-height sheet.
- The sheet puts the project at the top, above the keyboard, and docks the
composer to the bottom. Cancel, the system back gesture or pulling the bar
down closes it; sending swaps it for the new session.
- The project defaults to the list's project filter, otherwise to the
project of the most recently created task.
- On a tablet the button opens the new-session page beside the list.
Replace the drawer-based mobile layout with a single navigation stack. Home
is the session list grouped by status (waiting for you, working, then by
day) with the new-task composer docked under it; a session, Settings and
subagent or team-member pages push on top, and the system back gesture
moves up one level. Touch tablets keep the list beside the page.
- Composer: one row (+, model, send); permission mode, context usage and
the capability menu move into a + bottom sheet.
- Approvals: a waiting permission, question or plan takes the composer's
place; the transcript keeps a one-line marker. Denying can carry a reason.
- Messages: press and hold for copy, select text, edit and branch instead
of an always-visible icon row.
- Tool runs fold into one card per run and open the full timeline in a
sheet; the turn change card opens a file's diff full screen.
- Activity: a pill in the session bar opens the activity sheet instead of
auto-opening over the chat.
- Settings: one grouped list with the connection, provider, phone-editable
preferences and the computer-only settings shown read-only.
- Server: GET /api/sessions/live-status reports every running or waiting
session so the list can sort sessions the phone never opened.
* fix(server): reap background shell tasks on Stop and after runtime exit
A session's background shell task (Bash `run_in_background`, Dream, workflow)
is tracked as a non-Agent task, and nothing on the server ever bounded or
reaped one. A user Stop only interrupted the foreground turn and the Agent
tasks, so the shell kept running after the user had asked everything to stop.
A client disconnect let the CLI live forever, because the disconnect watcher
waits for `hasActiveBackgroundTasks` to clear and a task that never emits a
terminal notification never clears it. A hard runtime death published no
terminal state at all, so a later reconnect re-hydrated a ghost "Running"
entry that could never reach a terminal state.
This closes all three from one file of server bookkeeping:
- `handleStopGeneration` now reaps the session's non-Agent background tasks
through the same `requestControl` stop path the panel's per-task Stop uses,
and the 3-second force-kill guard watches them too. Programmatic stops
(turn stop, runtime-config restart) keep the previous narrow behavior.
- A new 31-minute ceiling bounds a session that a disconnected client keeps
alive through background tasks alone; when it elapses the shared runtime is
stopped and terminal bookends are published. A reconnect clears the ceiling.
- `close()` and `closeStoppedAgentsAfterRuntimeExit` publish terminal bookends
whenever the runtime is already gone but task records remain, instead of
leaving them to be re-hydrated as "Running".
The new timer registry is released by `cleanupSessionRuntimeState` and
classified as `cleared` in the session-state cleanup guard, so a ceiling
scheduled for a session cannot outlive that session.
会话的后台 shell 任务(Bash `run_in_background`、Dream、workflow)登记为非 Agent 任务,
服务端一直没有收掉它们的路径:按停止键只断前台轮和 Agent 任务,shell 还在跑;客户端断开后
CLI 因为 `hasActiveBackgroundTasks` 恒为真而一直存活;运行时崩溃又不写终态,重连后又
变成幽灵 Running。这次在 handler 一处收掉这三种:
- 按停止键时一并停掉本会话的非 Agent 后台任务,走面板单任务 Stop 同一个 stop 通路,
三秒强杀闸也一并盯住它们;程序化停止(停轮、重启换配置)仍走原来的窄路径。
- 新增 31 分钟硬上限:没有客户端、只靠后台任务续命的会话,到期就停掉共享运行时并写终态;
重连则撤销该计时。
- 运行时已经没了但任务记录还在时,`close()` 与 `closeStoppedAgentsAfterRuntimeExit`
直接写终态,不再留着让重连把 Running 灌回来。
Refs #1132
* fix(runtime): converge background tasks and reap orphaned shells
* fix(desktop): restore stopped shell tasks after cold history load
---------
Co-authored-by: gugugaga <267102352+omazili-guga@users.noreply.github.com>
Co-authored-by: 程序员阿江(Relakkes) <relakkes@gmail.com>
Recording:
- While dictating, the composer's toolbar row becomes a recording bar:
cancel, a scrolling trace of what the microphone heard, the clock
(a countdown in the last 15 s), stop, and send. Recording is no longer
painted in the error color.
- Stop writes the text into the draft. Send writes it, then submits the
draft through the composer's own path once the render carries the text.
Text held back because the draft changed is never sent.
- Dictated text is tinted for a moment where it landed.
- The settings transcription test uses the same trace, clock and stop
disc. VoiceWave is removed.
Default:
- Voice input is on by default for preferences that never saved it; a
saved `false` is kept.
- In the desktop app the microphone shows before the model is downloaded
and opens Settings > Voice input. The browser (H5) never offers it.
The edit-and-resend action disappeared whenever an edit could not start:
while a turn ran, while any background task (a dev server, a background
agent) was still running, while turn checkpoints loaded, or when they
failed. With nothing on screen it read as the feature being gone.
Completed prompts now keep the action, disabled, with the reason on hover
and as its accessible description. An open editor and its draft survive a
background task starting. Sessions past the checkpoint preview budget can
edit again by resolving each prompt through its transcript id; since that
mode already tells the user file undo is unavailable, those edits roll
back the conversation only.
A side chat forks the parent transcript, so the server rejects it for a
blank session. The new-session page still showed the side chat button in
its top-right corner, and every click only produced a "Start a
conversation before opening a side chat" toast.
Gate every way in on the session having a conversation: the top-right
button now only renders on mobile once the session has messages, and the
workspace launcher, its + menu and the /btw slash menu entry leave side
chat out of a blank session.
Rebuild every desktop page on one visual system, 「素 Porcelain」: a quiet
light theme with a tool-call timeline, bubbled user messages and a floating
composer card.
- Tokens: rewrite the white theme on Porcelain values and give all six
themes one radius scale (4/6/8/12/16/20), one type scale
(11/12/13/14/15/18/22/26), Inter up to semibold, JetBrains Mono at 400,
and status colours by meaning (running = info, waiting = warning,
done = success, failed = error). Terracotta is kept for the mark, the
send key and selection.
- Icons: draw every icon from lucide and drop the Material Symbols font.
The generic sparkle glyph is gone: a working turn shows the running
ring, skills use the same box the @ and / menus use for SKILL.md.
- Chat: tool calls read as a timeline, user turns as bubbles, the composer
as a floating card; cards, menus, dialogs and permission prompts share
the same surfaces and shadows.
- Shell: session tabs and workspace tabs use one shared tab chip; the
sidebar, workspace and settings follow the same grounds.
- Settings: the page frame is fluid up to 1120px for every pane, so
explorer-style panes (memory, skills) are no longer squeezed into 700px.
- Guard: designSystem.test.ts fails on off-scale font sizes, weights above
semibold, italic text, Material Symbols, Tailwind radius names and
sparkle icon imports.
* fix(desktop): show no error for images that are not there
A reply named the screenshots it had taken as
`/tmp/cc-haha-ui-review/0*.png`, and a red "Unable to load image" tile
with a retry button appeared under it. The inline gallery took the glob
for a file: its absolute-path pattern accepted any character but
whitespace and quotes. Wildcards and substitutions (`*`, `?`, `{name}`,
`${id}`, `$NAME`, `%03d`) now end a path there. Brackets stay, since real
directories use them.
The tile was the larger problem. In the local session history we
checked, 45 of the 115 pictures the gallery tried to show were red tiles
and only 18 existed: files cleaned out of /tmp, outputs deleted since,
web routes, example paths. A retry fixes none of them. Every failure
went red because an <img> error carries no status, but the
authenticated fetch that follows it does. A 400, 403, 404 or 413 from
the file routes now means the picture is not there to show: the inline
tile disappears, and a Markdown image falls back to its alt text. A
server fault, a dropped connection, a refused credential or bytes that
do not decode still show the retryable error, whose hint now names
those causes.
The guessed-name versus spelled-out-path split from #1429 goes: why the
load failed decides, not how the path was written. A server test pins
the 404 for a file missing from an allowed root, which the rule relies
on.
* fix(desktop): translate text typed into components
The header over the pictures in a reply read "1 IMAGE" in every locale.
The text was typed into JSX, and the type of the locale files only
proves that every key has a translation, not that components use one.
A scan of the 246 components found more: Mermaid's error, loading text
and preview hint, "Untitled" in the sidebar, Computer Use's "Failed to
check status.", token counts in /context, labels in the trajectory
view, and screen-reader labels such as the sidebar region and the diff
grid. They now go through t(), reusing a key where one existed and
adding 15 to all five locales. The English text is unchanged.
untranslatedText.test.ts keeps it that way. It reads every component
and fails on JSX text, text-bearing attributes and string literals
rendered as children that skip t(). Key caps, code, icon ligatures and
t() fallbacks are not prose. Product names, units and example values
are listed per file with the reason, and a listed text the source no
longer has fails too. Modal's "Close dialog" stays a known gap: a ui
primitive cannot read the locale, so the label has to come from its
callers.
When a provider closes a 200 stream before message_stop and every stream
retry fails, the error now names the upstream model provider as the cause,
states it is not a context-limit error, and records what the attempt
received: event count, last event, stop_reason, whether message_stop was
missing, open blocks, elapsed time and retries spent. That separates a
reply cut mid-block from one that only dropped message_stop.
The error carries a new upstream_stream_interrupted business code. The
desktop shows a localized explanation in all five locales and keeps the
raw evidence line under it.
* fix(agent-teams): keep a stopped team stopped and wake members after any failure
Stop killed every member, but the lead kept polling its mailbox: reports
sent before the Stop started hidden lead turns, and a single SendMessage
from one of them unpaused the whole team. Stop now sends team_plan_pause
so the lead holds teammate mail until the user's next message, and the
server only lets lead instructions written after that message resume the
team. A restart in flight when Stop arrives is stopped as well, and the
idle Stop button hides once every member is stopped.
Members that stopped on a usage limit, billing, network or crash failure
were only resumed if the lead happened to message them. The user's next
message now lists every stopped member with its failure and open tasks,
once per failure, whether or not Stop was pressed.
* feat(desktop): show where each team member stands on its tasks
A member that stopped or failed left its task reading "In progress" with
an animated bar, and its card said only "Stopped" or "Error". Tasks whose
owner stopped, failed or waits to retry now show that state without
animation, and the card says which task the member stopped on.
The member card chips tell done, current and next tasks apart, with the
subject on hover. The member drawer groups tasks into now, up next and
done, names the unfinished dependencies of blocked work, shows done as
3/5, and no longer prints +0:00 for a duration polling could not measure.
A reply that summarised commits quoted `开题报告2.docx` and
`D:/资料/测试文档1.docx` from their messages. The turn only ran `git log`,
yet three Word cards appeared under it and every name was a link, all
opening "file not found". Being mentioned is not being produced, and a
guessed path is not a file.
Output cards now need the turn to have produced the file. A changed file
in the turn checkpoint still counts. A turn whose checkpoint lists no
unverified write source (only read-only shell, no MCP writes) has no
unseen output, so other mentions are dropped without touching the disk.
Otherwise a mention is shown only once the disk reports it exists and
was modified since the turn began, compared on the server's clock via a
new checkpoint `startedAt`. Cards wait while the checkpoint loads.
File links guessed from code spans and prose render as plain code until
the workspace confirms the file exists, and stay plain when it does not.
Links the reply wrote in Markdown are kept as written. Clicking a file
that is gone shows the existing "could not open" toast instead of a tab.
Both checks share one stat channel: GET /api/sessions/:id/workspace/stat
(at most 20 paths, same boundary as the file route, the workspace root
resolved once per request). The client batches every message that mounts
in a tick into one request, dedupes in-flight paths and caches answers
for 10 s, so a remounted message renders its final state at once and a
link just verified opens without another round trip.
Sessions get a Chat / Trajectory switch in the header. The trajectory is a
dense one-line-per-event ledger (system prompt, user, injected context such
as skills and system reminders, assistant responses, tool calls) with a
three-lane minimap, turn folding, search, and a detail panel per row.
- CLI: record a deduplicated prompt snapshot sidecar for desktop sessions
(system prompt, tool catalog, user context) without touching the transcript.
- Server: project transcript records into trajectory rows with bounded,
cursor-paged reads, live appends, a turn index, row detail by byte range,
snapshot blobs, and a time-window lookup into the raw trace capture.
- Desktop: windowed ledger, minimap, detail panel with a raw request summary,
chat <-> trajectory navigation, subagent drill-in.
- Remove the Settings trace list, trace tabs and the standalone trace
window; migrate persisted trace tabs to session tabs.
The global effort default moves from max to low on both the server
(/api/effort fallback) and the desktop store. Users who already saved an
effort keep it.
ChatGPT Official and Grok Official are not in the saved provider list, so
their first draft selection carried no effort: the server ran the model
default while the selector showed the global value. Resolve the effort
against the model catalog for those providers too.
The workspace browser was a native WebContentsView, which always paints
above the DOM. Switching to the skills or settings page left the page
floating over it, and menus over the page needed a screenshot swap that
showed up clipped and flashed white.
Pages are now <webview> guests kept in a layer inside the session panel
that is never unmounted, so other pages, menus and dialogs draw over
them like any element. The main process adopts each guest and keeps all
in-page behaviour: navigation, history, find, zoom, capture, PDF,
downloads, shortcuts and annotation.
- Guard every attach in the main window: browser partition only, start
at about:blank, preload and sandbox pinned, reported ids validated.
- Match the native page: drop the blank history entry, keep page zoom
independent of app zoom, allow popups so they still become tabs.
- Keep app drags working over a page and close menus on a click into it.
- Tell the side dock it is off screen when the session page is hidden.
- Remove the overlay snapshot machinery and native bounds syncing.
The catalog redesign put the title, disclaimer, categories and search on
a white band ruled off from a tinted canvas holding the cards. The
extensions frame, the plugins tab and the skill detail all sit on the
plain page surface, so the seam read as two pages glued together and the
canvas flashed on every home/detail switch.
Lay the controls and the grid in one column on --color-surface and
separate them by spacing; the cards keep their own border and shadow.
* feat(desktop): edit a sent prompt and rerun from there (#1343)
Hovering a prompt the rewind API can already target now offers "Edit and
resend". The bubble turns into an inline editor; sending dry-runs the
existing rewind, confirms when later turns or restorable files are at
stake, rewinds with the same conversation/both modes as undo, reloads
history and sends the edited prompt. A failed rewind changes nothing and
keeps the draft; if the edit cannot be sent after a successful rewind it
is handed back to the composer.
Undo and edit share one rewind routine, the unused per-message
rewindAction prop is removed, and TextArea forwards its ref.
* fix(desktop): preserve edit-resend session and context
The prose scan excludes CJK so a verb is not swallowed into a path, which
cut an unquoted `测试文档1.docx` to `1.docx` and `D:/资料/测试文档1.docx` to
`1.docx`. A match that is recognisably the tail of a CJK name is now
widened to the whole token.
Names the text cannot bound are settled against what exists: `报告v2.docx`,
names glued together without a space, names with spaces or full-width
brackets, and names made only of CJK plus an extension. The extractor
offers the other readings; a changed file settles them, otherwise one
cached workspace listing per folder does, longest existing name first.
CJK-only names read like prose about formats (`后缀为.docx的文件`), so they
are never linked and appear as cards or images only once confirmed.
Fixes#1423
Output cards scanned inline code spans as prose, which excludes CJK on
purpose, so `开题报告2.docx` became a `2.docx` card that opened a missing
file. Code spans are now parsed whole with the same Unicode-aware parser
the rendered chip uses.
An image the reply only named by a bare filename, with no changed file
corroborating its location, was resolved at the workdir root and showed a
red load error when it was not there. Such guesses now disappear on
failure; paths spelled out with a directory keep the error and retry.
Long Agent Teams runs lost members for good: a truncated provider stream
ended a member's turn with nobody to wake it, the desktop Stop button and
every lead restart killed all members and marked the plan interrupted,
mail sent to a stopped member landed in an inbox nothing read, and a lead
kept inside one long turn never saw member reports. Aligned with the
official CLI 2.1.284 and verified with DeepSeek Flash through a
fault-injecting proxy.
Stream recovery
- Re-send a stream that breaks before any tool ran (proxy truncation,
transport errors), with the existing retry budget and backoff; the
desktop drops the discarded attempt's tool cards and todo update.
Desktop team runtime (teamPlanRuntime)
- The server supervises members: a stopped member restarts from its own
transcript when messaged; transient failures continue automatically
(15s/45s/2m/5m/10m) and only exhausted retries reach the lead; ready
dependent tasks wake their owner; a crash-loop guard ignores user stops.
- Stop pauses the team instead of ending it; the lead's next user message
is followed by a notice listing the stopped members and their open
tasks. Lead restarts (model/permission switch, crash) keep members;
server restarts re-own the team. Teams end on /clear or session delete.
- Approving a plan no longer races a concurrent plan read into
"Launch ownership was lost".
Mailbox and messaging
- Atomic inbox writes, identity-based read marking, read history files,
idle notifications with result/failureReason, and write failures
reported instead of "Message sent". External builds keep the official
between-turn delivery to the lead.
- SendMessage resumes non-running in-process teammates from their
transcript, notes restarting desktop members, queues mail for members
of a plan awaiting approval, and rejects unknown names.
CLI in-process teammates
- Compaction uses the teammate's own controller and real history and no
longer kills it on error; failed turns are classified and continued;
the turn-end mailbox drains as one batch; one durable transcript per
teammate.
Lead behaviour
- An unmet /goal ends the lead turn while members work, so member reports
arrive; WaitSessions on own team members returns immediately.
Desktop UI
- Member states for stopped, auto-retrying and failed, with reason,
countdown and recovery hint in all five locales.
Tests and tooling
- Regression tests for every behaviour above; module mocks in four test
files are restored after use so the single-process coverage run is not
polluted; the desktop smoke asserts the new Stop semantics.
Open the skills market on a bundled catalog of 398 curated ClawHub and
SkillHub skills in 13 categories instead of querying both registries live.
Live search stays available as an explicit "search all markets" scope
whose results are marked as not curated.
- Server: catalog scope (default) with category filter, filtering before
pagination and batched install state; catalog metadata overlaid on live
detail; per-scanner ClawHub reports, changelog and page URL; manual
catalog refresh script.
- Pin ClawHub reads and installs to the card's owner so a same-slug copy
by another author is never shown or installed in its place.
- Desktop: category chips, curated cards with tags, locale-aware summaries,
detail page with stats, security report, changelog, capability panel and
triggers; install confirmation requires acknowledgement for unaudited or
flagged skills.
- Docs: rewrite the skills market section and refresh screenshots.
The downloader races Hugging Face against hf-mirror (and npmjs against
npmmirror) and keeps whichever answers first, so behind a rule-based
proxy the domestic mirror usually wins even though the configured
network proxy is in use. Add an auto/official/mirror preference that
pins one host without silent fallback, and show which network proxy
downloads go through with a link to General settings.
The fit button used the same Maximize2 icon as the workspace panel's
maximize control and stayed pressed while already fitted, where clicking
it did nothing. Give it its own icon (fit width: MoveHorizontal, fit
window: Scan) and disable it while the viewer is fitted, so it is only
clickable when there is a manual zoom to undo.
The @ mention list only loaded disk and plugin skills, so bundled skills
such as imagegen never appeared. imagegen is also gated on image provider
env that is injected into each session's CLI rather than the server, so
the mentions endpoint now accepts the session's providerId and evaluates
imagegen against that provider's runtime env.
Adds a microphone button beside the composer. Click to record, click to stop; the audio is resampled to 16 kHz mono PCM16 WAV, posted to the local server and transcribed by a SenseVoice worker process, and the text lands in the draft without being sent. If the draft changed or an IME is composing, the result is held behind an insert button instead of overwriting the user's text.
Server: a small provider registry behind /api/voice/* (catalog, preferences, prepare/cancel/status/remove, transcribe). The engine and model are downloaded at runtime to <config>/cc-haha/voice with pinned sha256/sha512, HuggingFace plus hf-mirror and npm plus npmmirror, HTTP Range resume, automatic retry after interruptions, and a partial file kept across cancels. Recognition runs in a separate worker process (sidecar --voice-worker), started on demand and reclaimed when idle.
Desktop: an independent Voice input settings tab with enable switch, model download progress and resume, language, microphone selection and a transcription test with a live waveform. Uses the shared Dropdown/Card/Button components; adds a danger-ghost Button variant. Preferences live in desktop-ui.json (schemaVersion 6); the microphone device id stays in localStorage.
Electron: main-window media permission handler limited to app pages, main frame and audio only, plus the audio-input entitlement and NSMicrophoneUsageDescription.
Scope: Electron desktop only. Not verified on Windows or Linux, with a real microphone, or in a signed and notarized package.
* feat(desktop): show per-format file icons on chat file cards
Generated-file cards, the turn change card and message attachments used
one grey Material glyph for every file. They now render a folded-corner
document icon with a colored body and the extension label (PDF red,
Word blue, Markdown blue, Excel green, PowerPoint orange, archives
amber, code purple, ...), built on react-file-icon instead of
hand-drawn art. Extensions the library does not know are mapped to a
close sibling or to a category color, and jsonl now counts as code.
Brand colors live in lib/fileTypePalette.ts because react-file-icon
writes them into SVG attributes, where CSS variables are not reliable.
* fix(desktop): open workspace documents written through a symlinked path
Clicking an output card for /tmp/app/report.pdf opened the system app
instead of the workspace preview when the session's canonical workdir is
/private/tmp/app. The gate compared path strings, so a symlinked form,
a workdir that had not loaded yet, or a registered access root all
looked like "outside the workspace".
The string check stays as the fast path. When it says "outside", ask the
server through getWorkspaceFile, which resolves real paths: an accepted
document opens in the workspace, a 403 still goes to the system app.
* fix(desktop): load local images in the web UI and H5
A bare <img src> cannot send Authorization, and the server refuses a
credential-less cross-site subresource load, so the browser blocked the
response (net::ERR_BLOCKED_BY_ORB) and chat showed "unable to load
image". Only the Electron shell worked, because its main process injects
the credential for an allowlist of media routes.
When an <img> fails, retry once through the credentialed client and show
the result as a blob: URL; the failure notice appears only if that also
fails. Covers the inline image gallery, the lightbox, image generation
slots and Markdown images. The credential is only ever sent to the local
server's own origin.
* fix(desktop): keep spreadsheet numbers whole in the workspace preview
Columns without a stored width fell back to a fixed 72px, and stored
widths were chosen for Excel's font, which is narrower than the
preview's. Values such as 12,000.00 were clipped to "12,000...".
A column the file gives no width now fits its widest cell, and a number
column never ends up narrower than its numbers. Text in a column whose
width the file sets is still clipped, as Excel clips it, and merged
headings that spill over their span do not widen a column.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397