The CLI proxy now rejects clients below 1.0.13. Advertise 1.0.46, use the official interactive grok-pager/grok-shell user agent, and send the client identifier and authenticate-response headers.
A single transcript the index could not project flipped the whole index to
degraded, and every session list request then fell back to a full JSONL
scan (about 1.8 GB here, 3.6-5 s of CPU on each cold start).
- Stream records over 8 MiB (Read-tool image results store their base64
twice) into a bounded skeleton instead of rejecting the transcript. The
selection logic is shared with the metadata reader via
boundedJsonProjection.
- Treat budget, changed-during-read and transient I/O failures as
source-scoped: list and sidebar reads keep serving the index and read
only those transcripts from disk. Index-wide failures still fall back.
- Persist budget failures in source_files.state so the next launch knows
them before its first read, and stop rereading them on appends.
- A targeted entry read that misses one transcript no longer cools down
index reads for every other session.
Adds a microphone button beside the composer. Click to record, click to stop; the audio is resampled to 16 kHz mono PCM16 WAV, posted to the local server and transcribed by a SenseVoice worker process, and the text lands in the draft without being sent. If the draft changed or an IME is composing, the result is held behind an insert button instead of overwriting the user's text.
Server: a small provider registry behind /api/voice/* (catalog, preferences, prepare/cancel/status/remove, transcribe). The engine and model are downloaded at runtime to <config>/cc-haha/voice with pinned sha256/sha512, HuggingFace plus hf-mirror and npm plus npmmirror, HTTP Range resume, automatic retry after interruptions, and a partial file kept across cancels. Recognition runs in a separate worker process (sidecar --voice-worker), started on demand and reclaimed when idle.
Desktop: an independent Voice input settings tab with enable switch, model download progress and resume, language, microphone selection and a transcription test with a live waveform. Uses the shared Dropdown/Card/Button components; adds a danger-ghost Button variant. Preferences live in desktop-ui.json (schemaVersion 6); the microphone device id stays in localStorage.
Electron: main-window media permission handler limited to app pages, main frame and audio only, plus the audio-input entitlement and NSMicrophoneUsageDescription.
Scope: Electron desktop only. Not verified on Windows or Linux, with a real microphone, or in a signed and notarized package.
* feat(desktop): show per-format file icons on chat file cards
Generated-file cards, the turn change card and message attachments used
one grey Material glyph for every file. They now render a folded-corner
document icon with a colored body and the extension label (PDF red,
Word blue, Markdown blue, Excel green, PowerPoint orange, archives
amber, code purple, ...), built on react-file-icon instead of
hand-drawn art. Extensions the library does not know are mapped to a
close sibling or to a category color, and jsonl now counts as code.
Brand colors live in lib/fileTypePalette.ts because react-file-icon
writes them into SVG attributes, where CSS variables are not reliable.
* fix(desktop): open workspace documents written through a symlinked path
Clicking an output card for /tmp/app/report.pdf opened the system app
instead of the workspace preview when the session's canonical workdir is
/private/tmp/app. The gate compared path strings, so a symlinked form,
a workdir that had not loaded yet, or a registered access root all
looked like "outside the workspace".
The string check stays as the fast path. When it says "outside", ask the
server through getWorkspaceFile, which resolves real paths: an accepted
document opens in the workspace, a 403 still goes to the system app.
* fix(desktop): load local images in the web UI and H5
A bare <img src> cannot send Authorization, and the server refuses a
credential-less cross-site subresource load, so the browser blocked the
response (net::ERR_BLOCKED_BY_ORB) and chat showed "unable to load
image". Only the Electron shell worked, because its main process injects
the credential for an allowlist of media routes.
When an <img> fails, retry once through the credentialed client and show
the result as a blob: URL; the failure notice appears only if that also
fails. Covers the inline image gallery, the lightbox, image generation
slots and Markdown images. The credential is only ever sent to the local
server's own origin.
* fix(desktop): keep spreadsheet numbers whole in the workspace preview
Columns without a stored width fell back to a fixed 72px, and stored
widths were chosen for Excel's font, which is narrower than the
preview's. Values such as 12,000.00 were clipped to "12,000...".
A column the file gives no width now fits its widest cell, and a number
column never ends up narrower than its numbers. Text in a column whose
width the file sets is still clipped, as Excel clips it, and merged
headings that spill over their span do not widen a column.
Documents the agent writes open in the workspace panel instead of another
application, and local images the agent mentions show up in the conversation.
Workspace preview
- PDF (pdf.js with its own layout and text layer), Word (docx-preview inside a
scripts-disabled sandboxed iframe) and Excel (SheetJS; .xlsx, .xlsm, .xls) open
in the side panel with zoom and fit, per-file scroll/zoom/sheet memory, and a
refresh when the agent rewrites the file. The engines load lazily.
- Bytes come from a new GET /api/sessions/:id/workspace/raw route, with an
extension allowlist, size caps, the workspace boundary and canonical-path
checks. The file endpoint returns metadata and a version for documents. The
client fetches with the bearer credential, so it works in Electron, LAN H5 and
remote access alike.
- Chat links, output cards and the change card open pdf/docx/xlsx in the
workspace; documents outside the workdir still go to the system application.
- Image viewer with fit, zoom and pan, and "open in system app".
Chat images
- Markdown images outside the workdir, at ~/, C:\ and file:// paths render, open
in the viewer, and offer "open original" (pictures only).
- Images returned by tools such as Read appear as thumbnails under the call.
Hardening found in review
- previewFsUrl escapes each path segment; a double-escaped %2e%2e used to leave
/preview-fs/<session>/.
- The CORS, API timing and remote-access header decorators set headers in place.
Rebuilding the response buffered whole files in memory and dropped
Content-Length.
- The engine owns the pdf.js worker, so closing one document no longer fails the
next open.
- Office archives are inflated in steps to check their real sizes, not the sizes
they declare.
- A viewer that fails to load stays in its panel instead of taking the window down.
Adds pdfjs-dist, docx-preview, xlsx (SheetJS 0.20.3 tarball) and fflate as
renderer dev dependencies; Vite bundles them.
Refs #1397
The localized copy for a request-too-large rejection blamed the selected model
and told users to delete large files. The limit belongs to the provider or
relay, and the runtime now removes earlier images and documents by itself on
the next message, so say that and point to compacting or a new session if the
request still fails.
Refs #1399
A 413 from the API or a relay was reported as "Request too large (max 20MB)",
which is the PDF-only limit, and the only recovery stripped top-level media
from the single turn before the error. When the bytes sat in tool results,
@-mentioned images or older turns, nothing shrank and every later message
failed the same way.
The error now reports the size of what was sent and how much of it is images or
documents, names the provider or relay as the side that rejected it, and keeps
the upstream's own text in errorDetails. After the error, images and documents
in everything the failed request carried are replaced with placeholders,
including media nested in tool results and from @-mentioned attachments. The
rejection carries no sourceModel, so it still applies after a model switch.
Transcripts saved with the old wording keep working, and compaction shares the
placeholder logic instead of keeping its own copy.
Refs #1399
A UTF-8 BOM (PowerShell 5.x) or a zero-filled file (crash mid-write on NTFS)
made every read of scheduled_tasks.json and scheduled_tasks_log.json throw, so
listing and creating scheduled tasks failed with 500 until the file was fixed
by hand.
Both readers now strip the BOM and treat a blank file as empty. Content that
has data but does not parse still throws and is never overwritten.
Refs #1400
A tool call the server rejected as invalid stays in the transcript, and its
card is rebuilt from that input on every replay. An option description that
came back as an object threw React #31 and replaced the whole app with the
error page, again on every launch because the open tab is restored.
AskUserQuestion now renders only the string fields it can show, and each
transcript row sits in its own error boundary so a bad record no longer takes
the whole window down.
Refs #1400
The "Session running" dialog treats a running background task as the session
being busy, but Stop & Close only called stopGeneration, which interrupts the
foreground turn and Agent tasks. A background shell command (and its CLI) kept
running, so the reopened session spun again and asked the same question on the
next close.
Stop & Close now also stops every running background task, before the socket is
closed. "Is the session running" and "which tasks to stop" share one predicate
in backgroundTasks, so the dialog cannot offer a Stop that leaves the task that
raised it alive. Keep running sends nothing.
Refs #1398
Everything the assistant did between a <task-notification> turn and the next
real user message was hidden, so the reply a model gave once a background
command finished (and any tool work after it) never reached the chat, live or
after a reload, even though it was written to the transcript.
Drop that suppression from every session history projection (full, paged,
recovery, sub-agent lookup) and from the desktop store (history mapping and the
live-stream flag). Only the injected notification prompt stays hidden; the
background task cards still come from its notification data. The suppression
was added for #886 to quiet replies to stale notifications; the CLI already
avoids those at the source when the model has read the task's result.
Refs #1389
Keep explicit file identities across output cards and prose links, preserve
shell outputs without checkpoint evidence, and retain a file card when
video preview fails. Add cross-project path and opening regressions.