docs: rebuild the documentation site around two readers

The site had drifted from the product. Every screenshot predated the
v0.5.0 UI redesign, the reading experience shipped no search and no
syntax highlighting, and a third of the pages were internal process
artefacts — migration task lists addressed to agentic workers, a
release runbook, a proposal marked "historical".

Reorganise around the only two people who read this: someone getting
the desktop app running for the first time, and someone reading the
source. Five sections replace nine — start / desktop / im / cli /
internals — and the pages that served neither reader are gone.

Site rewrite:

- Palette lifted from the desktop app's 「纸·墨·印」 themes, so the
  site and the product read as one thing. Light mirrors 纯白, dark
  mirrors 墨夜, and dark mode exists at all now.
- Fonts are self-hosted. The old @import from Google Fonts is
  unreachable from mainland China, which left every heading in a
  fallback serif; it also only requested weight 600 while the CSS
  asked for 900, so Latin and CJK in the same heading disagreed.
- Docs were shipped as one 968KB manifest downloaded on every page
  view. Split into a 32KB index plus one lazily imported chunk per
  page; the entry bundle is now 101KB gzipped.
- Add search, syntax highlighting, per-route meta with canonical and
  hreflang, a sitemap, and an error boundary. Replace the 44vh
  mobile sidebar with a drawer.
- Image dimensions are read at build time and written into the tag,
  so lazy images reserve their space instead of collapsing.

Screenshots are recaptured from a real v0.5.0 build against a clean
demo project, with tokens, QR codes and paired accounts redacted.
The previous set is deleted rather than kept alongside.

Routes follow file paths, so the restructure would have broken every
inbound link; 37 old paths redirect, in both languages. The PR policy
gate and CODEOWNERS also hardcoded docs/guide/contributing.md.

Verified: check:docs 78 pages / 323 links / 0 problems, check:policy
127 pass. Walked every route at 1440 and 390 in both themes for
overflow, contrast, keyboard reachability and focus management.
This commit is contained in:
程序员阿江(Relakkes)
2026-07-27 17:32:41 +08:00
parent c64f69972d
commit c2cd615824
299 changed files with 9059 additions and 11274 deletions
-129
View File
@@ -1,129 +0,0 @@
# Claude Code Multi-Agent System Documentation
> Complete guide and technical reference for multi-agent orchestration
---
## Documentation Index
### [01-usage-guide.md](./01-usage-guide.md) — Usage Guide
A comprehensive user-facing manual covering:
- **Agent Tool**: Parameter reference, spawn methods, background execution
- **Six Built-in Agents**: general-purpose, Explore, Plan, verification, claude-code-guide, statusline-setup
- **Background Tasks**: Asynchronous execution, progress tracking, completion notifications
- **Agent Teams**: Team creation, member collaboration, message communication
- **Worktree Isolation**: Independent environments, branch management, secure contexts
- **Custom Agents**: Definition format, tool pool configuration, system prompts
**Target Audience**: All Claude Code users
---
### [02-implementation.md](./02-implementation.md) — Implementation Details
A deep technical reference for developers covering:
- **Architecture Overview**: 5 agent categories, 4 spawn paths
- **Agent Spawn Flow**: Detailed walkthrough of Sync / Async / Fork / Teammate paths
- **Tool Pool System**: Three-layer filtering, constant definitions, permission mapping
- **Context Passing**: CacheSafeParams, system prompt construction, fork cache optimization
- **Teams Internals**: TeamFile structure, mailbox system, inbox polling, message routing
- **Background Task Engine**: LocalAgentTask lifecycle, progress tracking, notification queue
- **Permission Synchronization**: Team-level permissions, mode propagation, bubble mode
- **End-to-End Data Flow**: From Agent Tool invocation to result delivery
**Target Audience**: Contributors, architects, and developers seeking deep implementation understanding
---
### [03-agent-framework.md](./03-agent-framework.md) — Agent Framework Deep Dive
Deconstructing the architecture behind Claude Code's agent framework from source code, covering:
- **Core Agent Loop**: AsyncGenerator state machine, five-phase while(true) loop
- **System Prompt Engineering**: Layered construction, cache boundary, CLAUDE.md loading
- **Tool System Design**: Full lifecycle management, three-stage registration, 7-step execution pipeline
- **Context Management & Compression**: Four-level progressive compression, system context injection
- **Skills & Plugin Ecosystem**: Skill definition and discovery, plugins, hooks, MCP integration
- **Permission & Security Model**: Layered permission model, rule pattern matching
- **Fault Recovery Mechanisms**: 6 built-in recovery strategies, model fallback
- **Comparison with LangChain/ReAct**: Architecture paradigm differences, why not ReAct
- **Why Claude Code Is So Good**: 7 core design principles
**Target Audience**: Architects studying AI agent design, AI application developers, technical researchers
---
## Illustration Notes
All diagrams use a dark background (#1a1a2e) with Claude Code Haha orange-blue accent (#FF7A00), consistent with Claude Code's official documentation style.
| Image | Description | Document |
|-------|-------------|----------|
| `01-agent-overview.png` | Multi-Agent System Overview — Architecture panorama | Usage Guide |
| `02-agent-types.png` | Six Built-in Agents — Type comparison matrix | Usage Guide |
| `03-spawn-flow.png` | Agent Spawn Flow — Four-path decision tree | Usage Guide |
| `04-agent-teams.png` | Agent Teams Collaboration — Team communication topology | Usage Guide |
| `05-architecture.png` | Implementation Architecture — Core module relationships | Implementation |
| `06-context-passing.png` | Context Passing — CacheSafeParams data flow | Implementation |
| `07-tool-pool.png` | Tool Pool System — Three-layer filtering pipeline | Implementation |
| `08-background-task.png` | Background Task Engine — Lifecycle state machine | Implementation |
| `09-teams-mailbox.png` | Teams Mailbox System — Message routing topology | Implementation |
| `10-fork-cache.png` | Fork Cache Optimization — Byte-level consistent sharing | Implementation |
| `11-agent-framework-overview.png` | Agent Framework Overview — Core component relationships | Framework Deep Dive |
| `12-agent-core-loop.png` | Core Agent Loop — Five-phase state machine | Framework Deep Dive |
| `13-system-prompt-pipeline.png` | System Prompt Pipeline — Layered cache pipeline | Framework Deep Dive |
| `14-context-compression.png` | Context Compression — Four-level progressive strategy | Framework Deep Dive |
---
## Quick Start
### For Users
1. Read the [Usage Guide](./01-usage-guide.md)
2. Learn about the six built-in agents and their use cases
3. Try spawning subagents using the Agent tool in a conversation
4. Explore multi-agent collaboration with Agent Teams
### For Developers
1. Read the [Implementation Details](./02-implementation.md)
2. Browse the source code:
- `src/tools/AgentTool/` — Agent tool implementation
- `src/tools/TeamCreateTool/` — Team creation
- `src/tools/SendMessageTool/` — Inter-agent communication
- `src/utils/swarm/` — Swarm collaboration infrastructure
- `src/utils/forkedAgent.ts` — Fork agent context
- `src/tasks/` — Task management system
3. Understand the four spawn paths and context passing mechanisms
---
## Core Concepts Quick Reference
| Concept | Description |
|---------|-------------|
| **Agent Tool** | Primary entry point — accepts a prompt + subagent_type to spawn a subagent |
| **Subagent** | An independent child agent that executes tasks with its own tool pool and permissions |
| **Fork Agent** | A forked agent that inherits the parent's full context and shares prompt cache |
| **Teammate** | A collaborative member within an Agent Team, communicating via mailbox |
| **Worktree** | Git worktree isolation mode providing an independent file environment |
| **LocalAgentTask** | Local agent task state, tracking running/completed/failed status |
| **DreamTask** | Automatic memory consolidation task that runs periodically in the background |
| **CacheSafeParams** | Cache-safe parameters ensuring byte-level consistency of API request prefixes |
| **TeamFile** | Team configuration file storing the member list and permissions |
| **Mailbox** | File-based message queue supporting asynchronous communication between teammates |
---
## Related Resources
- [Claude Code Haha Home](/)
- [Memory System Documentation](/en/memory/01-usage-guide)
- [Agent Tool Source Code](https://github.com/NanmiCoder/cc-haha/tree/main/src/tools/AgentTool/)
- [Swarm Infrastructure](https://github.com/NanmiCoder/cc-haha/tree/main/src/utils/swarm/)
- [Task Management System](https://github.com/NanmiCoder/cc-haha/tree/main/src/tasks/)
- [GitHub Issues](https://github.com/NanmiCoder/cc-haha/issues)
-56
View File
@@ -1,56 +0,0 @@
# Channel System Research
This section documents the upstream Claude Code Channel/MCP architecture. It is not the recommended setup path for the IM integrations shipped by this repository.
To connect WeChat, DingTalk, WhatsApp, Telegram, or Feishu to the current Desktop app, start with [IM Integrations](../im/).
## Current product path
The working integration in this repository is:
```text
Desktop settings
→ /api/adapters
→ ~/.claude/adapters.json
→ adapters/*
→ pairing and allowlist
→ HTTP session creation
→ /ws/:sessionId
→ Claude Code session
```
This adapter path is smaller and separate from the upstream Channel mechanism. It matches the Desktop server’s REST and WebSocket architecture and keeps platform authorization in dedicated sidecar processes.
## Documents
### [Channel System Architecture](./01-channel-system.md)
A source-level analysis of the upstream system, including:
- MCP Channel capability and notification flow;
- inbound XML wrapping and outbound tool calls;
- runtime, OAuth, organization, session, and plugin gates;
- permission relay and short request IDs;
- plugin registration and trust boundaries.
### Historical IM Gateway proposal
The Chinese documentation retains an early IM Gateway proposal as architecture history. It explains why the implementation moved toward independent adapters and `/ws/:sessionId`. It is not a current setup guide and is intentionally not presented as an English user workflow.
## When to read this section
Use these pages when you are:
- studying the upstream Claude Code Channel design;
- comparing Channel plugins with the repository’s adapter implementation;
- designing a future protocol or plugin integration;
- reviewing the security boundaries of remote Agent control.
For a working bot or linked account, use:
- [IM Integration overview](../im/)
- [Telegram](../im/telegram.md)
- [Feishu](../im/feishu.md)
- [WeChat](../im/wechat.md)
- [DingTalk](../im/dingtalk.md)
- [WhatsApp](../im/whatsapp.md)
@@ -1,3 +1,10 @@
---
title: Environment Variables
nav_title: Env Vars
description: Authentication, model, Azure, and local-runtime variables, plus the real precedence order.
order: 2
---
# Environment Variables
Claude Code Haha has two configuration paths:
@@ -39,7 +46,7 @@ Azure OpenAI uses a dedicated Responses API path:
Example:
```bash
```ini
CLAUDE_CODE_USE_AZURE_OPENAI=1
AZURE_OPENAI_BASE_URL=https://your-resource.cognitiveservices.azure.com
AZURE_OPENAI_API_VERSION=2025-04-01-preview
@@ -57,7 +64,7 @@ AZURE_OPENAI_CODEX_DEPLOYMENT=your_codex_deployment
| `DISABLE_TELEMETRY` | Set to `1` to disable telemetry |
| `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` | Set to `1` to disable non-essential network requests |
See [Local Server](../reference/local-server.md) for `SERVER_HOST`, `SERVER_PORT`, `SERVER_AUTH_REQUIRED`, and related variables.
See [Local Server](../internals/server.md) for `SERVER_HOST`, `SERVER_PORT`, `SERVER_AUTH_REQUIRED`, and related variables.
## Configuration methods
@@ -71,7 +78,7 @@ Desktop stores the provider index at:
Provider-managed environment data is written to an isolated Haha configuration. You do not need to copy it into `~/.claude/settings.json`. When the CLI finds an active provider, it reuses its credentials, models, and protocol settings. Providers using `openai_chat` or `openai_responses` automatically use a loopback proxy.
See [Third-Party Models](./third-party-models.md) for the setup flow.
See [Third-Party Models](../start/models.md) for the setup flow.
### Repository `.env`
@@ -83,7 +90,7 @@ cp .env.example .env
Minimal example for an Anthropic-compatible endpoint:
```bash
```ini
ANTHROPIC_AUTH_TOKEN=sk-example
ANTHROPIC_BASE_URL=https://provider.example.com/anthropic
ANTHROPIC_MODEL=provider-model
@@ -127,5 +134,5 @@ Keep one primary provider configuration source. Desktop users should use the Pro
- Never commit `.env`, provider configuration, or a `settings.json` containing secrets.
- Do not expose full tokens in screenshots, issues, logs, or diagnostic archives.
- Use `CLAUDE_CONFIG_DIR` to isolate tests from real user configuration.
- `--print` skips the workspace trust dialog and must only run in trusted directories. See [CLI Reference](./cli-reference.md).
- `--print` skips the workspace trust dialog and must only run in trusted directories. See [CLI Reference](./reference.md).
- Do not treat CORS as authentication when exposing the local server. Enable an H5 token or explicit authentication.
+115
View File
@@ -0,0 +1,115 @@
---
title: Install and Run
nav_title: Install and Run
description: Run the CLI from source - install dependencies, configure a provider, launch from any directory.
order: 0
---
# Install and Run
The CLI is the core of Claude Code Haha — every desktop session runs one underneath. If you only want the graphical app, installing that is enough; see [Download and install](../start/install.md). The steps below are for people who want a terminal workflow, `--print` automation, or a source checkout to read and contribute to.
The CLI runs from source only. There is no separate installer for it.
## Get the source
Install [Git](https://git-scm.com/downloads) and [Bun](https://bun.sh) first, then:
```bash
git clone https://github.com/NanmiCoder/cc-haha.git
cd cc-haha
bun install
```
## Configure a model provider
```bash
cp .env.example .env
```
Edit `.env` with at least one working authentication method, base URL, and model. A minimal Anthropic-compatible setup looks like this:
```ini
ANTHROPIC_AUTH_TOKEN=sk-example
ANTHROPIC_BASE_URL=https://provider.example.com/anthropic
ANTHROPIC_MODEL=provider-model
```
See [Environment variables](./env.md) for what each variable means, how the authentication headers differ, and how to configure Azure and other protocols. If you already configured a provider in the desktop app, the CLI reuses it and you do not need a `.env` at all.
Never commit a real API key, and never paste one into an issue, a screenshot, or a diagnostics bundle.
## Start and verify
macOS, Linux, or Git Bash:
```bash
./bin/claude-haha
./bin/claude-haha -p "Summarize the directory structure of this project"
```
Windows PowerShell or cmd:
```powershell
bun --env-file=.env ./src/entrypoints/cli.tsx
```
Once you see streaming output and tool calls, the provider, the project directory, and the CLI are connected.
## Run from any directory
`./bin/claude-haha` only works inside the checkout. Put it on your `PATH` and you can type `claude-haha` in any project directory — the CLI treats the current working directory as the project root.
On macOS and Linux, add this to `~/.bashrc` or `~/.zshrc`:
```bash
# Option 1: add to PATH (recommended)
export PATH="$HOME/path/to/claude-code-haha/bin:$PATH"
# Option 2: alias
alias claude-haha="$HOME/path/to/claude-code-haha/bin/claude-haha"
```
Reload the shell config:
```bash
source ~/.zshrc # or source ~/.bashrc
```
On Windows, add the same `PATH` line to `~/.bashrc` under Git Bash:
```bash
export PATH="$HOME/path/to/claude-code-haha/bin:$PATH"
```
To verify, start it from a different directory and ask what the current directory is:
```bash
cd ~/your-other-project
claude-haha
```
### Windows with a WSL toolchain
If `claude-haha` runs on Windows or Git Bash while Node, Python, uv, and bun live inside WSL, call them through WSL explicitly:
```bash
wsl -e bash -lc 'node --version && python3 --version'
```
When the CLI detects a `wsl` / `wsl.exe` invocation, it sets `MSYS2_ARG_CONV_EXCL=*` so Git Bash does not rewrite WSL paths such as `/home/...` into `C:/Program Files/Git/home/...`.
To route Bash tool commands through WSL by default, set this before startup:
```bash
export CLAUDE_CODE_SHELL_PREFIX='wsl -e bash -lc'
```
Computer Use still controls Windows desktop apps, so CLI tools inside WSL do not need an entry in `computer-use-config.json`. If you only need the WSL toolchain and no desktop control, pass `--no-computer-use` or turn it off under Settings → Computer Use.
## Next steps
- [Command reference](./reference.md): flags, headless mode, and recovery mode
- [Environment variables](./env.md): the full list of authentication, model, and runtime variables
- [Architecture overview](../internals/index.md): how the CLI, server, desktop shell, and adapters divide the work
- [Contributing and quality gates](../internals/contributing.md): what to run before opening a PR
@@ -1,4 +1,11 @@
# CLI Reference
---
title: Command Reference
nav_title: Commands
description: Interactive sessions, --print headless mode, recovery mode, and common command-line flags.
order: 1
---
# Command Reference
Claude Code Haha starts an interactive session by default. With `--print`, it can also run as a non-interactive agent in scripts, CI, or another program. The command in a source checkout is `./bin/claude-haha`; the installed executable name depends on the installation method.
-138
View File
@@ -1,138 +0,0 @@
# Desktop Quick Start
This walkthrough takes you from installation to a working session.
![The Main Session with its project and history sidebar expanded, keeping the conversation, composer, model, and permission controls at the center](../../images/desktop_ui/25_main_session.png)
## 1. Install and open the app
Download the package for your platform from [GitHub Releases](https://github.com/NanmiCoder/cc-haha/releases). Detailed package names and platform notes are in the [Installation Guide](./04-installation.md).
On first launch:
- macOS may show the standard downloaded-app confirmation. Official releases from `v0.4.3` onward are signed and notarized.
- An unsigned Windows build may show SmartScreen. Confirm that the package came from the project’s GitHub Release before choosing **More info → Run anyway**.
- Linux AppImage users may need FUSE, depending on the distribution.
## 2. Confirm the data location
Desktop uses `~/.claude` by default for sessions, providers, settings, Skills, Agents, memory, tasks, traces, and adapter state. Most users should keep this default. If you need portable or isolated storage, select an explicit custom directory in General settings and restart when prompted.
Do not point automated tests or temporary experiments at your real user data directory.
## 3. Connect a model provider
Open **Settings → Providers**. There are four main paths:
| Path | Use it when |
|---|---|
| Claude official login | You want models available to your Claude account |
| ChatGPT official login | You want the current GPT and Codex catalog available to your account |
| Grok official login | You want the xAI model catalog available to your account |
| Custom provider | You have an Anthropic-compatible or OpenAI-compatible endpoint and API key |
Custom providers can use `anthropic`, `openai_chat`, or `openai_responses` format. Presets are editable, and the final model list still depends on the endpoint.
Use **Test connection** before starting a session. Never paste API keys into chat messages or screenshots.
## 4. Create a session
Select **New session**, then choose:
- a project directory;
- the current working tree or an isolated Worktree when Git options are available;
- a model and supported reasoning effort;
- a permission mode.
Sessions are bound to their working directory. The sidebar groups history by project and time, and open sessions can be kept in separate tabs.
## 5. Choose a permission mode
The current desktop supports five modes:
![The New Session permission menu showing Ask, Accept Edits, Auto, Plan, and Bypass Permissions](../../images/desktop_ui/18_permission_modes.png)
| Mode | Behavior |
|---|---|
| Ask | Requests approval before protected operations |
| Accept Edits | Automatically accepts eligible file edits, while other protected actions can still ask |
| Plan | Lets the model investigate and propose a plan without carrying out normal implementation work |
| Auto | Reviews operations automatically and allows or blocks them according to the Auto policy |
| Bypass Permissions | Skips eligible approval prompts; explicit denials and tools that require user interaction still apply |
Auto and Bypass Permissions expand what can run without a click. Read the confirmation dialog and use them only in a project and environment you trust. The mode cannot be changed while an active turn makes that unsafe.
## 6. Chat and review work
The composer supports:
- text and multiline input;
- pasted images, drag-and-drop, and file selection;
- `/` commands;
- `@` workspace file search;
- model and effort selection when supported.
While a turn runs, review tool calls, permission requests, tasks, SubAgents, teams, sources, changed files, and previews. Use the workspace panel for files and diffs, and the activity panel for background work.
## 7. Explore the workspace
The workspace can search the complete project without requiring every directory to be expanded first. Packaged desktop builds include the appropriate ripgrep binary.
![The expanded workbench listing real Git changes, file types, and line counts](../../images/desktop_ui/22_workspace_changed_files.png)
You can:
- preview files and local attachments;
- open files with the system default app or a configured editor;
- select one or several diff lines and send review comments back to chat;
- inspect the current branch, Worktree, and changed-file summary;
- open the embedded terminal and browser preview.
![A real code diff with the inline local-comment editor open on a changed line](../../images/desktop_ui/23_workspace_diff_review.png)
## 8. Manage Skills, Agents, and pets
### Install Skills
**Skills** shows installed Skills and a marketplace for discovering, reviewing, installing, and removing supported third-party Skills. Review source and risk information before installation.
![The Skills marketplace with source status, the third-party warning, filters, security labels, and installation states](../../images/desktop_ui/21_skill_marketplace.png)
### Create or edit an Agent
Open **Settings → Agents**, then select **Create Agent**.
![Create Agent dialog with scope, model, reasoning effort, tools, and system prompt fields](../../images/desktop_ui/17_agent_create.png)
1. Choose **User** scope to reuse the Agent across projects, or **Project** scope to keep it with the current project.
2. Add a name, description, and system prompt that clearly define the Agent's responsibility and expected output.
3. Inherit the main session model and effort, or select a model and supported effort specifically for this Agent.
4. Keep all tools, disable tools, or choose searchable built-in, MCP, and custom tool rules.
5. Save the Agent. User Agents are stored in `~/.claude/agents/`; project Agents are stored in the current project's `.claude/agents/`.
Built-in, plugin, and policy Agents remain read-only. See the [Multi-Agent Usage Guide](../agent/01-usage-guide.md#6-custom-agents) for definition formats, precedence, and inheritance.
### Turn on the desktop pet
Open **Settings → Pets** and enable **Show desktop pet**.
![Desktop Pet settings showing four built-in pets and appearance controls](../../images/desktop_ui/14_pet_settings_overview.png)
1. Choose Dada, Huhu, Bubu, or Huihui.
2. Adjust the size and animation setting.
3. Enable the active-task panel if you want it to open automatically while a task is running. The panel remains hidden when no task is active.
4. Hover for a small reaction, click the pet to focus the main window, or drag it to another position. Right-click the pet to close it; return to **Settings → Pets** to show it again.
Select **Add pet** to make one of your own: use a transparent PNG or WebP you already have, or follow the in-dialog prompt to have any drawing AI produce an action sheet and pick that (sizes are aligned for you). Everything happens locally, without calling the current chat model or using chat quota. Pets run only in the Electron desktop app and are not supported in H5.
Follow the complete [Desktop Pet Guide](./pets.md) for custom-image requirements, task states, and interaction details.
## 9. Optional remote access
- Use [IM integrations](../im/) for explicitly paired private-chat users.
- Use [H5 Access](./06-h5-access.md) only on a trusted LAN or behind a reverse proxy you control.
- Local loopback Web UI access is separate from H5 and does not grant remote access.
## 10. If something fails
Open **Settings → Diagnostics** to inspect the desktop runtime, local index, sidecar, providers, and recoverable state. Read [FAQ](./05-FAQ.md) before manually editing stored JSON.
-143
View File
@@ -1,143 +0,0 @@
# Desktop Feature Guide
This page summarizes the current user-visible Desktop surface.
## Sessions and chat
- Multiple tabs and project-filtered session history
- Streaming Markdown, code, Mermaid, tool calls, thinking blocks, and permission requests
- Image, file, directory, and PDF attachments with context preserved across restart and resume
- Slash commands, `@` file search, model selection, and supported reasoning-effort levels
- Stop, resume, search, navigation outline, and checkpoint-based undo for the current turn
- Restored scroll position and stable search in long, virtualized conversations
## Permission controls
Desktop supports Ask, Accept Edits, Plan, Auto, and Bypass Permissions. Auto evaluates operations according to its policy; Bypass skips only approvals that are eligible to be skipped. Explicit denials, protected boundaries, and tools that require direct user interaction remain enforced.
![The New Session permission menu showing all five execution modes](../../images/desktop_ui/18_permission_modes.png)
The selected mode is visible in the session composer. Higher-autonomy modes require an explicit warning confirmation.
## Workspace and code review
The workspace panel combines:
![The expanded workbench listing changed files, file types, and Git line counts](../../images/desktop_ui/22_workspace_changed_files.png)
- full-project file search using the packaged ripgrep binary;
- a file tree, previews, and external-editor actions;
- changed-file summaries and inline diffs;
- single-line or contiguous-line review comments sent back into chat;
- embedded terminal and local browser preview.
Branch and Current/Isolated Worktree selection belongs to the new-session launch controls. The workspace then shows the selected session's files, Git status, and diffs.
![A real code diff with the inline local-comment editor open](../../images/desktop_ui/23_workspace_diff_review.png)
### Browser Preview
Switch the workbench from **Files** to **Browser** to open a local development page or HTTPS URL inside the app.
![The embedded Browser Preview showing a safe documentation example page, address bar, screenshot action, and element picker](../../images/desktop_ui/24_browser_preview.png)
- **Screenshot** sends the current page image back to the session.
- **Select element** captures a page element, selector, and screenshot as explicit Agent context.
- Treat authenticated pages, cookies, and external-site content as sensitive. Use a public demo page for documentation screenshots.
## Tasks, SubAgents, and activity
The activity panel collects foreground tasks, background tasks, SubAgents, team members, and sources without flooding the main transcript.
Running SubAgents can be opened before they finish. Their status, output, tool activity, duration, and terminal state continue to refresh. Completed, failed, and stopped terminal states are persisted so they do not revert to “running” after restart.
![The scheduled-task form with prompt, model, working directory, full-permission warning, frequency, and notification controls](../../images/desktop_ui/20_scheduled_task.png)
Scheduled tasks support cron-based runs, model selection, enable/disable controls, manual execution, notifications, and history. New scheduled tasks currently run with Bypass Permissions; the form does not provide a permission selector. Use the smallest practical working directory and review the prompt before saving. Tasks run only while the Desktop app and computer remain available.
## Providers and models
Provider settings support:
- Claude, OpenAI, and Grok official login;
- editable built-in presets;
- Custom Anthropic, OpenAI Chat, and OpenAI Responses endpoints;
- per-provider model mappings, context metadata, proxy mode, and supported effort levels;
- connection testing and runtime model discovery where available.
The visible model catalog and supported effort levels depend on the current account, provider, and runtime response. Release-specific model names belong in release notes rather than this evergreen guide.
## Skills and marketplace
The Skills area groups local Skills by source and exposes their metadata and files. The marketplace aggregates supported third-party sources with search, filters, details, file previews, installation confirmation, local installation state, and risk information.
![The Skills marketplace showing source health, the third-party warning, filters, security status, and install state](../../images/desktop_ui/21_skill_marketplace.png)
Marketplace metadata is not a security guarantee. Review the source and requested behavior before installing a Skill.
## Agent management
The Agents page shows built-in, user, project, plugin, and policy sources together with override and effective-state information.
![Create Agent dialog showing scope, prompt, model, effort, and tool controls](../../images/desktop_ui/17_agent_create.png)
Editable user and project Agents support:
- reusable user scope or current-project scope;
- model inheritance or an explicit model ID;
- supported effort from `low` through `max`;
- searchable tool selection plus custom and MCP tool rules;
- a system prompt and description.
User Agents are stored in `~/.claude/agents/`, while project Agents are stored in the current project's `.claude/agents/`. Built-in, plugin, and policy definitions remain read-only. Saving a definition and refreshing a running session are separate outcomes; a refresh failure does not silently discard a saved file.
## Desktop pets
The optional pet window is a native transparent Electron surface. Open **Settings → Pets**, enable the window, and choose one of the four built-in companions: Dada, Huhu, Bubu, or Huihui.
![Desktop Pet settings with the enable switch, built-in companions, and appearance controls](../../images/desktop_ui/14_pet_settings_overview.png)
The pet reflects working, waiting-for-you, failed, and idle states. While a task is active, its task panel can show the session title and status; selecting a row returns to that session. The panel remains hidden when no task is active. If automatic display is disabled, an activity-count badge lets you open it when work is running.
The window also supports direct interaction:
- hover to trigger a small reaction and idle gaze tracking;
- click to wave and focus the main window;
- drag to move the window, with its position restored after restart;
- right-click to close it, then use **Settings → Pets** to show it again.
### Add a custom pet
Select **Add pet** to choose a local creation path.
![Add pet dialog with three ways to make a pet](../../images/desktop_ui/15_pet_create_methods.png)
- **Use a picture you already have** turns a static transparent PNG or WebP into a lightweight locally animated pet. It will not run or track the cursor.
- **Draw one with AI that runs and jumps** supplies a copyable prompt, a reference template and a checklist, so any drawing AI can produce an 8-column × 9-row action sheet you then pick.
- **I already have an action sheet** skips the walkthrough and imports a sheet or finished atlas directly.
The last two paths slice, rescale and mirror the sheet locally on import, so exact source dimensions are not required; a finished `1536×2288` atlas is kept as-is.
Imports stay local and do not call the selected chat model. After a successful import, the custom companion appears under **Your pets** and becomes the selected pet.
![A locally imported custom pet selected in the Your pets list](../../images/desktop_ui/16_pet_custom_result.png)
Size, animation, active-task-panel, and window-position preferences are restored between launches. Pets run only in the Electron desktop app and are not supported in H5. See the [Desktop Pet Guide](./pets.md) for image limits, atlas layout, storage, and troubleshooting.
## Local data and search
A rebuildable SQLite projection accelerates session lists, global search, activity statistics, scheduled runs, teams, and traces. Original JSON and JSONL files remain authoritative. If the index is unavailable or damaged, readers can fall back to source files, and Diagnostics provides a confirmed rebuild action.
Rebuilding the index does not delete chats, settings, Skills, memory, tasks, or traces.
## IM and H5
The desktop can launch sidecar adapters for WeChat, DingTalk, WhatsApp, Telegram, and Feishu. Each platform is private-chat oriented and requires an allowlisted or paired user. See [IM Integrations](../im/).
[H5 Access](./06-h5-access.md) serves the browser UI from the local server for a trusted LAN or controlled reverse proxy. It uses a separate token and origin policy and is not a public multi-user account system.
## Diagnostics and recovery
Diagnostics surfaces local-index health, sidecar and runtime information, fallback state, storage information, and safe rebuild actions. Startup failures that cannot be recovered automatically should remain visible rather than leaving only a background process.
Doctor and repair flows are intentionally conservative: protected user data is not silently reset.
-103
View File
@@ -1,103 +0,0 @@
# Installation
Claude Code Haha Desktop is built with Electron and ships packages for macOS, Windows, and Linux. Official macOS releases from `v0.4.3` onward use Developer ID signing and notarization; older or temporary builds may still require manual approval.
## Download
Open [GitHub Releases](https://github.com/NanmiCoder/cc-haha/releases) and select the package for your platform:
| Platform | Package |
|---|---|
| macOS Apple Silicon | `Claude-Code-Haha-<version>-mac-arm64.dmg` |
| macOS Intel | `Claude-Code-Haha-<version>-mac-x64.dmg` |
| Windows x64 | `Claude-Code-Haha-<version>-win-x64.exe` |
| Windows ARM64 | `Claude-Code-Haha-<version>-win-arm64.exe` |
| Linux x64 | `...-linux-x86_64.AppImage` or `...-linux-amd64.deb` |
| Linux ARM64 | `...-linux-arm64.AppImage` or `...-linux-arm64.deb` |
On macOS, **Apple M-series** means arm64; **Intel** means x64.
## macOS
Open the DMG and drag the app to `Applications`. An official signed release should only show the normal downloaded-app confirmation.
For an older or explicitly unsigned temporary build, macOS may report that the app is damaged or cannot verify the developer. After confirming the package source, either use **System Settings → Privacy & Security → Open Anyway**, or run:
```bash
xattr -cr /Applications/Claude\ Code\ Haha.app
```
Do not use this workaround for an untrusted download.
## Windows
Run the `.exe` installer. If an unsigned package triggers SmartScreen, verify that it came from the expected GitHub Release, then select **More info → Run anyway**.
The installer supports x64 and ARM64 packages. Close running Claude Code Haha processes before an overwrite install. User data is kept separately from application files.
## Linux
For AppImage:
```bash
chmod +x Claude-Code-Haha-<version>-linux-x86_64.AppImage
./Claude-Code-Haha-<version>-linux-x86_64.AppImage
```
If FUSE is missing, Ubuntu 22.04 and earlier normally use `libfuse2`; Ubuntu 24.04 and later normally use `libfuse2t64`.
For a deb package:
```bash
sudo apt install ./Claude-Code-Haha-<version>-linux-amd64.deb
```
## Local Web UI
For source development, run the server from the repository root and Vite from `desktop/`:
```bash
SERVER_PORT=3456 bun run src/server/index.ts
```
In a second terminal:
```bash
cd desktop
bun run dev --host 127.0.0.1 --port 2024
```
Open `http://127.0.0.1:2024`.
True loopback access from the same machine does not require H5 to be enabled or an H5 token. This trust does not extend to LAN addresses or reverse proxies.
## Headless Linux over SSH
Keep both processes bound to loopback:
```bash
SERVER_PORT=3456 bun run src/server/index.ts
cd desktop
bun run dev --host 127.0.0.1 --port 2024
```
Forward both ports from your own computer:
```bash
ssh -L 2024:127.0.0.1:2024 -L 3456:127.0.0.1:3456 user@example.com
```
Then open:
```text
http://127.0.0.1:2024/?serverUrl=http%3A%2F%2F127.0.0.1%3A3456
```
This does not expose the service to the LAN. To serve the built Web UI to a trusted LAN or reverse proxy, follow [H5 Access](./06-h5-access.md#enable-h5-without-the-desktop-ui).
## Updates and data
Official releases check GitHub Releases for updates. In-place updates and overwrite installs are designed to preserve local settings and sessions.
The default data location is under `~/.claude`. A custom data directory can be selected in the app. Before changing storage manually, use the in-app controls and read the recovery guidance in [FAQ](./05-FAQ.md).
-122
View File
@@ -1,122 +0,0 @@
# Desktop FAQ
## Where is my data stored?
The default data root is under `~/.claude`. It contains user-owned sessions, provider settings, Skills, Agents, memory, tasks, traces, adapter configuration, and derived desktop state.
Use the General settings control when changing the data directory. Do not delete or replace the directory as a first troubleshooting step.
## Why does macOS say the app is damaged?
Official macOS releases from `v0.4.3` onward are signed and notarized. This message is more likely with an older or temporary unsigned package. Confirm that the file came from the project’s GitHub Release, then follow the macOS steps in [Installation](./04-installation.md#macos).
## Why does Windows show SmartScreen?
Windows signing is not a release requirement for every build. An unsigned installer can trigger SmartScreen even when the package is intact. Verify the GitHub Release and architecture before selecting **More info → Run anyway**.
## A Custom provider returns 401. What should I check?
Open **Settings → Providers**, verify the Base URL, API format, API key, and model ID, then run **Test connection**.
Authentication headers differ by provider:
- standard API-key providers normally use the API key field;
- bearer-token endpoints require the corresponding provider/auth configuration;
- OpenAI-compatible endpoints must use the correct `openai_chat` or `openai_responses` format;
- Kimi Code uses the current K3 Coding API preset and model metadata.
Prefer editing or recreating the provider in the UI over manually rewriting stored JSON. Never share the stored credential or a diagnostic screenshot containing it.
## Why is an official model missing?
Official-login catalogs are filtered by the models available to the account and by runtime discovery. A model listed in the app’s metadata is not a guarantee that every account can call it.
Sign out and back in, refresh the catalog, and confirm account access. For Custom providers, verify the endpoint’s real model ID and capability support.
## Why was my reasoning effort reduced or ignored?
Effort is applied only when the selected model and provider support it. Per-Agent effort follows the same rule. Unsupported combinations can be normalized, reduced, or omitted rather than forcing an invalid API request.
## My old sessions are missing from the list. Were they deleted?
Usually not. The SQLite local index is a rebuildable projection; original session files remain authoritative. Open **Settings → Diagnostics**, inspect index state, and use the confirmed rebuild action if needed.
Rebuilding the index does not delete source sessions, settings, Skills, memory, tasks, or traces.
## The app opens to a blank or startup error screen
Restart once and open Diagnostics if the app recovers. If a visible startup error remains:
1. record the exact error and platform;
2. confirm the package architecture;
3. check that security software did not quarantine a sidecar;
4. avoid deleting the data directory;
5. report the issue with the release version and sanitized logs.
The desktop has bounded renderer and sidecar recovery, but it intentionally leaves an unrecoverable startup failure visible.
## H5 asks for a token
Remote browser access requires H5 to be enabled and a current token. True same-machine loopback development access is separate and does not require H5.
Regenerating the token invalidates the previous one. See [H5 Access](./06-h5-access.md).
## An IM bot says I am unauthorized
Binding a bot or linked account does not authorize every chat user. The user must be in `allowedUsers` or complete pairing with the current six-character code. Codes expire after 60 minutes, are one-time use, and are rate limited after repeated failures.
See [IM Integrations](../im/).
## Will updates delete my providers or chats?
Official in-app and overwrite updates are designed to preserve the user data directory. Back up important user-owned data before unusual manual migrations, and do not treat application build directories as the data source.
## Desktop pets
See [Desktop Pets](./pets.md) for the complete enable, customization, and import workflow.
### Why is the active-task panel missing?
The panel lists only non-idle sessions, such as tasks that are running, waiting for you, or need attention. It hides automatically when every task is complete or idle; enabling **Show active task panel** does not keep historical tasks visible.
Confirm **Settings → Pets → Show active task panel** is enabled, then start a real task. Status refreshes periodically; if it still does not appear after a few seconds, hide and show the pet again.
### Why is my pet not moving?
Confirm **Settings → Pets → Play animations** is enabled. The app also respects the operating system’s reduced-motion preference, which can reduce or disable animation.
A custom pet made from one image receives lightweight breathing, floating, and task-state motion. It does not have the complete frame-by-frame actions of a v2 atlas pet. Hover over or click the pet to trigger interactions such as jumping or waving.
### Why did my custom-pet import fail?
Check each requirement:
- A single image must be PNG or WebP, 32–4096 pixels on each side, no larger than 8 MB, and no more than 16,777,216 total pixels.
- An action sheet must use an 8-column × 9-row layout. Exact pixel dimensions are not required — it is sliced and rescaled on import. A finished 1536×2288 atlas is kept as-is.
- An action sheet must have a transparent background. White or coloured backdrops are rejected because the pet would render as a square on the desktop.
- The pet ID may be at most 73 characters, contain only lowercase letters, numbers, and single hyphens, and must not duplicate an existing ID.
- Display name and description are required.
The import reads only the local file you confirm in the system picker. Fix the source image and create the pet again instead of manually editing its manifest.
### How do I restore a pet after closing it?
Right-clicking the pet and choosing close, or turning off **Settings → Pets → Show desktop pet**, saves the disabled state. Turn that setting on again to restore the pet; reinstalling the app is not necessary.
If the pet was dragged onto another display, hide and show it again. Its saved position is clamped into the currently visible work area when the pet window is recreated.
### How do I delete a custom pet?
There is no in-app delete button in the current version. Select a built-in pet first, choose **Settings → Pets → Open folder**, delete only the folder for the custom pet, then return to Settings and select **Refresh**.
Delete only the confirmed pet directory under `${CLAUDE_CONFIG_DIR:-~/.claude}/cc-haha/pets`. Do not delete the entire `~/.claude` directory. Built-in pets ship with the app and are not removed here.
### Does “Draw one with AI that runs and jumps” use my chat quota?
No. That path never calls the chat model selected for the session, and the app does not request any image-generation service on your behalf. It gives you a prompt and a reference template; you generate the picture in whichever drawing tool you like and then pick the file. Assembly and import both happen on your own machine.
### Can the pet approve permissions or send replies?
No. The pet window reports task state, focuses the main window, and opens the corresponding session. It has no message composer and cannot approve file, command, or Computer Use permissions.
When it shows **Waiting for you**, open the task in the main window, review the full context, approve or reject the request, and reply there. The pet never bypasses the selected permission mode.
-120
View File
@@ -1,120 +0,0 @@
# H5 Access
H5 is an optional browser entry point to the same local chat UI used by Desktop. It is intended for a trusted LAN or a reverse proxy you control.
It is not a public SaaS authentication system. Anyone with the reachable service URL and current H5 token can access the exposed chat capabilities.
## Enable H5
1. Open **Settings → H5 Access**.
2. Enable H5.
3. Generate a token.
4. Set the LAN access host or public base URL used to generate a launch link and QR code.
5. Choose the current port, or configure a fixed port when you need a stable bookmark or reverse proxy.
Desktop same-origin access automatically allows the origin derived from the configured public base URL. `allowedOrigins` is an advanced local-server API option for custom cross-origin deployments; it is not a field in the Desktop H5 settings page.
For a same-origin LAN deployment, a typical public base URL is:
```text
http://192.168.1.20:3456
```
The generated launch URL contains the server URL and token. Treat it like an API credential and do not post it publicly.
H5 settings are stored in `~/.claude/cc-haha/settings.json`. The full token is persisted so paired devices can reconnect after restart. Only trusted local control surfaces should be able to read it back.
## Enable H5 without the Desktop UI
Build the Web UI, then start the server from the repository root:
```bash
cd desktop
bun run build
cd ..
SERVER_HOST=0.0.0.0 SERVER_PORT=3456 bun run src/server/index.ts
```
The server serves `desktop/dist`. If it is started outside the repository root, set `CLAUDE_H5_DIST_DIR` to the absolute `desktop/dist` path.
From another terminal on the server itself, generate a token:
```bash
curl -sS -X POST http://127.0.0.1:3456/api/h5-access/enable
```
Configure exact origins and an optional public base URL:
```bash
curl -sS -X PUT http://127.0.0.1:3456/api/h5-access \
-H 'Content-Type: application/json' \
--data '{
"allowedOrigins": ["https://cc.example.com"],
"publicBaseUrl": "https://cc.example.com"
}'
```
Read the current configuration and full token only from server loopback:
```bash
curl -sS http://127.0.0.1:3456/api/h5-access
```
If SSH port forwarding is sufficient, keep the service on `127.0.0.1` and use the [headless installation path](./04-installation.md#headless-linux-over-ssh) instead of enabling H5.
## Browser launch URL
For same-origin hosting, the page and API share the public server root:
```text
https://cc.example.com/?serverUrl=https%3A%2F%2Fcc.example.com
```
The generated QR link also includes `h5Token`. On first connection, a manually opened page can ask for:
| Field | Meaning |
|---|---|
| Server URL | Reachable Desktop server or reverse-proxy URL |
| H5 Token | Current token generated by the trusted local settings surface |
The browser stores the Server URL and token in its local storage and sends authorization with REST and WebSocket traffic.
## Reverse proxy requirements
Use HTTPS and proxy all required routes, including:
- `/api/*`
- `/proxy/*`
- `/ws/*` with WebSocket upgrade
- the built Web UI
Preserve the public `Host` and normal proxy information. For example:
```nginx
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
```
Do not rewrite `Host` to loopback while stripping every forwarding header. The server needs enough information to distinguish a remote proxied request from a trusted local request.
## Mobile scope
H5 prioritizes the chat flow:
- session navigation uses a compact drawer;
- composer, stop, attachments, and permission actions remain touch accessible;
- `@` file results adapt to mobile width;
- desktop-only workspace and terminal panels are not the primary mobile surface;
- Desktop pets do not run in H5.
## Security checklist
- Keep H5 disabled unless it is needed.
- Use exact allowed origins; do not use a wildcard.
- Use HTTPS outside a trusted LAN.
- Share the token separately from the public URL.
- Regenerate the token when access should be revoked.
- Do not expose H5 directly to the public internet as if it were a multi-user account service.
Disabling H5 blocks remote access but keeps the current token for a later re-enable. Regenerating the token revokes the old token and QR links immediately.
+75
View File
@@ -0,0 +1,75 @@
---
title: Subagents
nav_title: Subagents
description: When to delegate, which agents ship built in, and how to write your own.
order: 3
---
# Subagents
A subagent is a copy of Claude sent off with one clearly scoped job. It works in its own context and reports back only the conclusion.
The point is that your main conversation doesn't get flooded. Ask "find every call site of `validateUser` in this repo" and, if the main agent searches itself, dozens of files end up in the context window. Delegate it and all that comes back is the list.
## When to delegate
- **Questions that require reading a lot of files** — finding usages, untangling dependencies, counting where a pattern occurs.
- **Independent work that can run in parallel** — frontend, backend, and tests at the same time.
- **The same thing from several angles** — several agents each reviewing the same code.
Conversely: if you already know the file and the line, just say so. No need for the detour.
Delegated agents appear under **SubAgents** in the Activity panel with their tool activity streaming live. Open one to read its full transcript and final result. Background agents work the same way — you don't have to wait for them to finish to see what they're doing.
## Built-in agents
Available without any setup:
| Name | What it's for |
|---|---|
| `general-purpose` | The catch-all. Researching complex questions, searching for code, multi-step tasks |
| `Explore` | Fast codebase exploration. Find files by pattern, search code by keyword, answer "how does this part work" |
| `Plan` | The architect. Designs an implementation strategy and returns step-by-step plans, critical files, and trade-offs |
| `claude-code-guide` | Answers questions about Claude Code, the Agent SDK, and the Claude API themselves |
| `verification` | Sign-off before you call it done. Runs builds, tests, and linters and returns PASS / FAIL / PARTIAL with evidence |
| `statusline-setup` | Configures the Claude Code status line |
You can name one directly ("use Explore to find…") or let Claude decide who to send.
## Seeing what's installed
![Settings → Agents: the agent browser grouped by source](../../images/app/settings-agents.webp)
Open **Settings → Agents**. Three cards at the top show total agents, how many are active, and how many sources are in play. Below that, agents are grouped by source in a fixed order:
**User** → **Project** → **Local** → **Managed** → **Plugin** → **CLI arg** → **Built-in**
When two agents share a name, the higher source wins and the shadowed one is tagged "Overridden by X". Day to day, two groups matter:
- **User** — the ones you wrote, available in every project, stored in `~/.claude/agents/`.
- **Project** — scoped to the current project, stored in its `.claude/agents/`, shipped with the repo.
Click any row for its detail page: model, effort, tool scope, and the full system prompt. Built-in and plugin agents are read-only and show a lock pill instead of edit controls.
## Writing your own
![The Create Agent dialog: scope, model, effort, tools, system prompt](../../images/app/agent-create.webp)
Click **Create Agent** in the top right. The fields:
1. **Scope** — user or project. Choosing project asks you to confirm the target path.
2. **Name** — 1–64 lowercase letters, digits, hyphens, or underscores, e.g. `code-reviewer`. This is what the main agent calls it by.
3. **Description** — when the main agent should delegate to it. **This is the field that matters most**: it's what the main agent reads to decide whether to call this agent at all. Write it vaguely and the agent will never be used.
4. **System prompt** — its responsibilities, boundaries, and expected output.
5. **Model** — inherit from the main agent, or pick Haiku / Sonnet / Opus / Fable, or enter a custom model ID. Simple repetitive work is faster and cheaper on Haiku.
6. **Effort** — inherit, or set low / medium / high / xhigh / max. Models that don't support a level downgrade or ignore it.
7. **Tools** — all tools, no tools, or a custom list. The custom picker groups built-in tools by read and search, modify files, execute commands, and workflow, with a free-text field below for MCP tool names or permission rules like `Bash(git:*)`.
8. **Color** — optional, purely for telling agents apart in the UI.
Saving writes a Markdown file into the matching directory and refreshes the current session. A failed refresh doesn't roll back the file — it still applies after a restart.
:::tip
Grant only the tools the job needs. An agent whose job is to read code and report back doesn't need Write or Bash. Narrow the permissions and you narrow how far it can go wrong.
:::
For the agent file format, source precedence, and inheritance rules, see [Agent internals](../internals/agent.md).
+92
View File
@@ -0,0 +1,92 @@
---
title: Computer Use
nav_title: Computer Use
description: Let Claude read your screen, move the mouse, and type into other apps.
order: 7
---
# Computer Use
With Computer Use enabled, Claude can take screenshots of your screen, move the mouse, click, and type — driving applications that have no API at all: system settings, native note apps, Finder, third-party desktop software.
It acts on this computer, so read what you're authorizing before you turn it on.
macOS and Windows are supported. There is no Linux executor yet.
## Preparing the environment
![Settings → Computer Use: environment checks, the two macOS permissions, authorized apps](../../images/app/settings-computer-use.webp)
Open **Settings → Computer Use**. The top of the page is a row of environment checks:
1. **Python 3** — required on your machine. If it isn't found, click **Download Python 3**. If your Python lives in conda, pyenv, or another custom environment, pick the executable under **Python interpreter path** and it will be preferred from then on.
2. **Virtual environment** and **Dependencies** — click **Install Environment** and the app creates an isolated venv and installs the platform dependencies. Your global Python is left alone.
3. When everything is green the page says all checks passed and Computer Use is ready.
**Re-check** re-runs the detection at any time.
## The two macOS permissions
macOS additionally requires two system permissions. Neither is optional:
| Permission | What it's for |
|---|---|
| Accessibility | Moving the mouse, clicking, typing |
| Screen Recording | Taking screenshots — i.e. letting it see |
The page has **Open accessibility settings** and **Open screen recording settings** buttons that jump straight to the right pane.
:::warning
After granting either one you must **fully quit and reopen the app**. macOS reads these permissions once at process start, so without a restart the page will keep reporting them as not granted.
:::
Make sure you're granting the permission to the app that actually launches Claude Code Haha. Screen Recording detection is occasionally unreliable — if the system settings clearly show it granted but the page still says otherwise, it generally works anyway.
## Pre-authorized apps
By default, every time Claude wants to control a new app it raises a "Computer Use wants to control these apps" prompt naming the apps and the reason. You can **Allow for session** or **Deny**.
For apps you keep approving, tick them under **Authorized Apps** in settings and Claude will control them without prompting. The search box filters your installed apps.
Two more grants are separate and never come along with an app authorization:
- **Clipboard access** — reading and writing the system clipboard.
- **System key combos** — sending system-level shortcuts.
:::danger
Pre-authorization is permanent approval. Keep password managers, banking apps, and corporate chat off that list — make it ask, every time.
:::
## Getting started
Start a session and describe the goal and the allowed apps in plain language. Begin with something small and reversible:
```text
Take a screenshot and tell me what you see.
Open Notes and create an empty note titled "test".
Find the Displays pane in System Settings, but don't change anything.
```
Claude works in a screenshot → decide → act → screenshot loop, so it's slower than you are and will occasionally misclick. Explicit boundaries ("only inside app X", "don't save") work far better than a broad goal.
Only one session can drive the mouse and keyboard at a time. If you see that another session holds it, stop or finish that session first.
## Known limits
- **There is no global abort hotkey.** Use the stop button in the session (`⌘.`).
- **Windows screenshots aren't filtered.** On macOS a screenshot keeps only authorized apps and the desktop; on Windows every visible window is captured. Close or minimize anything sensitive first.
- **Browsers and terminals are restricted.** Browsers are read-only (visible but not clickable) and terminals and IDEs are click-only (no typing). Use the browser extension for web pages and the Bash tool for commands.
- **Re-screenshot after the UI changes.** Old coordinates don't survive a page change.
## Troubleshooting
**The page keeps saying permissions are missing**
Confirm you granted them to the app that actually launches Claude Code Haha, fully quit and reopen, then click **Re-check**.
**The environment won't install**
Pick an explicit Python 3 under **Python interpreter path**, confirm it supports `venv`, and click **Install Environment** again. If it still fails, check the install log in **Settings → Diagnostics**.
**Screenshots work but clicks don't**
Make sure the target app is in the authorized list and is currently in the foreground. Browsers and terminals are subject to the tier restrictions above.
For the permission tiers, the Python bridge, and the executors, see [Computer Use architecture](../internals/computer-use.md).
+36 -24
View File
@@ -1,32 +1,44 @@
# Desktop App
---
title: Desktop feature map
nav_title: Feature map
description: Everything the desktop app can do, on one screen, with a link to each page.
order: 0
---
Claude Code Haha Desktop is the Electron-based workbench for local coding sessions. It combines chat, projects, workspace review, terminals, providers, Skills, Agents, background activity, scheduled tasks, IM adapters, and optional H5 access in one application.
# Desktop feature map
This guide describes the current Desktop experience. For exact changes between releases, see [GitHub Releases](https://github.com/NanmiCoder/cc-haha/releases).
The desktop app puts "talking to Claude" and "seeing what it actually changed" in the same window: projects and history on the left, the conversation in the middle, files, diffs, and a browser preview one click away on the right.
## Start here
This page is a map, not a manual. One line per feature — click through for the details. If you haven't installed it or connected a model yet, start with [Get started](../start/index.md).
- [Quick Start](./01-quick-start.md) — install the app, connect a provider, open a project, and send the first message
- [Feature Guide](./03-features.md) — sessions, workspace tools, Skills, Agents, pets, activity, search, and diagnostics
- [Installation](./04-installation.md) — platform packages, first-launch prompts, Web UI, and headless Linux
- [FAQ](./05-FAQ.md) — provider, data, startup, update, and recovery questions
- [H5 Access](./06-h5-access.md) — optional browser access from a trusted LAN or reverse proxy
- [Desktop Pets](./pets.md) — enable a companion, follow active tasks, customize its appearance, and import your own pet
- [IM Integrations](../im/) — connect WeChat, DingTalk, WhatsApp, Telegram, or Feishu
## Three places to know
## Typical workflow
- **Sidebar** — under the brand seal: New session, Scheduled, Skills Market. Below that, the search bar and your projects and past sessions. Settings sits at the bottom. Drag the edge to resize it.
- **Tab bar** — sessions side by side, switched like browser tabs. The four buttons on the right are Activity, Open project, Open terminal, and Show/hide workspace.
- **Composer toolbar** — left to right: attachments, permission mode, launch location, context usage ring, model and effort, and the run button.
1. Install the package for your operating system.
2. Complete first-run setup and choose where local data should be stored.
3. Sign in with Claude, OpenAI, or Grok, or add a Custom provider.
4. Create a session and select a project directory, branch, or isolated Worktree.
5. Review tool calls, file changes, previews, tasks, and SubAgent activity while the session runs.
6. Use Diagnostics when a provider, local index, sidecar, or desktop runtime needs attention.
## What you'll use daily
## Important boundaries
- [Sessions, permissions, and review](./sessions.md) — starting a session, reading tool cards, answering permission prompts, undoing a turn that went wrong.
- [Workspace](./workspace.md) — the right-hand panel: changed files, diff review, line comments, isolated worktrees, built-in browser.
- [Settings reference](./settings.md) — all 16 tabs, one paragraph each: themes, languages, proxy, terminal, MCP, token usage.
- Desktop data stays local unless a configured provider or integration needs to send it to an external service.
- H5 is not a public SaaS login system. It exposes the local service to anyone who has the configured URL and token.
- IM adapters deny access unless a user is explicitly allowlisted or paired.
- Desktop pets are an Electron feature and do not run in H5.
- Model availability, context size, and reasoning effort still depend on the selected account and provider.
## Getting more work out of it
- [Subagents](./agents.md) — when to delegate, which agents ship built in, how to write your own.
- [Skills and the Skills Market](./skills.md) — what a skill is, how it differs from an agent, what to check before installing one.
- [Scheduled tasks](./schedule.md) — have Claude review yesterday's commits every morning.
- [Computer Use](./computer-use.md) — let it read the screen, move the mouse, and type into other apps.
## Continuing somewhere else
- [Phone (H5) and IM](./remote.md) — scan a QR code to continue the same session in your phone's browser, or chat from WeChat, Feishu, or Telegram.
- [IM integrations](../im/index.md) — full setup steps for each of the five platforms.
## Making it yours
- [Desktop pet](./pets.md) — a little robot that floats on your desktop and shows you how the current task is going. Off by default.
## Curious how it works inside
The pages above cover usage. For architecture and source, start at the [architecture overview](../internals/index.md); the desktop process boundaries are in [Desktop Architecture](../internals/desktop.md).
+75 -158
View File
@@ -1,203 +1,120 @@
# Desktop Pets
---
title: Desktop pet
nav_title: Desktop pet
description: A little robot on your desktop that shows you how the current task is going.
order: 8
---
Desktop Pets are optional companions that live in a small transparent window outside the main app. They reflect a few useful session states, provide a quick way back to active sessions, and can be replaced with your own local artwork.
# Desktop pet
Pets are an **Electron Desktop feature**. They do not appear in H5 or the browser-only Web UI.
A small robot that floats above your desktop and acts out what's happening on this machine — head down and busy while a task runs, glancing around while it waits for your approval, slumped over when something failed. Glance at it while you're doing something else and you'll know whether to switch back.
![Desktop Pet settings with the four built-in companions and appearance controls](../../images/desktop_ui/14_pet_settings_overview.png)
It's a status indicator and a shortcut, nothing more. It can't approve permissions for you, and you can't type into it.
## Enable a pet
## Turning it on
1. Open **Settings → Pets**.
2. Turn on **Show desktop pet**.
3. Choose one of the four built-in companions:
- **Dada**, the coding companion;
- **Huhu**, the planning companion;
- **Bubu**, the fixing companion;
- **Huihui**, the building companion.
4. Adjust the size between `96px` and `192px`.
5. Choose whether to play animations and show the active-task panel.
The pet is **off by default**:
The selected pet, size, window position, animation preference, and task-panel preference are restored between launches.
1. Click **Settings** at the bottom of the sidebar.
2. Choose the **Pets** tab.
3. Pick one from **Built-in pets**.
4. Turn on **Show desktop pet**.
## Interact with the floating pet
![Settings → Pets: four built-in pets and appearance controls](../../images/app/settings-pets.webp)
| Action | Result |
## The four built-in pets
| Character | Who it is |
|---|---|
| Move the pointer near an idle pet | Its gaze follows the pointer; entering the pet also triggers a short jump |
| Click the pet | Bring the main Claude Code Haha window forward and play a wave |
| Drag the pet | Move the floating window; the pet runs in the drag direction |
| Right-click the pet | Open the menu for closing the pet window |
| Click the numbered task badge | Expand the active-session panel |
| Click a session in the panel | Return to that session in the main window |
| Click the arrow below the task list | Collapse the panel back to the numbered badge |
| Dada | A steady coding robot that helps build ideas one block at a time |
| Huhu | A pencil-and-notepad planning robot that maps a way through complex tasks |
| Bubu | A wrench-carrying repair robot with a knack for spotting and fixing cracks |
| Huihui | A gear-carrying build robot that perks up whenever a new reply arrives |
The pet can visibly distinguish working, waiting for you, failed, and idle states. The panel is a navigator, not an approval surface: return to the main app to handle any pending interaction, approve or deny a tool, stop work, or inspect full output.
Switching characters updates an already-open pet window immediately.
When there are no active sessions, the task panel stays hidden. Disabling **Show active task panel** keeps active work behind the numbered badge instead of removing or stopping it.
## Interacting with it
The animation switch affects the pet only. Turning it off does not stop sessions. Claude Code Haha also respects the operating system’s reduced-motion preference.
![The pet floating on the desktop](../../images/app/pet-desktop.webp)
## Create a custom pet
- **Hover** — while idle it hops, and its gaze follows your pointer.
- **Click** — brings the main window forward with a wave. Note that it only raises the window; it doesn't jump into a particular session.
- **Drag** — move it anywhere on the desktop. It tries to return to the same spot next launch.
- **Right-click** — choose **Close pet** from the system menu. That closes the floating window only; no task is stopped. To bring it back, return to **Settings → Pets**.
Select **Add pet** under **Your pets**. The dialog offers three ways to make one. Images are processed entirely on your own computer: nothing is uploaded, and no chat quota is used.
With **Show active task panel** enabled, a panel appears beside the pet whenever work is in flight, grouped by state:
![Custom pet creation dialog with three ways to make a pet](../../images/desktop_ui/15_pet_create_methods.png)
- **Working** — a session, background task, or agent is running.
- **Waiting for you** — something needs a permission approval or other input.
- **Needs attention** — the last run failed.
Before choosing a file, enter:
Clicking a row raises the main window and opens that session. Permissions are still approved in the main window. With the panel off, active work collapses to a numeric badge you can click to expand.
- a Pet ID of at most 73 characters, containing only lowercase letters, numbers, and single hyphens, such as `docs-helper`;
- a display name;
- a short description.
## Appearance controls
The ID must be unique among your custom pets.
- **Pet size** — anywhere between 96 and 192 pixels.
- **Play animations** — turn it off and the pet stays but stops moving. Useful when you want it quieter without dismissing it.
- **Show active task panel** — see above.
- **Collapsed by default** — show just the pet and expand the task panel only when you want it.
### Option 1: use a picture you already have
## Making your own
The quickest route, about a minute. Pick a static image with a transparent background and the app adds breathing, floating and status motion locally.
Click **Add pet** to the right of **Your pets**. Every image is processed on your own machine — nothing is uploaded, and nothing costs conversation credits.
The image must have:
All three routes ask for the same three fields first: a **pet ID** (lowercase letters, digits, and single hyphens, e.g. `moon-cat`), a **display name**, and a **description**.
- PNG or WebP format (APNG and animated WebP are rejected);
- width and height between `32px` and `4096px`;
- no more than `16,777,216` total pixels;
- a file size no larger than `8MB`.
### Route 1: use an image you already have
This kind of pet only sways gently. It **will not run, and it will not track your cursor**. For full motion, use option 2.
The quickest — about a minute. Pick a static PNG or WebP with a transparent background and the app adds gentle breathing and floating motion.
### Option 2: draw one with AI that runs and jumps
This kind of pet **won't run, wave, or track your cursor**. For that, use one of the routes below.
About ten minutes, and you need an AI that can draw. The dialog walks you through the prompt, the reference template and the checks.
### Route 2: have an AI draw an action sheet
#### Step 1: have an AI draw an action sheet
About ten minutes and an image-generating AI. The dialog already contains a complete prompt and a reference template; click **Copy this prompt** and paste it into the AI, replacing the two "character" lines with what you want.
Open any AI that can draw (Nano Banana, ChatGPT image generation, Midjourney, Stable Diffusion and Jimeng all work — these are examples, not endorsements) and send it the whole block below, replacing the two "Character" lines with what you want:
What it needs to draw is an **8-column by 9-row** action sheet, one frame per cell:
```text
Draw me a game character action sheet (sprite sheet).
[Character]
A round-headed orange kitten wearing a small blue scarf, chibi
three-heads-tall proportions, 3D cartoon render, soft glossy
surface, bright cheerful colours.
(Replace these lines with your own character - the more specific the better)
[Whole image]
- Fully transparent background: no backdrop colour, no grid lines, no text, no drop shadow
- Divide the image evenly into 8 columns x 9 rows, 72 equally sized cells
- One action frame per cell, character centred with a little margin around it
- Every cell must show the same character with identical proportions, colours and art style
- Leave unused cells fully transparent
[What to draw in each row]
Row 1, first 6 cells: standing still with a gentle breathing bob
Row 2, all 8 cells: a full run cycle facing right, always facing right
Row 3, first 4 cells: raising a hand and waving hello
Row 4, first 5 cells: crouch, leap, land
Row 5, all 8 cells: dejected and downcast, head lowered, sighing
Row 6, first 6 cells: waiting in place, glancing around
Row 7, first 6 cells: head down, busy working
Row 8, all 8 cells: head and gaze starting straight up, turning slowly to the right through upper-right, right and lower-right, ending near straight down
Row 9, all 8 cells: continuing from straight down, turning left through lower-left, left and upper-left, back to near straight up
```
Getting it wrong on the first try is normal. Ask for a redraw, or say "keep the character, redraw row 2 only".
#### Step 2: check it against the template
![Action sheet template: 8 columns by 9 rows, labelled with what each row should contain](../../images/desktop_ui/17_pet_action_sheet_en.png)
When the picture is ready, check three things before importing:
1. **The background is see-through, not white.** A white backdrop becomes a square on your desktop; this is the most common mistake.
2. **8 cells across, 9 rows down**, one action per cell.
3. **The same character throughout**, with no change of face, colours or proportions.
**Save the template** in the dialog writes this reference image to disk so you can lay out frames against it.
#### Step 3: pick the file
Fill in the ID, name and description, then select the picture. **You do not need to resize anything** — see "Sizes are aligned for you" below.
### Option 3: I already have an action sheet
If you have drawn a sheet already, or hold a finished atlas, this path skips the walkthrough and goes straight to the form. Validation is identical to option 2.
### What to draw in each row
You only draw **nine rows**; the app derives the rest:
| Row | Content | Frames needed |
|-----|---------|---------------|
| 1 | Idle: standing still with a gentle breathing bob | 6 |
| 2 | Run right: full run cycle, always facing right | 8 |
| 3 | Wave: raise a hand and greet | 4 |
| Row | Content | Frames used |
|---|---|---|
| 1 | Idle: standing still with a slight breathing motion | 6 |
| 2 | Run right: a full run cycle, always facing right | 8 |
| 3 | Wave: raising a hand in greeting | 4 |
| 4 | Jump: crouch, leap, land | 5 |
| 5 | Fail: discouraged, head down, sighing | 8 |
| 6 | Wait: looking around, shifting in place | 6 |
| 5 | Fail: dejected, head down, sighing | 8 |
| 6 | Wait: looking around, pacing in place | 6 |
| 7 | Work: head down, busy | 6 |
| 8 | Gaze, upper half: straight up turning clockwise to near straight down | 8 |
| 9 | Gaze, lower half: continuing from straight down back to near straight up | 8 |
| 8 | Gaze, upper arc: from straight up around to nearly straight down | 8 |
| 9 | Gaze, lower arc: from straight down back around to nearly straight up | 8 |
Notes:
"Run left" doesn't need drawing — the app mirrors row 2 horizontally. The last two rows are optional too; repeat the first idle frame and the pet simply won't track your cursor.
- **You do not draw "run left".** The app mirrors row 2 horizontally to produce it.
- **The last two rows are optional in practice.** Repeat the first idle frame and the pet simply will not track your cursor; everything else still works.
- Leave unused cells at the right of each row fully transparent.
When the image comes back, check three things against the template in the dialog: the background is genuinely transparent rather than white, it really is 8 across by 9 down, and it's the same character in all nine rows. If not, ask the AI to redraw — "keep the character, redraw row 2 only" works well.
### Sizes are aligned for you
### Route 3: I already have an action sheet
After you pick the file, the app assembles the runtime atlas locally: it slices the sheet on an 8 × 9 grid, rescales each cell to `192 × 208`, centres the character, mirrors the run row, and fills in the remaining runtime rows to reach `1536 × 2288`.
Skip the tutorial and go straight to the form. The validation rules are identical to route 2.
That means:
### You don't have to compute the size
- **Exact dimensions are not required.** Common AI output sizes such as `1024 × 1152` work; a ratio close to 8:9 gives the best result. The reference size is `1536 × 1872`.
- **A finished `1536 × 2288` atlas is kept byte-for-byte** and is never resampled.
- The assembled file stays under `8MB`; WebP encoding is used automatically if a lossless PNG would exceed it.
Once you pick the image, the app assembles the runtime atlas locally: it slices the 8×9 grid, scales proportionally, centers the character in each cell, mirrors the run-left row, and fills in the remaining rows.
### Common problems
So the **dimensions don't need to be exact**. Common AI output sizes like `1024 × 1152` work fine, and anything close to an 8:9 ratio works best. An existing `1536 × 2288` atlas is kept as-is and never rescaled.
| Message | Cause and fix |
|---------|---------------|
| This image has no transparent background… | The sheet was exported on a white or coloured backdrop. Ask for a transparent PNG, or remove the background with an editor. |
| This image cannot be sliced into 8 columns by 9 rows | The row or column count is off. Confirm 8 across and 9 down, or use a finished `1536 × 2288` atlas. |
| That image could not be read | An animated file (APNG / animated WebP) or a corrupt one. Use a static PNG or WebP. |
| That image is too big | The source exceeds 8 MB. Compress it, or ask for smaller output. |
| A pet with this ID already exists | Choose a different Pet ID, or remove the existing one (see "Storage and removal"). |
The image itself must be: a static PNG or WebP (no animated formats), between 32 and 4096 pixels on each side, at most 16,777,216 pixels total, and under 8 MB.
When the size, format, or image content does not meet these rules, the app refuses to create the pet and shows the matching error instead of adding an invalid entry.
:::warning
The most common failure is a **white background**. A white-backed image becomes a rectangle sitting on your desktop, so the app rejects it outright. Have the AI re-export a transparent PNG, or remove the background yourself.
:::
After a successful import, the new pet is selected automatically and appears under **Your pets**.
## Where they live, and deleting one
![A locally imported custom pet selected in Desktop Pet settings](../../images/desktop_ui/16_pet_custom_result.png)
Custom pet packages live in `${CLAUDE_CONFIG_DIR:-~/.claude}/cc-haha/pets`, and there's an **Open folder** button at the bottom of the settings page. Each pet gets its own subdirectory containing `pet.json` and its images.
## Storage and removal
There's no delete button in the UI yet. To remove one: select a built-in pet first, click **Open folder**, delete only that pet's subdirectory, then click **Refresh** back in settings. Don't delete the whole `pets` or `cc-haha` directory.
Custom pet packages are stored under:
A hand-edited `pet.json` or a swapped image may fail validation; invalid packages are skipped and reported in the settings page.
```text
${CLAUDE_CONFIG_DIR:-~/.claude}/cc-haha/pets
```
## Limits
Use **Open folder** in Pet settings to open the resolved directory. This is also the current removal path:
1. Select a built-in pet first if the package you are removing is active.
2. Open the custom pet folder.
3. Remove only the custom pet’s own package directory.
4. Return to Pet settings and select **Refresh**.
Closing or disabling the floating window does not delete a custom pet. If a selected package goes missing, the app falls back to a built-in pet.
The loader skips invalid or unsafe packages and reports how many folders could not be loaded. Do not replace package files while an import is still running, and avoid symbolic links or unsupported animated image formats.
## Privacy, safety, and boundaries
- Image selection, validation, copying, and lightweight animation are local operations.
- Importing a pet does not send the image to the selected chat model.
- A pet can show summarized task state and navigate to a session, but it cannot approve permissions, answer model questions, or control a task directly.
- Pets do not run in H5, IM integrations, or a browser-only deployment.
- The pet observes the local Desktop service; it is not a cloud monitor and does not stay active after the Desktop app exits.
- Always-on-top behavior, dragging across displays, and window placement can vary by operating system and desktop environment.
- A successful application build does not by itself prove pet-window behavior on every supported operating system.
If the pet does not move, check both **Play animations** and the operating system’s reduced-motion setting. If it does not appear at all, turn **Show desktop pet** off and on once, then restart the Desktop app before editing stored files manually.
The pet is desktop-only and never appears in H5. It stops working when you quit the app or the machine sleeps. Always-on-top behavior, dragging, and multi-monitor placement are all subject to what each operating system allows.
+96
View File
@@ -0,0 +1,96 @@
---
title: Phone (H5) and IM
nav_title: Phone and IM
description: Continue a session in your phone's browser, or chat from WeChat, Feishu, or Telegram.
order: 9
---
# Phone (H5) and IM
A task is running on your computer and you want to check on it, or add one more instruction, from your phone. Two routes:
- **H5 Access** — open the same interface in a mobile browser: sessions, messages, attachments, permission buttons, all of it.
- **IM Adapters** — talk to Claude directly inside WeChat, DingTalk, WhatsApp, Telegram, or Feishu.
Both require your computer to be on with the app running. They expose the local desktop service; they don't move anything to a cloud.
## H5 Access
![Settings → H5 Access: QR code and token](../../images/app/settings-h5.webp)
### Turning it on
1. Open **Settings → H5 Access**.
2. Turn on **Enable H5 access** and confirm after reading the warning.
3. Click **Generate token**. A QR code and an H5 link appear.
4. Scan it with your phone, or click **Copy launch URL** and send it to your own device.
The scanned link carries the server address and the token. Once your phone's browser connects, it remembers the connection and later visits go straight in.
![The mobile conversation view with a file-changes card](../../images/app/h5-session.webp)
### The token is the credential
Anyone holding that link can reach what your desktop exposes. So:
- Never paste a link containing the token into a group chat, a public issue, a log screenshot, or any public page.
- Only enable it on networks you trust. For anything reachable from the internet, put HTTPS, a VPN, or access control in front of it — don't rely on one long-lived token.
- If you suspect a leak, click **Regenerate token**. The old QR code and old token stop working immediately. Turning H5 off and on again does *not* rotate the token.
Turn it off when you're not using it. That immediately rejects remote access while keeping the token, so the same one still works next time.
:::warning
H5 is off by default, and it isn't a public service. Confirm you're on a network you trust before enabling it.
:::
### Access host and fixed port
**Access host / IP** takes your computer's current LAN IP, e.g. `192.168.1.20`, on the current service port. After switching Wi-Fi or unplugging Ethernet the old IP may no longer be yours; the page notices and offers the working one.
If you run your own reverse proxy, put the full URL here instead — `https://cc.example.com` — and add that origin to **Allowed origins**.
A **fixed port** is worth setting when a phone bookmark has to stay valid, a firewall only opens specific ports, or Nginx or Caddy forwards to one fixed upstream. The port must be between 1024 and 65535, and the change takes effect after a restart — the page tells you which port is currently live in the meantime.
### Locking your phone won't kill the task
That's what **Disconnect grace** is for: when your phone locks, backgrounds the tab, or drops off the network briefly, **a running task is not stopped**. It finishes in the background and the result is waiting when you reconnect.
Only when a task is idle *and* nothing is connected does the CLI process stop, after the grace period. The default is 30 seconds and the valid range is 5 seconds to 24 hours. Raise it — say to 600 — if you're operating remotely for a while.
### What works on a phone
Session list and project switching, sending messages, stopping, streaming replies, image and file attachments, permission buttons, questions from Claude, `@` file references, copy and fork — the whole conversation flow.
The desktop workspace, embedded terminal, native "open with", Computer Use authorization, and the desktop pet are not part of H5.
## IM Adapters
![Settings → IM Adapters: pairing and the five platforms](../../images/app/settings-im.webp)
**Settings → IM Adapters** supports five platforms, each connected differently:
| Platform | How to connect |
|---|---|
| WeChat | Generate a QR code in settings and scan it with WeChat to bind the account |
| DingTalk | Scan to create and authorize a bot in one step, or fill in Client ID / Secret manually |
| WhatsApp | Generate a QR code and scan it under **Linked devices** in WhatsApp |
| Telegram | Get a bot token from @BotFather and paste it in |
| Feishu | Enter an App ID and App Secret; if you don't have a bot, the page can create one from a template |
### Binding an account is not the same as allowing a person
This is the step people miss: **scanning only binds the account's messaging capability. Who is allowed to talk to the bot is a separate question.**
Under **Pairing**, click **Generate pairing code**, then send that code to the bot in a direct message from your own IM account. That completes the binding. Alternatively, list user IDs under **Allowed users**.
When both are empty, everyone is denied — that's deliberate, not a bug.
Paired users are listed below and can be unbound at any time; unbinding requires pairing again.
### Other settings
- **Default project** — the working directory for new IM sessions. Left empty, it uses your current user working directory.
- **Streaming card mode** — updates the message content live, so it reads more like watching it type.
- **Permission requests** — DingTalk can use an interactive card template ID for button-based approval. Without it, every platform falls back to the `/allow`, `/always`, and `/deny` text commands.
For each platform's application flow, permission configuration, and troubleshooting, see [IM integrations](../im/index.md).
+56
View File
@@ -0,0 +1,56 @@
---
title: Scheduled tasks
nav_title: Scheduled tasks
description: Run a saved prompt on a schedule — a code review every morning, for example.
order: 5
---
# Scheduled tasks
Save a prompt and have it run on a schedule. The usual jobs: review yesterday's commits every morning, tidy the backlog weekly, check a service's logs hourly.
## Where to start one
Click **Scheduled** in the sidebar to open the task list, then **+ New task** in the top right.
The list page carries a standing notice about the following.
:::warning
Scheduled tasks only run while the desktop app is open and your computer is awake. Close the lid, quit the app, or shut down and nothing fires — and nothing is caught up afterwards. This is not a cloud scheduler.
:::
## Every field in the form
![The New scheduled task dialog: name, description, prompt, frequency, notifications](../../images/app/schedule-create.webp)
- **Name** (required) — the identifier in the list. A hyphenated form like `daily-code-review` reads well.
- **Description** (required) — one line about what it does, so you recognize it later.
- **Prompt** — what actually gets sent to Claude. Give it the full context: what to look at, what to care about, what format to output. You won't be there to clarify.
- **Permissions** — fixed at **Full permissions (fixed)**; the form doesn't offer a choice. The prompt editor shows the directory scope of the run right below. So: **pick the narrowest possible working directory, and read your own prompt once before saving.**
- **Model** — set per task, independent of your session default. Routine checks are fine on a cheaper model.
- **Working directory** — where the task runs. You can also enable **Isolated worktree** so the task works in a separate Git worktree and never touches your main branch.
- **Frequency** — every N minutes, every N hours, daily, weekdays (Mon–Fri), specific days, monthly, or a custom cron expression. Time-of-day options get their own picker beside the frequency. Custom cron is "minute hour day month weekday" and invalid expressions are flagged as you type.
- **Push notification on completion** — pick channels: native desktop notifications and any IM channels you've configured. If no IM channel is set up, the form points you to **Settings → IM Adapters**; setup steps are in [Phone (H5) and IM](./remote.md).
Desktop notifications need **System Notifications** enabled in **Settings → General** and permission granted at the OS level.
Tasks fire with a small random delay so a dozen jobs don't all start on the same tick.
## Once it's running
Each task card offers:
- **Run now** — don't wait for the next trigger. Do this once right after creating a task to confirm the prompt works.
- **Logs** — the execution log, one row per run, with status (running, completed, failed, timeout) and duration. **Summary** shows the output excerpt; **View conversation** jumps to the full session for that run.
- **Edit** — change any field, including the frequency.
- **Disable / Enable** — pause without deleting.
- **Delete** — removes the task and all of its logs permanently.
The three numbers at the top of the list are total tasks, active, and disabled.
## A few habits worth having
1. **Run it manually before setting a frequency.** Prompt problems show up on the very first run.
2. **Keep the working directory narrow.** The task runs with full permissions, so the directory is its boundary.
3. **Don't schedule it too tightly.** A task that runs every minute burns tokens fast and floods the log.
4. **Ask for a conclusion, not a transcript.** Put "summarize in three lines at the end" in the prompt, or the notification will be unreadable.
+100
View File
@@ -0,0 +1,100 @@
---
title: Sessions, permissions, and review
nav_title: Sessions
description: A session from first question to reviewed diff, and which button to press when Claude asks.
order: 1
---
# Sessions, permissions, and review
A session is one complete collaboration: you describe what you want, Claude reads files, runs commands, and edits code, and every step stays in the conversation where you can go back and check it. This page explains what each part of the session screen does.
## Starting a session
Click **New session** in the sidebar, or press `⌘N` (`Ctrl+N` on Windows and Linux). The empty session asks you for exactly one thing: a project directory. After that you can start typing — the model and permission mode come from your defaults in Settings.
Each session opens as a tab, and you can run many side by side. A dot on the tab means that session is still running; closing a running tab asks whether you want to **Keep running** or **Stop and close**.
The small line under the session title is metadata: project path, branch, model. A session is bound to one directory — to work on a different project, start a new session.
## Reading the conversation
![A full session: a question, tool cards, a thinking block, a file edit with an inline diff](../../images/app/session-main.webp)
Claude doesn't just reply with a paragraph. Several kinds of card appear along the way:
- **Tool cards** — reading files, searching, running commands. Consecutive operations of the same kind collapse into one line, like "Read 6 files" or "edited 3 files"; expand it to see the details. Skim them normally, open them when something goes wrong.
- **Thinking blocks** — the reasoning before it acts, labelled **Thinking** while it runs and **Thought** once done. Collapsed by default.
- **File edits** — shown as an inline diff right in the conversation, so you don't have to look anywhere else.
- **Claude needs your input** — when it's genuinely unsure it asks, with buttons for the likely answers plus a free-text box.
In a long conversation, `⌘F` opens find-in-page and jumps between matches in the current session. `⌘K` is global search across every session you've ever had.
## The permission prompt: which button?
![The "Allow Claude to Edit index.html?" prompt with a preview of the change](../../images/app/session-permission.webp)
In the default permission mode, Claude stops and asks before editing a file or running a risky command. The dialog previews the change, then offers three buttons:
- **Allow** — just this once. The same operation will ask again next time.
- **Allow for session** — stop asking for this kind of operation in this session. It resets when the session closes.
- **Deny** — don't run it. Claude gets the refusal and tries another approach.
When in doubt, pick **Allow** — being asked a few extra times costs nothing. If you can't tell what it's about to do, click **Show full input** to see the raw arguments.
### The five permission modes
The permission button in the composer toolbar sets the overall strictness:
| Mode | What it does |
|---|---|
| Ask permissions | Confirm file edits and higher-risk commands when CLI asks |
| Auto accept edits | Claude writes to disk without asking |
| Auto mode | Claude reviews tool calls and runs actions it considers safe |
| Plan mode | Architecture and reasoning only, no files |
| Bypass permissions | Full tool access for shell and file system |
**Auto mode** and **Bypass permissions** each require a one-time confirmation. In Plan mode Claude produces a plan without touching files; when it's done you get a "Ready to code?" prompt where you can approve the plan or send it back for changes.
The permission mode is locked while a turn is running and unlocks when the turn finishes.
:::warning
**Bypass permissions** hands over your shell and your entire file system. Use it only in an isolated environment you can restore.
:::
## Undoing a turn
After each turn, a card appears in the conversation reading "**{n} files changed**", listing every file that turn touched. It offers two actions:
- **Undo current turn** — roll back the latest reply and restore the files it changed.
- **Roll back to before this turn** — for older turns: rewind both the conversation and the files to that checkpoint.
Both ask for confirmation first. Some turns have no file checkpoint; in that case only the conversation rolls back, and the message says so.
## The Activity panel
The first button on the right of the tab bar opens the Activity panel, which lists everything running in parallel for this session:
- **Tasks** — the to-do list Claude maintains for itself, with "Task progress 3/7" at the top.
- **SubAgents** — the agents it delegated to. Open one to read its full transcript.
- **Background tasks** — commands and workflows running in the background; each can be stopped individually.
- **Team** — when an Agent Team is in play, one row per member, and you can message a member directly.
Tool activity from background subagents bubbles up here too, so you don't have to wait for one to finish to see what it's doing.
## What the composer can do
![The slash-command panel that opens when you type `/`](../../images/app/composer-slash.webp)
- **`/` slash commands** — type `/` for the command panel. `/status` for session state and usage, `/context` for context breakdown, `/compact` to compress, `/review` to review changes, `/commit`, `/memory` to open project memory, `/doctor` to open the diagnostics check.
- **`@` file references** — type `@` for file search; the file you pick is attached to the message as a path.
- **Attachments** — click `+`, drag files in, or paste a screenshot. Images, PDFs, and directories all work.
- **Context usage ring** — the small ring shows how much of the context window is used; hover it for used, free, and window size. When it fills up, run `/compact`.
- **Model and effort** — switch models at any time. Effort has five levels — low, medium, high, xhigh, max — and models that don't support a level ignore it.
- **Location** — shows the current project and branch. In a Git project you can switch branches here, or turn on **Isolated worktree** to keep an experiment off your main branch. See [Workspace](./workspace.md).
Enter sends and Shift+Enter inserts a newline by default; **Settings → General** can swap that to `Ctrl/Cmd+Enter`. `⌘.` stops the current generation.
## Forking a conversation
Every past message has **Fork a new conversation**. It branches a new session from that point: everything before it is kept, everything after is up for grabs. Use it when you want to try a different approach without losing the thread you already have.
+127
View File
@@ -0,0 +1,127 @@
---
title: Settings reference
nav_title: Settings
description: All 16 settings tabs — what each one configures and when you'd need it.
order: 6
---
# Settings reference
Click **Settings** at the bottom of the sidebar. Sixteen tabs on the left, in a fixed order. This page walks through them in that order: what each one configures, and when you'd actually need to touch it.
## Providers
Model access. Sign in to Claude, ChatGPT, or Grok with an account (no API key required), or add any Anthropic- or OpenAI-compatible service with an API key.
You'll come here once during setup and rarely again. Full steps in [Connecting a model](../start/models.md).
## General
The tab you'll open most often — everything about how the app feels.
![Settings → General: color themes, language, output style, default permissions](../../images/app/settings-general.webp)
- **Color theme** — six of them: Pure White (default), Paper, Warm Classic, Celadon, Ink Night, Ink Blue. There's also **Follow the system**, which lets you pick which theme to use in light mode and which in dark mode.
- **Language** — the interface language.
- **Response Language** — makes Claude always reply in a given language, set independently of the interface.
- **Output Style** — Default, Explanatory, or Learning. Explanatory adds reasoning about implementation choices and codebase patterns; Learning pauses and asks you to write small pieces yourself. Running sessions keep their current prompt; the change applies to new sessions.
- **Default Session Permissions** — which permission mode new sessions start in. Each session can still be changed individually.
- **Effort Level** and **Thinking Mode** — defaults for new sessions. Turning thinking off sends an explicit non-thinking parameter to providers like DeepSeek that need one.
- **Message Sending** — Enter to send (Shift+Enter for a newline), or `Ctrl/Cmd+Enter` to send.
- **System Notifications** — route permission prompts, completed replies, and scheduled task results to the OS notification center. Enabling it requests system permission.
- **Network** — three modes: Direct connection (explicitly bypass the system proxy), System proxy (follow system or PAC rules per destination), or Manual proxy (a URL like `http://user:password@127.0.0.1:7890`). Below that, **AI request timeout**, which can go up to 1800 seconds when a provider is slow to first byte. App updates use their own proxy setting, over in About.
- **WebSearch** — how web search is routed. Auto prefers Claude's native WebSearch for Claude models and falls back to Tavily or Brave otherwise; those two need API keys you supply.
- **Auto-dream** — periodically tidies and compresses memory files in the background. Off by default, because it spends tokens.
- **UI Zoom** — scale the whole interface, also bound to `⌘+` / `⌘-`, with `⌘0` back to 100%.
- **Data Storage Location** — an advanced, rarely-touched setting. Defaults to the system directory `~/.claude`, or point it at an absolute path of your own. After switching, sessions, skills, MCP, plugins, and provider config are all read from the new directory; it needs a restart, and the two directories are never merged or migrated automatically.
![The same session in the Ink Night dark theme](../../images/app/session-dark.webp)
## H5 Access
Continue the same session in your phone's browser. Off by default. See [Phone (H5) and IM](./remote.md).
## IM Adapters
Talk to Claude from WeChat, DingTalk, WhatsApp, Telegram, or Feishu, and manage paired users. See [Phone (H5) and IM](./remote.md) and [IM integrations](../im/index.md).
## Terminal
A real host shell embedded in the app, for installing plugins, skills, MCP servers, and anything else that needs a command line. The desktop app bundles `claude-haha`, so anywhere the docs say `claude <args>` you can run `claude-haha <args>`.
On Windows you can also choose the startup shell (system default, PowerShell 7, Windows PowerShell, Command Prompt, or a custom executable) and set a Bash path — used when a tool calls Unix commands like `grep` or `sed`, usually pointing at Git Bash.
## MCP
External tools and data sources. STDIO, Streamable HTTP, and SSE transports are supported, and the scopes match the CLI:
- **Local** — only for you, but bound to one project.
- **Project** — written to the project's `.mcp.json` and shared with the team.
- **User** — written to your global config, active in every project.
The three numbers at the top are total servers, currently connected, and needs attention. STDIO commands run directly on your machine, so runtimes like Node, Python, and Bun must be installed and on your `PATH`.
## Agents
Browse installed agents and create your own. See [Subagents](./agents.md).
## Skills
Every skill available on this machine, grouped by source, with the prose and source files readable in place. See [Skills and the Skills Market](./skills.md).
## Memory
View and edit the Markdown memory files Claude keeps per project. Pick a project on the left, a file in the middle, then edit or preview the rendered output. The files live in `~/.claude/projects/<project>/memory/` and are loaded by the CLI at runtime.
`/memory` in any session jumps straight here. For how memory is written and recalled, see [Memory system](../internals/memory.md).
## Plugins
A plugin bundles skills, agents, hooks, and MCP servers together. This tab shows installed plugins, their health, and what capabilities each one exposes, with enable, disable, update, and uninstall — including multi-select for bulk operations.
After enabling or disabling anything, click **Apply changes** to push the change into the running runtime.
## Pets
A little robot floating on your desktop. Off by default. See [Desktop pet](./pets.md).
## Computer Use
Let Claude read the screen, click, and type. Unusable until you install the runtime environment and grant system permissions. See [Computer Use](./computer-use.md).
## Token usage
![Settings → Token usage: heatmap and stat cards](../../images/app/settings-usage.webp)
A usage dashboard computed from the Claude Code session records on this machine. Everything is calculated locally and nothing is uploaded.
- Across the top: total tokens, peak tokens, longest task, current and longest streak.
- In the middle: a heatmap with daily, weekly, and cumulative views. Click a day for that day's sessions, tokens, messages, and tool calls.
- Below: activity insights — active rate, most-used model, skills used, fresh versus cache-hit token split, and estimated cost. Estimated cost excludes models with no known price, and the UI says how many were skipped.
## Trace
Records the model request chain for each session — requests, responses, status events, timings — for debugging stalls, failures, and unexplained waits. The switch isn't on this tab: turn on **Agent trace** in **Settings → General** first, and this tab fills up.
Once enabled, new sessions write condensed records to a local traces directory. Existing records stay readable after you turn it off; only new ones stop being written. The trace list supports search, filtering (all / LLM / tools / errors), opening a trace in its own window, and deleting a session's trace without touching its chat history.
## Diagnostics
Where to go when something breaks. Logs server and CLI startup, provider, and session runtime errors.
- Across the top: log size, event count, warnings in the last 24 hours, retention policy.
- **Recent events** lists the actual errors, each with an event ID you can copy on its own.
- **Export Bundle** / **Copy error summary** / **Copy issue report** — use the latter two when filing an issue; they're already formatted.
- **Doctor** — checks user and current-project configuration state, read-only, returning a healthy / not configured / missing / invalid list. `/doctor` in a session opens it directly.
- **Reset safe UI state** — clears only regenerable interface keys like open tabs, theme, and zoom. Chat history, model config, skills, MCP, IM, and OAuth are always protected and never touched.
- **Local index** — status and size of the derived SQLite index, with a rebuild button. Rebuilding only affects the index; source conversations are never deleted.
:::info
Issue reports and exported bundles are redacted on a best-effort basis — chat contents, file contents, full environment variables, and API keys are omitted. Still give them a read before sharing, in case an internal hostname, username, or path slipped through.
:::
## About
Version, changelog, GitHub repo, feedback link.
**App Updates** checks GitHub Releases, downloads, and restarts to install. Updates use their own proxy setting, separate from **Settings → General** — if updates stall on a corporate network, configure the advanced update proxy here.
+76
View File
@@ -0,0 +1,76 @@
---
title: Skills and the Skills Market
nav_title: Skills
description: A skill is a ready-made procedure for Claude. What to check before installing one.
order: 4
---
# Skills and the Skills Market
A skill is a procedure someone else already worked out. Install it and Claude follows it whenever the matching situation comes up. A "PDF handling" skill, for instance, tells it which library to use, what order to split pages in, and what to do with an encrypted file — so you don't have to explain it every time.
**How it differs from an agent**: an agent is a worker with its own context and tool scope; a skill is knowledge and procedure that an agent loads. Delegating to an agent is hiring someone; installing a skill is handing them a manual.
## The Skills Market
![The Skills Market: cards, security badges, source filters](../../images/app/skill-market.webp)
Click **Skills Market** in the sidebar. It aggregates two sources, **ClawHub** and **SkillHub**, and loads more as you scroll — there is no "load more" button.
Three filters across the top:
- **Source** — all sources, ClawHub, or SkillHub.
- **Security** — see below.
- **Install status** — all skills, installed, or not installed.
The search box matches names and keywords.
### Reading the security badges
Each card carries a security badge. It reports what the **source** scanned, not an audit by this app:
| Badge | Meaning |
|---|---|
| Verified | Publisher is verified and the skill passed the source's security scan |
| Scanned safe | The source's security scan found no risks |
| Not audited | The source provided no security audit data |
| Flagged | The source's scan flagged potential risk |
"Not audited" doesn't mean unsafe — it means nobody checked. "Flagged" means don't install it unless you've read the files and understand exactly what they do.
## Before you install
A skill can bring new tools, scripts, and external dependencies, and once installed Claude may run them. The disclaimer at the top of the market means what it says: these skills come from third-party community sources and this app does not audit their contents.
Recommended:
1. Open the **Files** tab on the detail page and read `SKILL.md` and any accompanying scripts.
2. If you're unsure, have Claude look first — "check this skill's files for anything suspicious".
3. Confirm you actually need it. More skills isn't more capability; every skill costs context.
**Install** opens a confirmation with the install location, the security note, and a reminder that the skill takes effect in new sessions. **Sessions already open won't pick it up** — start a new one.
Uninstall lives in the same place and deletes the local files under that skill's directory.
## Installed skills
**Settings → Skills** lists everything available on this machine, grouped by source:
- **User** — installed by you, in `~/.claude/skills/`.
- **Project** — shipped with the repo.
- **Plugin** — bundled by a plugin.
- **Built-in** — shipped with the app.
Each entry shows its entry file, file count, and estimated token cost. Open one to switch between **doc mode** and **code mode** and read the skill's prose and source files directly.
### The `.agents/skills` convention
Besides `~/.claude/skills/`, the desktop app also reads `~/.agents/skills/`. That's an open standard directory shared with Codex, Cursor, Gemini CLI, and other clients — install once, use it everywhere. Skills from that directory are tagged `.agents` in the list.
A project's own `.agents/skills/` works the same way: it ships with the repo rather than being something you installed.
:::warning
A shared directory means skills another client installed also apply here. Scan **Settings → Skills** occasionally and make sure nothing on the list is a stranger.
:::
For how skills are loaded and what the file format is, see [Skills internals](../internals/skills.md).
+77
View File
@@ -0,0 +1,77 @@
---
title: Workspace
nav_title: Workspace
description: The right-hand panel — see what changed, review the diff line by line, preview a page in-app.
order: 2
---
# Workspace
The conversation tells you what Claude said. The workspace tells you what it actually changed. It's a panel on the right that you can pull out next to the conversation, so you never have to switch windows.
## Opening it
Click the folder icon on the right of the tab bar. Click it again to collapse. Drag the panel's left edge to resize.
At the top of the panel is a **Files / Browser** switch:
- **Files** — project files, Git changes, and diff review.
- **Browser** — a built-in browser for previewing the page you just changed.
## Changed files and All files
![The "Changed files" list, each row with a status marker and line counts](../../images/app/workspace-changes.webp)
In Files mode there are two views:
- **Changed files** — only files with uncommitted changes in this Git repo, each row showing its status (modified, added, deleted, renamed, untracked) and lines added or removed. This is where you'll spend review time.
- **All files** — the full directory tree. The search box above it matches file names across the whole project, including directories you haven't expanded yet.
Click a file name to open a preview: a diff in two columns, or the file itself. Previews accumulate as tabs, like an editor.
When the directory isn't a Git repo, **Changed files** says so plainly — that isn't an error.
Any file can have its path copied, or be pushed back into the composer as context with **Add to chat**.
## Diff review: leaving a note on a line
![Diff review with syntax highlighting, old and new side by side](../../images/app/workspace-diff.webp)
The diff keeps old and new lines with full syntax highlighting. The genuinely useful part is line-level comments:
1. Click a line — a comment box opens beside it.
2. To comment on a range, hold `Shift` and click the first and last line. The selection must stay on one side of the diff and inside one hunk.
3. Describe what you want changed there and click **Submit**.
4. The comment goes back into the composer along with the file path, line numbers, and the code itself, ready to send.
This is far more precise than describing "the null check in that one function" in prose.
If Claude changes the file while you're writing a comment, the panel tells you the diff has updated and asks you to reselect — that's there to stop a comment from landing on the wrong lines.
:::tip
Denied tool calls never reach disk, so they never show up in Changed files. Even so, give `git diff` one last read before you ship.
:::
## Isolated worktree: keeping experiments caged
The **Location** control in the composer offers **Isolated worktree**. Turn it on and the session gets its own Git worktree; everything Claude does happens there, and your current branch and working directory are untouched.
When it earns its keep:
- You want Claude to attempt a big rewrite you may not keep.
- Your working directory has uncommitted changes, so Git would block a branch switch.
- The branch you want is already checked out in another worktree.
The temporary worktree is cleaned up when you're done. History stays readable, but to continue you'll need a new session in the original project — the app tells you when that's the case.
## Built-in browser
![The built-in browser previewing a page that was just edited](../../images/app/workspace-preview.webp)
Switch the workspace panel to **Browser** and type a local dev address or any URL. Three buttons here exist specifically so Claude can see what you see:
- **Capture** — send the current rendering back into the conversation.
- **Pick element** — click an element on the page; its selector, position, and a screenshot go to Claude as context. This saves an enormous amount of back-and-forth on styling.
- **Zoom** — change the preview scale to check responsive layouts.
Logins and cookies in this browser are real, same as any browser. Before you demo or screenshot anything publicly, switch to a page that doesn't require signing in.
-188
View File
@@ -1,188 +0,0 @@
# Computer Use Guide
Computer Use lets a model inspect the screen and operate the mouse, keyboard, applications, and clipboard. It acts on the current computer, so review the authorization scope and platform differences before enabling it.
## Supported platforms
| Platform | Status | Important difference |
|---|---|---|
| macOS Apple Silicon / Intel | Supported | Requires Accessibility and Screen Recording; supports native screenshot filtering |
| Windows | Supported | Uses the Windows Python runtime; screenshots are not filtered to authorized windows |
| Linux | Unsupported | There is currently no Linux executor |
The packaged desktop app does not require Bun. Computer Use does require a working Python 3 installation. The app creates an isolated virtual environment under the managed user configuration directory and installs the platform dependencies there.
## Quick start
### Desktop app
1. Open Settings → Computer Use.
2. Enable Computer Use.
3. Check Python status. If automatic discovery fails, select a Python executable.
4. Run Install or Repair so the app can create the virtual environment and install dependencies.
5. On macOS, grant Accessibility and Screen Recording, then check again.
6. Choose the applications that may be controlled and, if needed, enable clipboard and system key permissions.
7. Start a session and describe both the goal and the applications involved.
Start with a small, reversible task:
```text
Take a screenshot and tell me what you can see.
Open Notes and create an empty note titled "Test".
Find the settings entry in the authorized app, but do not change anything.
```
### CLI
Source mode requires the project dependencies and Python 3:
```bash
bun install
python3 --version
./bin/claude-haha
```
Disable the dynamic Computer Use MCP with:
```bash
CLAUDE_COMPUTER_USE_ENABLED=0 ./bin/claude-haha
```
Or set the managed configuration in `~/.claude/cc-haha/computer-use-config.json`:
```json
{
"enabled": false
}
```
The desktop Settings page updates the same managed configuration. Prefer the UI and do not overwrite unknown fields by hand.
## Tools and Teach capability
The codebase defines 27 Computer Use tools:
| Category | Tools |
|---|---|
| Authorization | `request_access`, `list_granted_applications` |
| Screenshot | `screenshot`, `zoom` |
| Mouse | `left_click`, `right_click`, `middle_click`, `double_click`, `triple_click`, `left_click_drag`, `mouse_move`, `left_mouse_down`, `left_mouse_up`, `cursor_position`, `scroll` |
| Keyboard | `type`, `key`, `hold_key` |
| Applications | `open_application`, `switch_display` |
| Clipboard | `read_clipboard`, `write_clipboard` |
| Control flow | `wait`, `computer_batch` |
| Teach | `request_teach_access`, `teach_step`, `teach_batch` |
There are 24 base control tools. The three Teach tools are exposed only when the host enables the Teach capability; otherwise a session sees only the base tools.
Teach is intended for requests where the user wants to learn a workflow:
1. `request_teach_access` requests access to the applications used by the tour.
2. `teach_step` shows an anchored explanation and waits for the user to choose Next.
3. `teach_batch` groups predictable steps to reduce model round trips.
Teach authorization is separate from regular control authorization. Actions still pass the application allowlist and input safety checks. After the user exits a tour, the model must stop calling Teach tools.
## How it works
Computer Use follows a screenshot → analyze → act → screenshot loop:
```text
Model
→ Computer Use MCP tool
→ TypeScript dispatch and safety checks
→ Python bridge
→ macOS / Windows system operation
→ Screenshot or action result returned to the model
```
- Tool definitions and authorization logic live in `src/vendor/computer-use-mcp/`.
- CLI integration and the Python bridge live in `src/utils/computerUse/`.
- Platform executors live in `runtime/mac_helper.py` and `runtime/win_helper.py`.
- Desktop setup, permissions, and preauthorization are managed by `src/server/api/computer-use.ts` and `desktop/src/pages/ComputerUseSettings.tsx`.
## Authorization model
### Application access
The model first calls `request_access` with the required applications and a reason. The user can approve or deny the request. Applications preauthorized in Settings provide managed defaults; they do not grant arbitrary access to every application. Adding an application during a run still follows the relevant authorization flow.
Applications use three permission tiers:
| Tier | Capability |
|---|---|
| `read` | Inspect screenshots without input |
| `click` | Click, move, and scroll without typing or higher-risk input |
| `full` | Allow keyboard, drag, and other complete input after the remaining checks pass |
### Clipboard and system keys
Clipboard read, clipboard write, and system-level key combinations are separate grants. Authorizing an application does not automatically authorize these capabilities.
### Concurrency
A session lock prevents multiple sessions from competing for the mouse and keyboard. If Computer Use reports that another session owns the lock, finish or stop that session instead of deleting the lock file.
## Platform safety boundaries
### macOS
- Accessibility is required for application input.
- Screen Recording is required for screenshots.
- Screenshots support native window filtering, leaving authorized applications and the desktop visible.
### Windows
- The current screenshot filtering capability is `none`, so screenshots may include every visible window.
- The application allowlist still rejects input aimed at an unauthorized frontmost application.
- Because screenshot content is not filtered, close or minimize windows containing sensitive information before starting.
### Protections that are not currently available
- **There is no global Escape abort hotkey.** Use the current task's Stop action in the desktop app; a CLI run can be interrupted from its terminal.
- **Unauthorized windows are not automatically hidden before every action.** Do not rely on auto-hide for privacy.
- **Pixel staleness validation is disabled by default.** After the UI changes, the model should take another screenshot before clicking.
These are current implementation boundaries and must not be described as completed safeguards.
## Python runtime
During first install or repair, the app:
1. Synchronizes the helper and requirements for the current platform into the managed user directory.
2. Creates a venv with the detected or selected Python.
3. Installs or upgrades pip.
4. Uses the requirements content hash to decide whether dependencies need reinstalling.
5. Calls the platform helper with a JSON payload and parses a uniform JSON response.
macOS primarily uses `mss`, Pillow, PyAutoGUI, and PyObjC. Windows also uses pywin32, psutil, pyperclip, and screeninfo. See `runtime/requirements*.txt` for the current version constraints.
## Troubleshooting
### macOS still reports missing permissions
- Make sure the authorization applies to the app that actually launches Claude Code Haha.
- Fully quit and reopen the app after changing permissions.
- Run the permission check again from Settings.
### Python installation fails
- Select a specific Python 3 executable in Settings.
- Confirm that Python supports `venv`.
- Run Install or Repair again.
- Review the Computer Use installation log in Diagnostics.
### Screenshots work but clicks fail
- Confirm that the target application is authorized.
- Confirm that it is the frontmost application.
- Check whether the permission tier allows the requested action.
- Take a new screenshot after any UI change instead of reusing old coordinates.
### Other windows appear in Windows screenshots
This is a known boundary of the current Windows screenshot capability. The input allowlist does not filter screenshot content. Close or minimize sensitive windows first.
## Learn more
- [Computer Use architecture](./computer-use-architecture.md)
-153
View File
@@ -1,153 +0,0 @@
# FAQ and Support
## What is the fastest way to get help?
First, confirm that you are using [GitHub Releases Latest](https://github.com/NanmiCoder/cc-haha/releases/latest), then retry the shortest sequence that reproduces the problem reliably.
If the Desktop app still opens:
1. Go to **Settings → Diagnostics**.
2. Select **Copy Issue Report**.
3. Search [GitHub Issues](https://github.com/NanmiCoder/cc-haha/issues) for the same problem, and create a new issue only if there is no duplicate.
4. Paste the report and add the reproduction steps, expected result, and actual result.
If the report is not enough, you can also select **Export Bundle**. Issue reports and bundles attempt to omit chat content, file content, complete environment variables, and API keys. They can still contain private metadata such as local paths or provider hostnames, so **review everything before sharing it**.
Include:
- The Claude Code Haha version
- Operating system, CPU architecture, and installer type
- Provider type and API format, but never the API key
- The shortest reproduction steps and complete error text
- Whether the problem occurs in Desktop, H5, or the CLI
If the app cannot start, provide the installer version, a screenshot of the system error, and the last visible action before the failure. Do not delete or overwrite `~/.claude` as a troubleshooting shortcut.
## Providers and OAuth
### Do OpenAI-compatible services always require LiteLLM?
No. A Custom Provider in the Desktop app supports:
- Anthropic Messages (native)
- OpenAI Chat Completions (local proxy translation)
- OpenAI Responses API (local proxy translation)
Consider an external gateway such as LiteLLM only when the upstream service uses a protocol or custom fields that the app does not support.
### A custom provider returns 401 or “API Key invalid”
Edit the provider under **Settings → Providers**, then check:
1. The base URL is the API root required by the provider, without a missing or duplicated path.
2. The API format matches the real upstream endpoint.
3. The authentication strategy is correct. Some Anthropic-compatible services use a Bearer token, while the official Anthropic API uses `x-api-key`.
4. The main model ID exists and the current credential can access it.
5. Select **Test Connection**. Resolve the first connectivity failure before the second proxy-translation result shown for OpenAI formats.
Do not expose URL query credentials, tokens, or API keys in screenshots, issues, or diagnostic attachments.
### Claude, ChatGPT, or Grok sign-in does not complete
- Keep Claude Code Haha running while the sign-in flow is open.
- Allow the system browser to open the authorization page and complete authorization with the intended account.
- Confirm that the system clock is correct and that a proxy, firewall, or browser extension is not blocking the provider page or local OAuth callback.
- If the browser reports success but the app does not update, return to **Settings → Providers** and start the sign-in flow again.
If it still fails, copy an Issue Report and state whether the browser did not open, the authorization page failed, or the callback did not return to the app. Never share an OAuth code, access token, or browser cookie.
### The connection test succeeds, but chat still fails
Confirm that the current task selected the new Provider and model; saving a Provider does not necessarily select it for every task. Also check that the account supports the configured main and role-model mappings.
Create a new task and retry with a short plain-text message. If it still fails, copy an Issue Report from Diagnostics. A successful connection test proves basic endpoint, authentication, and translation behavior; it does not verify every model capability, tool call, or long-context combination.
## Installation and Updates
### macOS says the app cannot be verified, is damaged, or will not open
Confirm that the installer came from the [official Latest Release](https://github.com/NanmiCoder/cc-haha/releases/latest) and that you selected the correct Intel or Apple Silicon architecture. A signed and notarized public release should normally show only the standard download-source confirmation.
Older, draft, or temporary unsigned builds may need extra approval. See the [Desktop installation guide](/en/desktop/04-installation#macos). Do not bypass system security prompts for packages from unknown mirrors.
### Windows shows SmartScreen
Confirm that the installer came from the official Release. An unsigned Windows installer may trigger SmartScreen. Expand “More info,” verify the file name and source, and then decide whether to run it.
Quit Claude Code Haha completely before an in-place update. If the installer reports that a process is still running, close the relevant window and background process before retrying. Do not delete the user configuration directory first.
### A Linux AppImage does not start
Make it executable:
```bash
chmod +x Claude-Code-Haha-<version>-linux-<architecture>.AppImage
```
Some distributions also require FUSE. See the [Desktop installation guide](/en/desktop/04-installation#linux) for distribution-specific instructions.
## H5 Access
### A phone or another computer cannot connect to H5
Check each item in order:
1. **H5 Access** is enabled in the Desktop app.
2. You are using the Server URL currently shown in Settings or encoded in the current QR code, not an old LAN address.
3. The client has the current H5 token.
4. Both devices are on networks that can reach each other, and the system firewall allows the configured port.
5. A reverse proxy forwards both HTTP and WebSocket traffic and uses an allowed origin.
H5 is a remote entry point to the current Desktop service, not the complete Desktop environment. Terminals, native previews, pet windows, and some system capabilities remain Desktop-only.
See [H5 access](/en/desktop/06-h5-access) for deployment and security details.
## Git Branches and Worktrees
### Creating an isolated worktree fails
Common causes include:
- The selected directory is not a Git repository
- The branch does not exist or is already checked out by another worktree
- Uncommitted changes prevent the requested Git operation from running safely
- The target worktree path already exists or is not writable
Read the exact error shown in the app. You can use the current working tree, select a different branch, or handle the existing Git changes safely before retrying. Do not automatically delete an existing directory or discard uncommitted changes to make a worktree.
### Should I use the current working tree or an isolated worktree?
- **Current working tree**: continue work that already exists in the selected directory.
- **Isolated worktree**: run a parallel task or keep a branch separate from the current checkout.
If the current directory contains important uncommitted changes, inspect the Git status and make a backup before choosing.
## Computer Use
### Computer Use is unavailable or cannot control an app
Open **Settings → Computer Use** and confirm:
- The global feature switch is enabled
- Python and dependency checks pass
- Required macOS or Windows system permissions are granted
- The target app is in the list of apps allowed for control
- The current session's permission request was explicitly approved
After granting a system permission, reopen Claude Code Haha or the target app. Computer Use is not currently supported on Linux. A successful Desktop installation does not mean that Computer Use is fully configured.
If Settings still shows an error, copy an Issue Report and include the status shown on that page. Avoid uploading a full-screen screenshot that exposes content from other apps.
## CLI
### `bun install` or the CLI fails to start
Confirm that the shell is in the repository root and that Bun is current enough for the project:
```bash
bun --version
bun install
./bin/claude-haha --help
```
If the error mentions a missing Bun built-in such as `bun:bundle`, upgrade Bun. See [Get Started in 3 Minutes](./quick-start.md#option-2-run-the-cli-from-source) for installation and the [CLI reference](./cli-reference.md) for commands.
-58
View File
@@ -1,58 +0,0 @@
# Global Usage (Run from Any Directory)
If you want to run `claude-haha` directly from any project directory, set up one of the following. Once configured, `claude-haha` will automatically recognize your current working directory.
## macOS / Linux
Add to `~/.bashrc` or `~/.zshrc`:
```bash
# Option 1: Add to PATH (recommended)
export PATH="$HOME/path/to/claude-code-haha/bin:$PATH"
# Option 2: Alias
alias claude-haha="$HOME/path/to/claude-code-haha/bin/claude-haha"
```
Then reload the config:
```bash
source ~/.bashrc # or source ~/.zshrc
```
## Windows (Git Bash)
Add to `~/.bashrc`:
```bash
export PATH="$HOME/path/to/claude-code-haha/bin:$PATH"
```
### Windows + WSL Toolchains
If `claude-haha` runs on Windows / Git Bash but tools such as Node, Python, uv, or bun are installed inside WSL, call them through WSL explicitly:
```bash
wsl -e bash -lc 'node --version && python3 --version'
```
When cc-haha detects `wsl` / `wsl.exe`, it automatically sets `MSYS2_ARG_CONV_EXCL=*` so Git Bash does not rewrite WSL paths such as `/home/...` into `C:/Program Files/Git/home/...`.
To route Bash tool commands through WSL by default, set this before startup:
```bash
export CLAUDE_CODE_SHELL_PREFIX='wsl -e bash -lc'
```
Computer Use still controls Windows desktop apps. CLI tools running inside WSL do not need to be added to `computer-use-config.json`. If you only need the WSL toolchain and do not need desktop control, disable Computer Use with `--no-computer-use` or the Settings > Computer Use switch.
## Verify
After setup, navigate to any project directory and test:
```bash
cd ~/your-other-project
claude-haha
# Ask "What is the current directory?" — it should show ~/your-other-project
```
-109
View File
@@ -1,109 +0,0 @@
# Get Started in 3 Minutes
Claude Code Haha has two ways to run. Install the **Desktop app** for everyday development. Run the **CLI from source** only when you need a terminal workflow, scripting, or local development.
| Option | Best for | What you need |
|--------|----------|---------------|
| **Desktop (recommended)** | Managing projects, sessions, worktrees, code diffs, and permission reviews | The installer for your platform; **Bun is not required** |
| **CLI** | Terminal users, `--print` automation, and contributors | Git, Bun, and access to a model provider |
## Option 1: Install the Desktop App
### 1. Download the current stable release
Open [GitHub Releases Latest](https://github.com/NanmiCoder/cc-haha/releases/latest), then select the installer for your operating system and CPU:
- macOS on Apple Silicon (M-series): `mac-arm64.dmg`
- macOS on Intel: `mac-x64.dmg`
- Windows: `win-x64.exe` or `win-arm64.exe`
- Linux: the matching `.AppImage` or `.deb`
See the [Desktop installation guide](/en/desktop/04-installation) for platform-specific installation prompts.
### 2. Configure a model provider
Launch the app and open **Settings → Providers**.
The shortest path is to select an official option and follow its on-screen connection or sign-in flow:
- **Claude Official**
- **ChatGPT Official**: sign in with a ChatGPT account through OAuth
- **Grok Official**: sign in with an xAI account through OAuth
These official options do not require you to paste an API key. You can also select **Add Provider**, choose a built-in preset or Custom, and enter the API key, base URL, API format, and model mapping.
Before saving a custom provider, select **Test Connection**. When the test succeeds, make the provider the default. The app can translate OpenAI Chat Completions and OpenAI Responses requests through its built-in local proxy, so LiteLLM is not required by default.
See [Third-party models and custom providers](./third-party-models.md) for configuration details.
### 3. Create your first task
Select `+` in the sidebar:
1. Choose a local project directory.
2. If it is a Git repository, choose the branch to use.
3. Decide whether to use the current working tree or create an isolated worktree.
The current working tree shares any uncommitted changes already in that directory. An isolated worktree is better for parallel tasks or changes that should not affect your current checkout.
### 4. Confirm the model and permissions
Before sending the first message, confirm the Provider, model, and effort. For a first run, keep **Default** permission mode: the app will ask before sensitive tools or commands run.
Use automatic approval or bypass modes only when you understand their impact. See [Desktop quick start](/en/desktop/01-quick-start#5-choose-a-permission-mode) for the permission modes.
### 5. Send the first message
Start with a small, verifiable request:
```text
Inspect this project in read-only mode. Explain how to start it and what its main directories do. Do not modify files.
```
If you see a streaming response, tool calls, and permission requests, the Desktop app, provider, and project directory are connected.
## Option 2: Run the CLI from Source
### 1. Clone the repository and install Bun
Install [Git](https://git-scm.com/downloads) and [Bun](https://bun.sh), then run:
```bash
git clone https://github.com/NanmiCoder/cc-haha.git
cd cc-haha
bun install
```
### 2. Configure a model provider
```bash
cp .env.example .env
```
Edit `.env` with at least one valid authentication method, base URL, and model. See [Environment variables](./env-vars.md) for the variable definitions and authentication-header differences.
Never commit a real API key to Git or expose it in an issue, screenshot, or diagnostic attachment.
### 3. Start and verify
On macOS, Linux, or Git Bash:
```bash
./bin/claude-haha
./bin/claude-haha -p "Summarize the current project structure"
```
On Windows PowerShell or Command Prompt:
```powershell
bun --env-file=.env ./src/entrypoints/cli.tsx
```
See the [CLI reference](./cli-reference.md) for command options, headless mode, recovery mode, and global usage.
## Next Steps
- [Desktop quick start](/en/desktop/01-quick-start): sessions, permissions, attachments, and workspace actions
- [Third-party models and custom providers](./third-party-models.md): providers, API formats, and model mapping
- [FAQ](./faq.md): installation, OAuth, H5, worktree, and Computer Use troubleshooting
- [Global usage](./global-usage.md): start the CLI from any directory
-98
View File
@@ -1,98 +0,0 @@
# Third-Party Models
Claude Code Haha can connect to Anthropic Messages, OpenAI Chat Completions, and OpenAI Responses APIs. For Desktop users, the reliable entry point is **Settings → Providers**, not a hand-written environment file. The app stores authentication and model mappings and starts a local protocol proxy when one is required.
## Recommended setup
1. Open **Settings → Providers** in Desktop.
2. Choose a built-in preset or create a custom provider.
3. Enter the base URL, authentication details, and protocol format.
4. Configure at least the primary model. Add Haiku, Sonnet, and Opus mappings when the service uses different model IDs.
5. Test the provider before activating it.
6. Start a new session and verify a tool call. A text-only response does not prove that the full agent workflow works.
After activation, Desktop and the CLI processes it launches reuse the same provider configuration. Do not keep a second, stale key in `.env` or `~/.claude/settings.json`.
## Choose the correct protocol
| Provider format | Upstream API | Use it when |
|-----------------|--------------|-------------|
| `anthropic` | Anthropic Messages | The service natively accepts Anthropic requests and responses |
| `openai_chat` | `/v1/chat/completions` | The service implements OpenAI Chat Completions |
| `openai_responses` | `/v1/responses` | The service implements OpenAI Responses |
The `anthropic` format calls the service directly without changing the protocol. For `openai_chat` and `openai_responses`, Claude Code Haha starts a loopback proxy that translates Claude agent traffic into the selected protocol. This proxy belongs to the local runtime and does not need to be exposed to a LAN.
If a provider advertises both “OpenAI compatible” and “Anthropic compatible,” choose the API that it implements completely and that handles tool calls reliably. Do not infer the protocol only from a `/v1` path.
## Authentication and model mappings
The service determines the authentication header:
- Anthropic API keys usually use `x-api-key`.
- Bearer tokens usually use `Authorization: Bearer`.
- OpenAI-compatible services usually use bearer tokens, but private gateways can differ.
Model slots are logical tiers requested by the Claude agent, not additional downloads. The primary model must support tool use. Other slots can map to that same model or to separate models chosen for cost and capability. Leave a slot empty when the provider does not support it.
Model names and provider capabilities change frequently, so this page intentionally does not maintain a soon-to-be-stale model catalog. Use the current model ID reported by the provider and confirm it with the test action in Settings.
## Built-in runtimes
Claude Code Haha also includes Claude, OpenAI, and Grok runtimes. Available sign-in methods depend on the current build and local account state; follow the authorization flow shown on the Providers page.
Built-in runtimes and custom compatible endpoints are separate paths. Prefer a built-in runtime for an existing official account. Create a custom provider for a relay service, private gateway, or local model server.
## CLI-only configuration
CLI-only users can configure an Anthropic Messages-compatible endpoint directly:
```bash
ANTHROPIC_AUTH_TOKEN=sk-example
ANTHROPIC_BASE_URL=https://provider.example.com/anthropic
ANTHROPIC_MODEL=provider-model
./bin/claude-haha
```
This path does not translate Anthropic requests into an OpenAI protocol. Configure OpenAI Chat Completions or Responses services as Desktop providers so the app can manage the protocol proxy. See [Environment Variables](./env-vars.md) for the complete variables and effective precedence.
Azure OpenAI uses the project's dedicated Responses path. See [Azure OpenAI environment variables](./env-vars.md#azure-openai).
## LiteLLM: an advanced compatibility layer
Add LiteLLM only when the service has no reliable Anthropic endpoint and the built-in provider translation is not suitable. It introduces another service, another protocol conversion, and another troubleshooting boundary.
Minimal example:
```yaml
model_list:
- model_name: provider-model
litellm_params:
model: openai/provider-model
api_base: https://provider.example.com/v1
api_key: os.environ/PROVIDER_API_KEY
```
After starting LiteLLM, connect its Anthropic-compatible URL as an `anthropic` provider. Follow the [official LiteLLM documentation](https://docs.litellm.ai/) for deployment, authentication, and model prefixes.
## Capability boundaries
A third-party model needs at least the following behavior for a stable agent workflow:
- correct handling of multi-turn messages, system content, and tool calls;
- preservation of tool-call IDs and acceptance of matching tool results;
- sufficient context and output limits;
- complete, correctly ordered events in streaming mode.
Thinking modes, effort, prompt caching, image input, and structured output depend on the protocol, provider, and exact model. They cannot be labeled universally supported or unsupported. When diagnosing a failure, disable optional capabilities, verify basic text plus tool use, and restore features one at a time.
## Troubleshooting
| Symptom | Check first |
|---------|-------------|
| `401` / `403` | Authentication header, token permissions, and whether the base URL belongs to that credential |
| `404` | Provider format and whether the base URL duplicates an API path |
| Model not found | Use the provider's real model ID, not an alias from another platform |
| Text works but tools never run | Confirm that both the model and gateway fully support tool calls |
| Old endpoint after switching providers | Remove duplicate configuration from `.env`, `settings.json`, or Desktop |
| Streaming stops early | Disable optional capabilities and check whether a gateway buffers or rewrites the event stream |
+69 -46
View File
@@ -1,80 +1,85 @@
---
title: DingTalk Integration
nav_title: DingTalk
description: Authorize a DingTalk bot by QR code and run over DingTalk Stream, with no public callback URL required.
order: 4
---
# DingTalk Integration
The DingTalk adapter uses DingTalk Stream, so it does not require a public callback URL. It supports private chats only.
For organizations already running on DingTalk: one QR scan authorizes a bot, and the adapter connects over DingTalk Stream, so no public callback URL or tunnel is needed. Replies stream through an AI Card by default. Only private chats are supported, and tappable approval cards need one extra template ID — otherwise approval is text replies.
Current behavior includes text, image attachments, project selection, status, stop, AI Card streaming, and text or card-based permission approval.
## Bind the bot by QR code
## Bind a DingTalk bot
1. Open **Settings → IM Adapters** and select the **DingTalk** tab.
2. Under **Bind DingTalk bot with QR**, select **Scan to bind**.
3. Scan with the DingTalk mobile app and confirm bot creation and authorization.
4. Wait for Desktop to report the bound state.
Open **Settings → IM Integration → DingTalk**.
The recommended flow is:
1. Select **Scan to bind**.
2. Scan with the DingTalk mobile app.
3. Confirm bot creation and authorization.
4. Wait for Desktop to show the bound state.
5. Save the configuration.
Desktop stores `clientId` and `clientSecret` in `~/.claude/adapters.json` and restarts the adapter sidecar.
**Client ID** and **Client Secret** are filled in and written to local configuration automatically, and the adapter reconnects with them.
## Manual credentials
If QR binding is unavailable, enter:
If QR binding is unavailable, fill in the fields below the QR area yourself:
- **Client ID** — DingTalk application `appKey`
- **Client Secret** — application `appSecret`
- **Client ID** — the DingTalk application `appKey`
- **Client Secret** — the application `appSecret`
- **Stream Endpoint** — optional; defaults to `https://api.dingtalk.com`
- **Permission card template ID** — optional
- **Allowed Users** — optional explicit user IDs
- **Permission Card Template ID** — optional; when set, permission requests prefer an interactive card
- **Allowed Users** — optional; explicit DingTalk user IDs
Saving starts a DingTalk Stream connection.
Select **Save**, and the adapter opens a DingTalk Stream connection with the new credentials.
## Authorize a user
Binding the bot does not authorize every DingTalk account. Generate a six-character pairing code in Desktop and send it to the bot in a private chat.
Binding the bot does not authorize every DingTalk account.
Codes expire after 60 minutes, are one-time use, and are rate limited after repeated failures. Paired users can be removed from Desktop at any time.
1. Back at the top of the page, under **Pairing**, select **Generate Code**.
2. Send the six-character code to the bot in a private chat.
3. Once pairing is confirmed, you can start chatting.
## Projects and commands
Codes are valid for 60 minutes and work once. Paired accounts appear under **Paired Users** and can be revoked at any time; a revoked user needs a fresh code.
If no default project is configured, the adapter returns recent projects. Reply with a number, project name, or absolute path. The mapping is persisted in `~/.claude/adapter-sessions.json`.
## Projects and sessions
Supported commands:
With **Default Project** set, the first accepted message opens a session in that directory. Without it, the bot lists recent projects; reply with a number, project name, or absolute path.
- `/help` or `帮助`
- `/status` or `状态`
- `/projects` or `项目列表`
- `/new` or `新会话`
- `/new <number, project name, or absolute path>`
- `/clear` or `清空`
- `/stop` or `停止`
Later messages in the same chat reuse that session. `/new` picks another project; `/clear` empties the context while keeping the project binding.
## Permission approval
## Commands
Without a published card template, reply:
DingTalk has no bot menu like Feishu, so the chat box is the only entry point. After pairing, the bot suggests `/help`.
- `/allow <requestId>`
- `/always <requestId>`
- `/deny <requestId>`
- `/help` or `帮助` — show available commands
- `/status` or `状态` — project, branch, model, run state, and task summary
- `/projects` or `项目列表` — list recent projects again
- `/new` or `新会话` — clear the current binding and choose a project
- `/new <number, project name, or absolute path>` — start directly in that project
- `/clear` or `清空` — clear context, keep the project binding
- `/stop` or `停止` — stop the current generation
With a configured permission-card template, the adapter prefers an interactive card and receives its callback over DingTalk Stream. Text commands remain the fallback.
## Approval and reply behavior
## Replies and attachments
By default a permission request arrives as text. Reply with one of:
- Private messages arrive through `dingtalk-stream`.
- Normal replies use the message `sessionWebhook`.
- AI Card is preferred for streaming output.
- Image attachments are downloaded and passed as inline image input.
- Shared attachment size limits are enforced.
- `/allow <requestId>` — allow once
- `/always <requestId>` — persist the matching approval
- `/deny <requestId>` — deny
With a published template entered in **Permission Card Template ID**, the adapter prefers an interactive card and receives the button callback over DingTalk Stream. Text commands stay available whenever a card fails to send or cannot be seen.
Normal replies prefer AI Card streaming. Image attachments are downloaded and passed as inline image input; oversized attachments are reported back in the chat.
## Unbind
Unbinding the bot clears DingTalk credentials, allowlisted and paired users, and the permission-card template ID. Removing a single paired user revokes only that user.
Two levels:
- **Unbind bot account** clears the DingTalk credentials, the permission card template ID, and this platform's authorization lists. Binding again means scanning or entering credentials again.
- **Unbind** next to a name under **Paired Users** revokes only that person.
## Development
Packaged Desktop starts the sidecar automatically. For source development:
Packaged Desktop starts the sidecar automatically. Run it by hand only when working from source:
```bash
cd adapters
@@ -91,3 +96,21 @@ export DINGTALK_STREAM_ENDPOINT="https://api.dingtalk.com"
export DINGTALK_PERMISSION_CARD_TEMPLATE_ID="..."
export ADAPTER_SERVER_URL="ws://127.0.0.1:3456"
```
Normal Desktop use needs none of these; QR binding or manual entry writes the local configuration.
## Troubleshooting
**Still unauthorized after a successful scan.** That is the expected flow. Scanning only stores bot credentials; the individual user still has to send a pairing code or appear in **Allowed Users**.
**Scanning fails or waits forever.** Select **Scan to bind** again for a fresh QR code. If it still fails, enter **Client ID** and **Client Secret** manually.
**The adapter reports missing credentials.** Neither the environment variables nor `dingtalk.clientId` / `dingtalk.clientSecret` in `~/.claude/adapters.json` took effect. Complete QR binding or manual entry in Settings first.
**No permission card appears.** Confirm the template ID belongs to a published template, that the template supports the callback route the adapter uses, and that the bot has published a new version. Text approval works regardless.
**No replies.** Confirm Desktop is running, the tab shows the bound state, the sender is paired or allowlisted, and `~/.claude/adapters.json` is writable. When running from source, also confirm `ADAPTER_SERVER_URL` points at the running Desktop server.
## Source
`index.ts`, `helpers.ts`, `ai-card.ts`, `permission-card.ts`, and `media.ts` under `adapters/dingtalk/`, plus `pairing.ts`, `session-store.ts`, `ws-bridge.ts`, and `http-client.ts` under `adapters/common/`.
+53 -30
View File
@@ -1,59 +1,70 @@
---
title: Feishu Integration
nav_title: Feishu
description: Create a Feishu bot from the official template and drive Desktop sessions from a private chat with card-based approval.
order: 1
---
# Feishu Integration
The Feishu adapter connects an enterprise custom app to local Desktop sessions. It handles `p2p` private chats only; group chats are not supported.
Best for teams already on Feishu: an official template creates a bot with every required permission preconfigured, and permission requests arrive as interactive cards you can tap. Common commands can be exposed as a bot menu. It handles private (`p2p`) chats only, and changing the bot configuration means publishing a new version in the developer console.
## Create the Feishu app
## Create the bot
Feishu provides a template with the messaging, event, and card capabilities needed by this integration:
Open [Create a Feishu bot](https://open.feishu.cn/page/openclaw?form=multiAgent). This is the official OpenClaw template, with messaging, event subscription, and card callback permissions already granted — no scopes to add by hand. The **Create Feishu bot** button under **Settings → IM Adapters → Feishu** opens the same page.
[Create a Feishu bot from the template](https://open.feishu.cn/page/openclaw?form=multiAgent)
Choose a name, create the app, then keep its **App ID** and **App Secret** for the next step.
Choose a name, create the app, and save its **App ID** and **App Secret**.
## Configure the bot menu
In the [Feishu developer console](https://open.feishu.cn/app?lang=en-US), create and publish a bot version. Optional menu entries can invoke:
This step is optional. With a menu, you can switch projects and start sessions by tapping instead of typing.
- `/projects`
- `/new`
- `/clear`
In the [Feishu developer console](https://open.feishu.cn/app?lang=en-US), open your bot and go to its bot menu configuration. Add three entries, each with a label of your choice and one of these commands:
The app must be published before users can reliably receive the configured bot behavior.
- `/projects` — list recent projects and switch
- `/new` — start a new session
- `/clear` — clear the current context
## Configure Desktop
Save the menu, then publish a new application version. Menu changes take effect only after publishing.
Open **Settings → IM Integration → Feishu**:
## Enter credentials in Desktop
1. Enter the App ID and App Secret.
2. Generate a six-character pairing code.
3. Save the configuration.
4. Send a private message to the bot and provide the code.
1. Open **Settings → IM Adapters** and select the **Feishu** tab.
2. Paste the values into **App ID** and **App Secret**.
3. Leave **Encrypt Key** and **Verification Token** empty; the template bot does not need them.
4. Enable **Streaming Card Mode** if you want long replies to update one card in place.
5. Select **Save**.
Pairing codes expire after 60 minutes, work once, and are rate limited after repeated failures. App credentials do not authorize every organization user.
**Allowed Users** can stay empty. When it is, only paired accounts are accepted, which is usually what you want.
## Pair your account
At the top of the page, under **Pairing**, select **Generate Code**. The six-character code is written to local configuration immediately — no separate save.
Send any message to your new bot in Feishu, then send the code when prompted. Once pairing is confirmed you can talk to Claude Code directly.
Codes are valid for 60 minutes, work once, and are invalidated when a new one is generated.
## Commands
Alongside the menu buttons, these work in the chat box at any time:
- `/help` or `帮助`
- `/status` or `状态`
- `/clear` or `清空`
- `/projects` or `项目列表`
- `/new` or `新会话`
- `/clear` or `清空`
- `/stop` or `停止`
## Permission approval
## Approval and reply behavior
Permission requests are sent as interactive cards. Selecting allow or deny returns the result to the pending Desktop session.
Permission requests arrive as interactive cards. Selecting allow or deny returns the result to the pending Desktop session.
If card actions do not work, confirm that the latest application version is published and includes the required card-action capability.
## Reply behavior
- Normal text uses Feishu post messages.
- Permission approval uses cards.
- Streaming output prefers patching the same message.
- Long completed content is split to respect platform limits.
Normal replies use Feishu post messages. Streaming output prefers patching the same message, and long completed text is split to respect platform limits.
## Development
Packaged Desktop starts the sidecar automatically. For source development:
Packaged Desktop starts the sidecar automatically. Run it by hand only when working from source:
```bash
cd adapters
@@ -69,4 +80,16 @@ export FEISHU_APP_SECRET="xxx"
export ADAPTER_SERVER_URL="ws://127.0.0.1:3456"
```
If messages are missing, confirm that the app is published, the conversation is a private chat, and the sender is paired or explicitly allowlisted.
## Troubleshooting
**No messages arrive.** Confirm the app is published — menu edits require a new version — and that the conversation is a private chat rather than a group.
**Card buttons do nothing.** The card action capability usually did not ship with the published version. Publish again from the developer console.
**Still unauthorized.** Check that the code is within its 60-minute window, that it is the current one, and that the account now appears under **Paired Users** in Desktop.
**Session not restored after a restart.** Verify that `~/.claude/adapter-sessions.json` is writable and that the session still exists in Desktop.
## Source
`adapters/feishu/index.ts`, plus `pairing.ts`, `session-store.ts`, `ws-bridge.ts`, and `http-client.ts` under `adapters/common/`.
+67 -49
View File
@@ -1,76 +1,94 @@
---
title: IM Integrations
nav_title: Overview
description: Bridge Feishu, Telegram, WeChat, DingTalk, or WhatsApp private chats into the Desktop app and continue the same session from your phone.
order: 0
---
# IM Integrations
Claude Code Haha can bridge private messages from WeChat, DingTalk, WhatsApp, Telegram, and Feishu into local Desktop sessions.
A session running in the Desktop app can be reached from a private chat on your phone. Once bound, a message in Feishu, Telegram, WeChat, DingTalk, or WhatsApp drives the Claude Code session on your own machine: start a long task before you leave, then follow the progress, approve permissions, and switch projects from the road.
These integrations do not make the local agent public. A user is accepted only when they are explicitly listed in `allowedUsers` or have completed pairing.
The chat partner is a bot or account you bound yourself. Messages reach your local Desktop app; no intermediate service holds your code.
## Current architecture
![Settings shows pairing management on top and one tab per platform](../../images/app/settings-im.webp)
The supported path is the Desktop adapter architecture:
## What you get
```mermaid
flowchart LR
A["Desktop Settings"] --> B["/api/adapters"]
B --> C["~/.claude/adapters.json"]
C --> D["Platform adapter sidecar"]
D --> E["Pairing and allowlist check"]
E --> F["HTTP session creation"]
F --> G["/ws/:sessionId"]
G --> H["Claude Code session"]
```
- **The same session, continued.** Messages sent from your phone enter the Claude Code session on your computer, where file edits, commands, and reads really happen.
- **Project switching.** `/projects` lists recent projects and switches to the one you pick; `/new` starts a fresh session.
- **Permission approval.** When Claude wants to write a file or run a risky command, the request is pushed to the chat. Feishu and DingTalk send interactive cards, Telegram sends buttons, WeChat and WhatsApp expect a text reply.
- **Status and stop.** `/status` reports the current project, model, and run state; `/stop` interrupts the current turn.
Platform adapters are separate local processes. They read their platform configuration, enforce authorization, map a private chat to a project session, and bridge messages over the Desktop server.
The Desktop app has to stay running. The chat side is only a remote control.
This is different from the upstream Claude Code Channel/MCP research retained under [Channel System](../channel/).
## Choosing a platform
## Set up an integration
All five expose the same capabilities. They differ in setup cost and approval experience.
1. Keep the Desktop app running.
2. Open **Settings → IM Integration**.
3. Configure or bind one platform.
4. Set a default project if desired.
5. Generate a six-character pairing code.
6. Send the code in a private chat with the bot or bound account.
| Platform | How you connect | Best for | Known limits |
|---|---|---|---|
| Feishu | Create a bot from the official template, paste its App ID and App Secret | Teams that want one-tap permission approval | Private (`p2p`) chats only; menu changes require publishing a new app version |
| Telegram | Ask `@BotFather` for a Bot Token, paste it into Settings | Individuals who can reach Telegram; fastest setup | Private chats only |
| WeChat | Scan a QR code in Settings to log in a bot account | People who only want WeChat | Private chats only; permission approval is text replies |
| DingTalk | Scan a QR code in Settings; credentials are filled in for you | Organizations already on DingTalk | Private chats only; interactive approval cards need an extra template ID |
| WhatsApp | Scan from **Linked devices** on your phone | Users outside mainland China | Personal linked-device login, not the official Cloud API; personal private chats only |
The code expires after 60 minutes, can be used once, and is invalidated when a new code is generated. Repeated failures are rate limited.
If you have no preference, start with Telegram or Feishu — their approval flows are the most comfortable.
## Platforms
## Pairing flow
- [WeChat](./wechat.md) — QR-bound bot account, private chats
- [DingTalk](./dingtalk.md) — DingTalk Stream, private chats
- [WhatsApp](./whatsapp.md) — personal linked-device session through WhatsApp Web
- [Telegram](./telegram.md) — BotFather token, private chats
- [Feishu](./feishu.md) — enterprise custom app, `p2p` chats
Binding happens in two layers: first the Desktop app gets platform credentials, then your personal account is authorized with a pairing code. The second layer is identical everywhere.
## Local state
1. Open **Settings → IM Adapters**.
2. Bind one platform in its tab: Feishu and Telegram take credentials, WeChat, DingTalk, and WhatsApp use a QR code.
3. Pick a directory under **Default Project**.
4. Select **Save**.
5. Back at the top, in **Pairing**, select **Generate Code** to get a six-character code.
6. Send that code to your bot in a private chat on the matching platform.
7. Once pairing is confirmed, anything you type goes to Claude Code.
`~/.claude/adapters.json` stores platform configuration, allowlists, and pairing state. Sensitive fields returned by the settings API are masked.
A code is valid for 60 minutes, works once, and is invalidated the moment a new one is generated. The code itself is platform-neutral — it binds whichever account sends it. Five failed attempts within five minutes trigger rate limiting.
`~/.claude/adapter-sessions.json` stores chat-to-session mappings, including the session ID, working directory, and update time. This allows an adapter to reconnect to an existing session after restart.
Generating a code and QR binding are written to local configuration immediately. **Save** is only needed for typed values such as App ID, Bot Token, **Allowed Users**, and **Default Project**.
Both paths follow `CLAUDE_CONFIG_DIR` when a custom data directory is active.
Paired accounts appear under **Paired Users**, where **Unbind** revokes one of them. A revoked user needs a fresh code.
## Authorization model
## Default project decides where work happens
- `allowedUsers` and `pairedUsers` are combined.
- If both are empty, access is denied.
- Binding a bot or linked account does not authorize its contacts.
- Removing a paired user requires that user to pair again.
- Removing platform credentials stops that platform from connecting.
**Default Project** is the working directory for new IM sessions. With it set, the first message from your phone opens a session in that directory. Left empty, the bot lists recent projects and asks you to choose.
Later messages in the same chat reuse that session, and the mapping survives a Desktop restart. `/new` changes the directory; `/clear` empties the context while keeping the project binding.
## Common commands
The exact presentation differs by platform, but adapters support the same core operations:
Entry points differ slightly per platform — Feishu can expose commands as a bot menu — but these work everywhere:
- `/help` — show available commands
- `/status` — show the current project, model, and run state
- `/projects` — list or switch recent projects
- `/new` — start a new session or choose another project
- `/clear` — clear current context while keeping the project binding
- `/help` — list available commands
- `/status` — current project, model, and run state
- `/projects` — list recent projects and switch
- `/new` — start a new session, optionally with a project number or path
- `/clear` — clear context, keep the project binding
- `/stop` — stop the current generation
Permission requests are returned as platform buttons, cards, or explicit text commands. A request ID must match the pending Desktop session request.
WeChat, DingTalk, and Feishu also accept Chinese aliases such as `帮助`, `状态`, `项目列表`, `新会话`, `清空`, and `停止`.
## Development
## Security
Packaged Desktop starts configured adapter sidecars automatically. Manual `bun run <platform>` commands are for source development and isolated troubleshooting.
::: warning This is a remote control for your computer
A paired account can make Claude read files, write files, and run commands on your machine. Send pairing codes only to yourself, never post one in a group, and never commit bot credentials.
:::
Authorization is the union of **Allowed Users** and paired users. When both are empty, every sender is rejected. Binding a bot or a linked account does not authorize its contacts.
Platform credentials, pairing state, and allowlists live in `~/.claude/adapters.json`; chat-to-session mappings live in `~/.claude/adapter-sessions.json`. Both stay on your machine, both contain material that can drive it, and neither should be shared. Sensitive fields are masked when the settings page reads the configuration back. Both paths follow `CLAUDE_CONFIG_DIR` when a custom data directory is active.
For a full mobile interface rather than a chat window, see [H5 access](../desktop/remote.md).
## Per-platform guides
- [Feishu](./feishu.md) — template bot, card approval
- [Telegram](./telegram.md) — BotFather token, button approval
- [WeChat](./wechat.md) — QR-bound account, text approval
- [DingTalk](./dingtalk.md) — QR authorization, AI Card streaming
- [WhatsApp](./whatsapp.md) — personal linked device, text approval
+47 -36
View File
@@ -1,60 +1,59 @@
---
title: Telegram Integration
nav_title: Telegram
description: Get a Bot Token from BotFather, paste it into Desktop, and approve permissions with native buttons.
order: 2
---
# Telegram Integration
The Telegram adapter connects a BotFather bot to local Desktop sessions. It accepts private chats only; groups are not supported.
The fastest of the five to set up: ask `@BotFather` for a token, paste it into Desktop, done. Permission requests come back as native buttons. It accepts private chats only; groups are not supported.
## Create a bot
In Telegram, open the official **@BotFather** account:
In Telegram, open the official `@BotFather` account and send `/newbot`. Then:
1. Send `/newbot`.
2. Choose a display name.
3. Choose a username ending in `_bot`.
4. Copy the Bot Token returned by BotFather.
1. Choose a display name, for example `ClaudeCodeHaha Bot`.
2. Choose a username in Latin letters ending in `_bot`, for example `jiang_cc_hah_bot`.
3. Copy the **Bot Token** that BotFather returns.
Treat the token as a credential.
That token is the bot's password. Do not paste it anywhere public.
## Configure Desktop
## Enter the token in Desktop
Open **Settings → IM Integration → Telegram**:
1. Open **Settings → IM Adapters** and select the **Telegram** tab.
2. Paste the value into **Bot Token**.
3. Select **Save**.
1. Paste the Bot Token.
2. Generate a six-character pairing code.
3. Save the configuration.
4. Send a message to the new bot and provide the pairing code when prompted.
**Allowed Users** can stay empty. When it is, only paired accounts are accepted. To allowlist a known account directly, enter its numeric Telegram user ID; separate several with commas.
The code expires after 60 minutes, works once, and is rate limited after repeated failures. Bot configuration alone does not authorize all Telegram users.
## Pair your account
At the top of the page, under **Pairing**, select **Generate Code**. The six-character code takes effect immediately — no separate save.
Send any message to your new bot, then send the code when prompted. Once pairing is confirmed you can talk to Claude Code directly.
Codes are valid for 60 minutes, work once, and are invalidated when a new one is generated. Repeated failures are rate limited; wait a few minutes before retrying.
## Commands
- `/start` — show help
- `/start` — show help and available commands
- `/help` — show available commands
- `/projects` — list or switch recent projects
- `/status` — show project, model, run state, and task summary
- `/clear` — clear context while keeping the project
- `/new` — start a new session and choose a project
- `/projects` — list recent projects and switch
- `/status` — project, model, run state, and task summary
- `/new` — clear the current binding and choose a project again
- `/clear` — clear context, keep the project binding
- `/stop` — stop the current generation
## Permission approval
## Approval and reply behavior
Telegram presents buttons for:
A permission request arrives as a message with three buttons: allow once, always allow the matching operation, and deny. The choice is converted to a `permission_response` for the pending Desktop session.
- allow once;
- always allow the matching operation;
- deny.
The callback is converted to a `permission_response` for the pending Desktop session.
## Reply behavior
The adapter buffers streaming output:
- a placeholder can be sent during thinking;
- text deltas are accumulated;
- completed text is split into platform-sized messages.
Replies pass through a streaming buffer: a placeholder can be sent while Claude is thinking, text deltas accumulate in place, and completed text is split into platform-sized messages.
## Development
Packaged Desktop starts the adapter automatically. For source development:
Packaged Desktop starts the sidecar automatically. Run it by hand only when working from source:
```bash
cd adapters
@@ -69,4 +68,16 @@ export TELEGRAM_BOT_TOKEN="123456:ABC-DEF..."
export ADAPTER_SERVER_URL="ws://127.0.0.1:3456"
```
If a sender is rejected, verify that the current pairing code was sent to the correct bot private chat and that the sender is now present in the paired or allowed list.
## Troubleshooting
**The adapter reports a missing token.** Neither `TELEGRAM_BOT_TOKEN` nor `telegram.botToken` in `~/.claude/adapters.json` took effect. Re-enter the token in Settings and save.
**Settings opens but the bot does nothing.** When running from source, the web app only writes configuration; it does not launch `bun run telegram`. Packaged Desktop starts the sidecar for you.
**A sender is rejected.** Confirm a code was generated, that it is within its 60-minute window, and that it was sent to the correct bot in a private chat.
**Session not restored after a restart.** Verify that `~/.claude/adapter-sessions.json` is writable and that the session still exists in Desktop.
## Source
`adapters/telegram/index.ts`, plus `pairing.ts`, `session-store.ts`, `ws-bridge.ts`, `message-buffer.ts`, and `format.ts` under `adapters/common/`.
+57 -45
View File
@@ -1,79 +1,75 @@
---
title: WeChat Integration
nav_title: WeChat
description: Scan a QR code in Settings to log in a WeChat bot account, then authorize individual users with a pairing code.
order: 3
---
# WeChat Integration
The WeChat adapter binds a bot account by QR code, then authorizes individual private-chat users through pairing or an allowlist.
For people who only want to use WeChat: the whole setup is one QR scan, with no developer platform to register on. The trade-off is that permission approval is text replies rather than tappable buttons, and only private chats are supported.
It supports text, transcribed voice content, image and file attachments, project selection, status, stop, and text-based permission approval. Group chats are not supported.
Besides text, this path also accepts transcribed voice content, images, and file attachments.
## Bind the WeChat bot account
## Bind the bot account
Open **Settings → IM Integration → WeChat**:
1. Open **Settings → IM Adapters** and select the **WeChat** tab.
2. Select **Scan to Bind**.
3. Scan the QR code with WeChat and confirm the login on your phone.
4. Wait for the status to read **WeChat is bound**.
1. Select **Scan to bind**.
2. Scan the QR code with WeChat.
3. Confirm the login in WeChat.
4. Wait for Desktop to show the bound state.
5. Save the configuration.
Credentials are written to local configuration and the adapter restarts on its own; no separate save is needed. The tab then offers **Rescan** and **Unbind WeChat account**.
Desktop stores the returned `accountId`, `botToken`, `baseUrl`, and `userId` in the `wechat` section of `~/.claude/adapters.json`, then restarts the adapter sidecar.
Binding credentials does not authorize every WeChat user.
Binding the bot account is not the same as allowing every WeChat contact. Who may use it is decided in the next step.
## Authorize a user
Generate a pairing code at the top of **IM Integration**, then send the six-character code to the bound bot in a private chat.
1. Back at the top of the page, under **Pairing**, select **Generate Code**.
2. Send the six-character code to the bound bot in WeChat.
3. Once pairing is confirmed, you can send messages directly.
The code:
Codes are valid for 60 minutes, work once, and are invalidated when a new one is generated.
- expires after 60 minutes;
- can be used once;
- is replaced immediately when a new code is generated;
- is rate limited after repeated failures.
Known WeChat user IDs can instead be entered in `Allowed Users`.
Known WeChat user IDs can instead be entered in **Allowed Users**, separated by commas. That suits a fixed allowlist but is less convenient than pairing.
## Projects and sessions
With a default project configured, the first accepted message creates or resumes a session in that directory.
With **Default Project** set, the first accepted message opens a session in that directory. Without it, the bot lists recent projects; reply with a number, project name, or absolute path.
Without a default project, the adapter returns recent projects. Reply with a number, project name, or absolute path. The resulting chat-to-session mapping is stored in `~/.claude/adapter-sessions.json`.
Use `/new` to choose another project and start a new session.
Later messages in the same chat reuse that session, and the mapping survives a Desktop restart. `/new` picks another project; `/clear` empties the context while keeping the project binding.
## Commands
- `/help` or `帮助`
- `/status` or `状态`
- `/projects` or `项目列表`
- `/new` or `新会话`
- `/new <number, project name, or absolute path>`
- `/clear` or `清空`
- `/stop` or `停止`
WeChat has no configurable menu, so the chat box is the only entry point. After pairing, the bot suggests `/help`.
## Permission approval
- `/help` or `帮助` — show available commands
- `/status` or `状态` — project, branch, model, run state, and task summary
- `/projects` or `项目列表` — list recent projects again
- `/new` or `新会话` — clear the current binding and choose a project
- `/new <number, project name, or absolute path>` — start directly in that project
- `/clear` or `清空` — clear context, keep the project binding
- `/stop` or `停止` — stop the current generation
Reply to the text approval message with:
## Approval and reply behavior
A permission request arrives as text containing a request ID. Reply with one of:
- `/allow <requestId>` — allow once
- `/always <requestId>` — persist the matching approval
- `/deny <requestId>` — deny
The adapter sends the response back to the pending Desktop session.
## Attachments and replies
- WeChat messages are received through long polling.
- Long text is split into platform-sized messages.
- Images enter model input as inline images.
- Other files are downloaded to a local temporary path for the session.
- Shared attachment size limits are enforced.
Messages are received through long polling and long replies are split into platform-sized messages. A typing indicator is shown while Claude thinks or runs tools. Images enter model input inline; other files are downloaded to a local temporary path. Oversized attachments are reported back in the chat.
## Unbind
Unbinding the WeChat account clears its credentials and platform authorization lists. Removing one paired user only revokes that user.
Two levels:
- **Unbind WeChat account** clears the bot credentials and this platform's authorization lists. Binding again means scanning again.
- **Unbind** next to a name under **Paired Users** revokes only that person. They need a fresh pairing code afterwards.
## Development
Packaged Desktop starts the sidecar automatically. For source development:
Packaged Desktop starts the sidecar automatically. Run it by hand only when working from source:
```bash
cd adapters
@@ -91,4 +87,20 @@ export WECHAT_USER_ID="..."
export ADAPTER_SERVER_URL="ws://127.0.0.1:3456"
```
If messages are rejected, confirm that Desktop is running, the account is bound, the sender is paired or allowlisted, and both adapter state files are writable.
Normal Desktop use needs none of these; QR binding writes the local configuration.
## Troubleshooting
**Still unauthorized after a successful scan.** That is the expected flow. Scanning only stores bot credentials; the individual user still has to send a pairing code or appear in **Allowed Users**.
**The QR code expired.** Select **Scan to Bind** again. WeChat login QR codes are short-lived.
**The adapter reports a missing WeChat account.** Neither the environment variables nor `wechat.accountId` / `wechat.botToken` in `~/.claude/adapters.json` took effect. Complete QR binding in Settings first.
**No replies.** Confirm Desktop is running, the tab reads **WeChat is bound**, the sender is paired or allowlisted, and `~/.claude/adapters.json` is writable. When running from source, also confirm `ADAPTER_SERVER_URL` points at the running Desktop server.
**Session not restored after a restart.** Verify that `~/.claude/adapter-sessions.json` is writable and that the session still exists in Desktop.
## Source
`index.ts`, `protocol.ts`, and `media.ts` under `adapters/wechat/`, plus `pairing.ts`, `session-store.ts`, `ws-bridge.ts`, and `http-client.ts` under `adapters/common/`.
+50 -43
View File
@@ -1,83 +1,80 @@
---
title: WhatsApp Integration
nav_title: WhatsApp
description: Scan from Linked devices on your phone to attach the Desktop app as a logged-in WhatsApp Web device.
order: 5
---
# WhatsApp Integration
The WhatsApp adapter uses a personal WhatsApp Web linked-device session through `@whiskeysockets/baileys`. It is not the official WhatsApp Business Platform or Cloud API.
For users outside mainland China: no Meta developer account is required, and one QR scan lets your personal WhatsApp number drive the Desktop app. In exchange, this is a WhatsApp Web linked-device login rather than the official Cloud API — no official SLA, template messages, or broadcast capability. It handles personal private chats only, not groups, channels, or status. Permission approval is text replies.
It handles personal private chats only. Groups, channels, and status broadcasts are not supported.
## This is not "creating a bot"
## What linked-device access means
What the integration really does is attach the Desktop app as one more logged-in Web device on your WhatsApp account. So there is no Meta for Developers app, no WhatsApp Business Account, and no Phone Number ID, access token, webhook URL, or message template review.
You do not need a Meta developer app, WABA, Phone Number ID, Cloud API access token, webhook, or message template.
The person on the other side is talking to the WhatsApp account you bound, not to a separately created bot.
Instead:
1. Desktop creates a WhatsApp Web login QR code.
2. You scan it from **WhatsApp → Linked devices**.
3. Local Baileys auth state is saved.
4. The adapter observes private messages received by that account.
5. Only allowlisted or paired senders can reach a Claude Code session.
The chat recipient is the WhatsApp account you bound, not a separately created bot.
For customer service, template messages, an official SLA, or broadcasting, you would need the official Cloud API instead, which is a separate implementation. See the [Linked Devices help page](https://faq.whatsapp.com/1317564962315842/) and the [WhatsApp Cloud API overview](https://developers.facebook.com/docs/whatsapp/cloud-api/overview).
## Bind the account
Open **Settings → IM Integration → WhatsApp**:
1. Open **Settings → IM Adapters** and select the **WhatsApp** tab.
2. Select **Scan to Bind**.
3. On your phone, open **WhatsApp → Settings → Linked devices**.
4. Scan the QR code shown in Desktop.
5. Wait for the status to read **WhatsApp is bound**.
1. Select **Scan to bind**.
2. Open **Linked devices** on the phone.
3. Scan the QR code.
4. Wait for Desktop to confirm the binding.
The default local auth directory is:
The login state is saved locally and the adapter restarts on its own; no separate save is needed. The default directory is:
```text
~/.claude/whatsapp-auth/default
```
Do not publish or share this directory.
That directory is equivalent to a credential that can send and receive messages as you. Do not share it and do not commit it.
## Authorize a sender
Linked-device login does not authorize all contacts.
Linked-device login attaches the account; it does not authorize your contacts.
Generate a six-character pairing code in Desktop, then send it in a private WhatsApp chat to the bound account. The sender’s JID is added to `whatsapp.pairedUsers`.
1. Back at the top of the page, under **Pairing**, select **Generate Code**.
2. From the account you want to authorize, send the six-character code to the bound account in a private chat.
3. Once pairing is confirmed, that JID is recorded in the local authorization list.
Known JIDs can be entered in `Allowed Users`, for example:
Known JIDs can instead be entered in **Allowed Users** as a country code plus phone number:
```text
<country-code><phone-number>@s.whatsapp.net
15551234567@s.whatsapp.net
```
Codes are valid for 60 minutes, work once, and are invalidated when a new one is generated.
## Commands
- `/start` or `/help`
- `/projects`
- `/status`
- `/clear`
- `/new [project]`
- `/stop`
- `/start` or `/help` — show help and available commands
- `/projects` — list recent projects and switch
- `/status` — current project, model, and run state
- `/new [project]` — start a new session or switch projects
- `/clear` — clear context, keep the project binding
- `/stop` — stop the current generation
## Permission approval
## Approval and reply behavior
WhatsApp uses explicit text replies:
WhatsApp has no tappable buttons, so a permission request is answered by replying as the message instructs:
- `1` or `/allow <requestId>` — allow once
- `2` or `/always <requestId>` — persist the matching approval
- `3` or `/deny <requestId>` — deny
## Reply behavior
- Thinking can produce a short status message.
- Completed text is split into platform-sized messages.
- Recognized Markdown image output can be sent as an image message.
- The adapter does not depend on editing one WhatsApp message for token-level streaming.
The adapter does not rely on repeatedly editing one message for token-level streaming: a short status message is sent while Claude thinks, completed text is split into platform-sized messages, and Markdown image output is recognized and sent as an image message.
## Unbind
Use the WhatsApp settings page to unbind, remove local auth state, and scan again. Removing only a paired user revokes that sender without unlinking the account.
**Unbind WhatsApp account** removes the locally stored login state, after which you have to scan again. To revoke a single person instead, select **Unbind** next to their name under **Paired Users**; the account link itself is unaffected.
## Development
Packaged Desktop starts the sidecar automatically. For source development:
Packaged Desktop starts the sidecar automatically. Run it by hand only when working from source:
```bash
cd adapters
@@ -89,8 +86,18 @@ Optional overrides:
```bash
export WHATSAPP_AUTH_DIR="$HOME/.claude/whatsapp-auth/default"
export WHATSAPP_ACCOUNT_JID="<country-code><phone-number>@s.whatsapp.net"
export WHATSAPP_ACCOUNT_JID="15551234567@s.whatsapp.net"
export ADAPTER_SERVER_URL="ws://127.0.0.1:3456"
```
If the adapter reports that no account is bound, complete QR binding in Desktop first; the manual adapter command does not provide a separate login UI.
## Troubleshooting
**The adapter reports no bound account.** Complete QR binding in Desktop first; `bun run whatsapp` has no login UI of its own.
**Still unauthorized after binding.** Scanning binds the account, not the sender. Generate a pairing code and send it from the chat you want to authorize.
**WhatsApp reports a logout.** Unbind in Settings and scan again. Removing the device under **Linked devices** on your phone also logs it out.
## Source
`index.ts`, `protocol.ts`, `session.ts`, and `media.ts` under `adapters/whatsapp/`, plus `pairing.ts`, `session-store.ts`, and `ws-bridge.ts` under `adapters/common/`.
@@ -1,31 +1,30 @@
# Claude Code Agent Framework Deep Dive
---
title: Agent Framework Deep Dive
nav_title: Agent Framework
description: The core agent loop, prompt engineering, tool system, and context compression, read from source.
order: 6
---
> Deconstructing the architecture behind the world's most popular AI code editor — from source code to design philosophy.
# Agent Framework Deep Dive
<p align="center">
<a href="#1-the-core-agent-loop">Core Loop</a> · <a href="#2-system-prompt-engineering">Prompt Engineering</a> · <a href="#3-tool-system-design">Tool System</a> · <a href="#4-context-management-compression">Context Management</a> · <a href="#5-skills-plugin-ecosystem">Skills & Plugins</a> · <a href="#6-permission-security-model">Permissions</a> · <a href="#7-fault-recovery-mechanisms">Recovery</a> · <a href="#8-how-it-differs-from-langchain-react">vs LangChain</a> · <a href="#9-why-claude-code-is-so-good">Why It Works</a>
</p>
Deconstructing the architecture behind the world's most popular AI code editor — from source code to design philosophy.
![Agent Framework Architecture Overview](./images/11-agent-framework-overview.png)
---
## What This Framework Solves
## Preface: A Fundamental Question
If you observe Claude Code closely, you'll notice some remarkable behaviors:
Watch Claude Code closely and a few behaviors need explaining:
- It can modify dozens of files in a single conversation with extremely few errors
- It automatically recovers from edge cases (token overflow, API timeouts, tool failures)
- It can simultaneously manage multiple subagents collaborating on complex tasks
- Long conversations don't degrade — they actually become more precise over time
Behind these capabilities lies a carefully engineered Agent framework. This document deconstructs that framework from the source code level, revealing its core design philosophy.
None of that comes from the model alone; it is designed into the framework. The rest of this page walks the source in order.
---
## The Core Agent Loop
## 1. The Core Agent Loop
### 1.1 Not ReAct — An Async Generator State Machine
### Not ReAct — An Async Generator State Machine
Most agent frameworks (including LangChain) adopt the classic **ReAct** pattern:
@@ -42,7 +41,7 @@ export async function* query(params: QueryParams): AsyncGenerator<...>
This function is the heart of the entire agent. It's not a simple "think-act-observe" loop but a **streaming state machine** that yields messages in real-time and drives iteration through state assignment (not recursive calls).
### 1.2 The State Structure
### The State Structure
```typescript
// src/query.ts:204-217
@@ -60,7 +59,7 @@ type State = {
}
```
### 1.3 Five Phases of the Core Loop
### Five Phases of the Core Loop
The entire `while (true)` loop (`src/query.ts:307-1728`) consists of five phases:
@@ -140,11 +139,9 @@ No recursion, no callback hell — just simple `state = next` followed by `conti
- **State traceability**: Every transition reason is recorded
- **Controllable recovery**: Errors at any phase can be recovered by modifying state
---
## System Prompt Engineering
## 2. System Prompt Engineering
### 2.1 Layered Construction Architecture
### Layered Construction Architecture
The system prompt isn't a static string — it's dynamically assembled through a **layered pipeline** (`src/constants/prompts.ts:444-577`):
@@ -171,7 +168,7 @@ The **cache boundary (`SYSTEM_PROMPT_DYNAMIC_BOUNDARY`)** is a critical design e
This means Claude Code's system prompt **doesn't need to be reprocessed every time** — the static portion is shared globally, dramatically reducing latency and cost.
### 2.2 Two Section Types
### Two Section Types
```typescript
// src/constants/systemPromptSections.ts
@@ -187,7 +184,7 @@ DANGEROUS_uncachedSystemPromptSection('mcp_instructions', async () => {
}, 'MCP servers can connect/disconnect mid-session')
```
### 2.3 CLAUDE.md Loading Mechanism
### CLAUDE.md Loading Mechanism
CLAUDE.md is the custom instruction system, loaded by **priority from low to high** (`src/utils/claudemd.ts`):
@@ -205,7 +202,7 @@ project-root/CLAUDE.local.md ← Local private instructions (highest prio
Supports `@path` syntax for recursive file inclusion, with automatic circular reference prevention.
### 2.4 System Prompt Priority Resolution
### System Prompt Priority Resolution
The final system prompt is determined through `buildEffectiveSystemPrompt()` (`src/utils/systemPrompt.ts:41-123`):
@@ -216,11 +213,9 @@ The final system prompt is determined through `buildEffectiveSystemPrompt()` (`s
5. **Default prompt** — Standard system prompt
6. **Append prompt** — Always appended at the end
---
## Tool System Design
## 3. Tool System Design
### 3.1 Tools: More Than Function Calls
### Tools: More Than Function Calls
Claude Code's tools aren't simple "name + params + execute". Each tool is a **complete lifecycle management unit** (`src/Tool.ts:362-695`):
@@ -258,7 +253,7 @@ type Tool<Input, Output> = {
This design makes every tool **self-describing, self-validating, and self-rendering** — the framework doesn't need to understand tool internals, just call standard interfaces.
### 3.2 Tool Registration: Three-Stage Pipeline
### Tool Registration: Three-Stage Pipeline
Tool discovery and registration happens in three stages (`src/tools.ts`):
@@ -278,7 +273,7 @@ Stage 3: MCP Merge (assembleToolPool)
Sorting (cache stability)
```
### 3.3 Tool Execution Pipeline
### Tool Execution Pipeline
Each tool invocation passes through a **7-step pipeline** (`src/services/tools/toolExecution.ts`):
@@ -290,7 +285,7 @@ Each tool invocation passes through a **7-step pipeline** (`src/services/tools/t
Each step can **interrupt, modify, or enhance** the execution flow. This isn't a simple `try { tool.call(input) } catch` — it's a full middleware pipeline.
### 3.4 Deferred Tool Loading
### Deferred Tool Loading
Claude Code has 48+ built-in tools. Sending all tool definitions to the model on every API call would waste massive tokens. The solution:
@@ -305,11 +300,9 @@ Claude Code has 48+ built-in tools. Sending all tool definitions to the model on
The model dynamically retrieves full definitions via the `ToolSearch` tool when needed. This dramatically reduces system prompt size.
---
## Context Management & Compression
## 4. Context Management & Compression
### 4.1 The Secret Behind Unlimited Conversations
### The Secret Behind Unlimited Conversations
Claude Code claims "conversations have no context limit." Behind this is a **four-level compression system**:
@@ -331,7 +324,7 @@ Staged summarization of historical messages. Not all-at-once summarization, but
When all local optimizations are insufficient, Claude itself generates a complete conversation summary that replaces all historical messages.
### 4.2 System Context Injection
### System Context Injection
Before every API call, two types of context are automatically injected (`src/context.ts`):
@@ -350,7 +343,7 @@ getUserContext() → {
}
```
### 4.3 System Reminders
### System Reminders
System reminders are special **attachment messages** injected into tool results or user messages (`src/utils/attachments.ts`):
@@ -366,11 +359,9 @@ Use cases include:
- Accompanying information for side questions
- Availability notices for deferred tools
---
## Skills & Plugin Ecosystem
## 5. Skills & Plugin Ecosystem
### 5.1 Skills System
### Skills System
Skills are one of Claude Code's most powerful extension mechanisms. They're not simple "command aliases" but **complete AI behavior definitions**.
@@ -411,7 +402,7 @@ Project skills (.claude/skills/) ← Project-level
Policy skills (policy) ← Organization-managed
```
### 5.2 Plugin System
### Plugin System
Plugins are higher-level extension units that can contain **skills, hooks, MCP servers, and LSP servers**:
@@ -430,7 +421,7 @@ type BuiltinPluginDefinition = {
The key plugin design: **users can toggle enable/disable**, unlike directly registered skills.
### 5.3 Hooks System
### Hooks System
Hooks are **programmable interception points** across the entire lifecycle:
@@ -449,7 +440,7 @@ Hooks execute as shell commands, with exit codes controlling behavior:
- **2**: stderr content shown to model or user
- **Other**: Shown to user only
### 5.4 MCP: Model Context Protocol
### MCP: Model Context Protocol
MCP is the standard protocol for Claude Code's interaction with the external world. Tool naming convention:
@@ -462,11 +453,9 @@ Supported transports: `stdio`, `sse`, `http`, `websocket`, `sdk`
MCP tools are discovered at runtime and **seamlessly merged** into the unified tool pool alongside built-in tools.
---
## Permission & Security Model
## 6. Permission & Security Model
### 6.1 Layered Permission Model
### Layered Permission Model
```
┌─────────────────────────────────────┐
@@ -486,7 +475,7 @@ MCP tools are discovered at runtime and **seamlessly merged** into the unified t
└─────────────────────────────────────┘
```
### 6.2 Permission Decision Flow
### Permission Decision Flow
Permission check for every tool invocation:
@@ -504,7 +493,7 @@ Decision reason traceability:
- `type: 'hook'` — Hook interception
- `type: 'classifier'` — ML classifier decision
### 6.3 Permission Rule Pattern Matching
### Permission Rule Pattern Matching
```javascript
// Exact match
@@ -518,9 +507,7 @@ Decision reason traceability:
{ tool: 'File*', behavior: 'allow' } // Allow all File* tools
```
---
## 7. Fault Recovery Mechanisms
## Fault Recovery Mechanisms
This is one of Claude Code's most sophisticated designs. The core loop in `src/query.ts` has **6 built-in recovery strategies**:
@@ -545,25 +532,23 @@ if (error.type === 'prompt_too_long') {
}
```
### 7.1 Model Fallback
### Model Fallback
When the primary model's stream fails, the system:
1. Cleans up orphaned incomplete messages
2. Switches to a fallback model
3. Retries with the new model
### 7.2 Media Size Recovery
### Media Size Recovery
When images or other media cause token overflow:
- Triggers reactive compaction
- Automatically strips image content
- Retains text information and retries
---
## How It Differs from LangChain/ReAct
## 8. How It Differs from LangChain/ReAct
### 8.1 Architecture Paradigm Comparison
### Architecture Paradigm Comparison
| Dimension | LangChain | Claude Code |
|-----------|-----------|-------------|
@@ -577,7 +562,7 @@ When images or other media cause token overflow:
| **Extension Mechanisms** | Python class inheritance | Skills + Plugins + Hooks + MCP |
| **Caching Strategy** | None | Global / session / per-turn three-level cache |
### 8.2 Why Not ReAct?
### Why Not ReAct?
The ReAct pattern has several inherent limitations:
@@ -593,7 +578,7 @@ Claude Code's Async Generator pattern solves all these problems:
- **Cache optimization**: Static prompts cached globally, dynamic parts minimized
- **Parallel capability**: Read-only tools auto-parallelize, write tools serialize for ordering
### 8.3 Specific Differences from LangChain Agents
### Specific Differences from LangChain Agents
```
LangChain Agent:
@@ -616,7 +601,7 @@ Key differences:
- LangChain requires an OutputParser to parse tool calls from model output
- Claude Code directly uses Anthropic API's native `tool_use` capability — no parsing needed
### 8.4 Comparison with LangGraph
### Comparison with LangGraph
LangGraph is LangChain's evolution, introducing graph structures:
@@ -630,13 +615,11 @@ LangGraph is LangChain's evolution, introducing graph structures:
Claude Code's advantage is **simplicity** — no need to define graph structures; a single while loop handles everything.
---
## 9. Why Claude Code Is So Good
## Why Claude Code Is So Good
From source code analysis, we can distill these core design principles:
### 9.1 Streaming First
### Streaming First
The entire architecture is designed around `AsyncGenerator` — everything is streamed:
- Model responses are streamed
@@ -646,7 +629,7 @@ The entire architecture is designed around `AsyncGenerator` — everything is st
Users **never have to wait** — they see the model thinking, tools executing, and results emerging.
### 9.2 Intelligent Caching
### Intelligent Caching
Three-level prompt caching system (`src/services/api/claude.ts:3213-3237`):
@@ -660,7 +643,7 @@ Section Cache (per-turn) ← systemPromptSection memoization
This dramatically reduces latency and cost for every API call.
### 9.3 Graceful Degradation
### Graceful Degradation
Six recovery strategies ensure Claude Code **almost never interrupts the user's workflow due to technical issues**:
- Token overflow? Auto-compress
@@ -668,7 +651,7 @@ Six recovery strategies ensure Claude Code **almost never interrupts the user's
- Model failure? Fall back to alternate model
- Tool failure? Log error, continue conversation
### 9.4 Minimal Abstraction Principle
### Minimal Abstraction Principle
Unlike LangChain's "abstract everything" philosophy, Claude Code's core has only:
- **One loop** (`while (true)` in `query()`)
@@ -677,7 +660,7 @@ Unlike LangChain's "abstract everything" philosophy, Claude Code's core has only
No Agent → AgentExecutor → Chain → Memory → Callback nesting layers. This makes the code **easy to understand, debug, and extend**.
### 9.5 Native API Integration
### Native API Integration
Claude Code directly leverages Anthropic API's native capabilities:
- **Native tool calling**: No OutputParser needed, directly uses `tool_use` blocks
@@ -687,7 +670,7 @@ Claude Code directly leverages Anthropic API's native capabilities:
This avoids the "framework tax" — the abstraction layer that frameworks like LangChain add between the LLM and the developer.
### 9.6 Tool-Driven Agent
### Tool-Driven Agent
Claude Code's philosophy: **an agent's capability equals the capability of its tools**.
@@ -698,7 +681,7 @@ Claude Code's philosophy: **an agent's capability equals the capability of its t
**All capabilities are exposed through the unified tool interface**, and the model uses natural language reasoning to decide which tool to use. No explicit orchestration logic needed — the model itself is the orchestrator.
### 9.7 Deep Developer Experience Integration
### Deep Developer Experience Integration
Claude Code isn't "generic agent + code plugin" — it's **deeply optimized for coding scenarios from the ground up**:
@@ -708,11 +691,7 @@ Claude Code isn't "generic agent + code plugin" — it's **deeply optimized for
- **LSP integration**: Language Server Protocol provides type information and diagnostics
- **MCP ecosystem**: Connects to various external tools via standard protocol
---
## 10. Architecture Summary
### Core Component Relationships
## Core Component Relationships
```
User Input
@@ -742,23 +721,17 @@ query() async generator loop (src/query.ts)
└─ Teammate (mailbox communication)
```
### One-Line Summary
> **Claude Code's agent framework is a streaming state machine powered by AsyncGenerator, exposing all capabilities through a unified tool interface, combined with four-level context compression, three-level prompt caching, and six fault recovery strategies — an AI system that autonomously completes complex programming tasks without explicit orchestration.**
---
## 11. Key Source File Index
## Key Source File Index
| Component | File Path | Description |
|-----------|-----------|-------------|
| Core Loop | `src/query.ts` | Main agent loop (~1730 lines) |
| Query Engine | `src/QueryEngine.ts` | High-level wrapper (~687 lines) |
| Query Engine | `src/QueryEngine.ts` | High-level wrapper |
| Tool Definition | `src/Tool.ts` | Tool type system (~792 lines) |
| Tool Registry | `src/tools.ts` | Tool discovery and registration (~389 lines) |
| Tool Execution | `src/services/tools/toolExecution.ts` | Execution pipeline (~1500 lines) |
| Tool Execution | `src/services/tools/toolExecution.ts` | Execution pipeline |
| Tool Orchestration | `src/services/tools/toolOrchestration.ts` | Parallel/serial strategy |
| System Prompt | `src/constants/prompts.ts` | Prompt assembly (~577 lines) |
| System Prompt | `src/constants/prompts.ts` | Prompt assembly |
| Prompt Sections | `src/constants/systemPromptSections.ts` | Section caching |
| Context Management | `src/context.ts` | System/user context |
| CLAUDE.md | `src/utils/claudemd.ts` | User instruction loading |
@@ -780,11 +753,9 @@ query() async generator loop (src/query.ts)
| Remote Sessions | `src/remote/RemoteSessionManager.ts` | CCR connection management |
| Bridge | `src/bridge/bridgeMain.ts` | Remote bridge |
---
## Further Reading
## 12. Further Reading
- [Usage Guide](./01-usage-guide.md) — User-facing multi-agent manual
- [Implementation Details](./02-implementation.md) — Technical deep dive into multi-agent orchestration
- [Usage Guide](./agent.md) — User-facing multi-agent manual
- [Implementation Details](./agent-internals.md) — Technical deep dive into multi-agent orchestration
- [Anthropic API Docs](https://docs.anthropic.com/) — Native API capabilities
- [MCP Protocol Spec](https://modelcontextprotocol.io/) — Model Context Protocol
@@ -1,20 +1,21 @@
# Claude Code Multi-Agent System — Implementation Details
---
title: Multi-Agent Internals
nav_title: Multi-Agent Internals
description: Spawn paths, tool pool filtering, context passing, and the background task engine.
order: 5
---
> A deep dive into the architecture, spawn flow, context passing, and collaboration mechanisms of multi-agent orchestration.
# Multi-Agent Internals
<p align="center">
<a href="#1-architecture-overview">Architecture</a> · <a href="#2-agent-spawn-flow--four-paths">Spawn Flow</a> · <a href="#3-tool-pool-system--three-layer-filtering">Tool Pool</a> · <a href="#4-context-passing-mechanism">Context Passing</a> · <a href="#5-agent-teams-internals">Teams Internals</a> · <a href="#6-background-task-engine">Task Engine</a> · <a href="#7-dreamtask--automatic-memory-consolidation">DreamTask</a> · <a href="#8-worktree-isolation-implementation">Worktree Isolation</a> · <a href="#9-permission-synchronization">Permission Sync</a> · <a href="#10-agent-lifecycle-end-to-end-data-flow">Lifecycle Data Flow</a> · <a href="#11-key-source-file-index">Source Index</a> · <a href="#12-feature-flags">Feature Flags</a>
</p>
A deep dive into the architecture, spawn flow, context passing, and collaboration mechanisms of multi-agent orchestration.
![Implementation Architecture Overview](./images/05-architecture.png)
---
## 1. Architecture Overview
## Architecture Overview
Claude Code's multi-agent system consists of the following core modules:
### 5 Core Modules
### Core Modules
| Module | Responsibility | Key Files |
|--------|---------------|-----------|
@@ -24,7 +25,7 @@ Claude Code's multi-agent system consists of the following core modules:
| **Task System** | State tracking, progress updates, notification queue | `src/tasks/LocalAgentTask/` |
| **Swarm Infrastructure** | Team management, mailbox communication, permission sync | `src/utils/swarm/` |
### 5 Agent Categories
### Agent Categories
```
┌─────────────────────────────────────────────────┐
@@ -50,9 +51,7 @@ Claude Code's multi-agent system consists of the following core modules:
└───────────┘
```
---
## 2. Agent Spawn Flow — Four Paths
## Agent Spawn Flow — Four Paths
### Entry Point: `AgentTool.call()`
@@ -200,9 +199,7 @@ runAgent(promptMessages, toolUseContext, options)
└─ totalTokens
```
---
## 3. Tool Pool System — Three-Layer Filtering
## Tool Pool System — Three-Layer Filtering
![Tool Pool System](./images/07-tool-pool.png)
@@ -271,9 +268,7 @@ All available tools
└─ Final tool pool
```
---
## 4. Context Passing Mechanism
## Context Passing Mechanism
![Context Passing](./images/06-context-passing.png)
@@ -393,9 +388,7 @@ agentDefinition.model ← Agent definition
'inherit' ← Inherit parent model
```
---
## 5. Agent Teams Internals
## Agent Teams Internals
### TeamFile Structure
@@ -516,9 +509,7 @@ SendMessage({ to, message })
└─ sendToUdsSocket(socketPath, message)
```
---
## 6. Background Task Engine
## Background Task Engine
![Background Task Engine](./images/08-background-task.png)
@@ -618,9 +609,7 @@ function enqueuePendingNotification(taskId, result) {
| Write method | Asynchronous queue writes to prevent memory buildup |
| Safety measure | O_NOFOLLOW to prevent symlink attacks |
---
## 7. DreamTask — Automatic Memory Consolidation
## DreamTask — Automatic Memory Consolidation
DreamTask is a special background agent used for cross-session memory consolidation.
@@ -649,9 +638,7 @@ type DreamTaskState = {
}
```
---
## 8. Worktree Isolation Implementation
## Worktree Isolation Implementation
### Creation Flow
@@ -684,9 +671,7 @@ async function createAgentWorktree(slug) {
- If no changes: automatically deletes the worktree (`removeAgentWorktree()`)
- On abnormal exit: cleanup is ensured via `registerTeamForSessionCleanup()`
---
## 9. Permission Synchronization
## Permission Synchronization
### Team-Level Permissions
@@ -728,9 +713,7 @@ Teammate needs permission
└─ Handled by Leader's useSwarmPermissionPoller
```
---
## 10. Agent Lifecycle End-to-End Data Flow
## Agent Lifecycle End-to-End Data Flow
```
1. User triggers Agent Tool
@@ -778,9 +761,7 @@ Teammate needs permission
└─ Asynchronous: processes after receiving <task-notification>
```
---
## 11. Key Source File Index
## Key Source File Index
### Agent Tool Core
@@ -840,9 +821,7 @@ Teammate needs permission
|------|---------------|
| `src/coordinator/coordinatorMode.ts` | Coordinator mode configuration |
---
## 12. Feature Flags
## Feature Flags
| Flag | Controls |
|------|----------|
@@ -1,16 +1,17 @@
# Claude Code Multi-Agent System — Usage Guide
---
title: Multi-Agent Usage Guide
nav_title: Multi-Agent Usage
description: Built-in agents, spawn options, background tasks, Agent Teams, and custom agent definitions.
order: 4
---
> Let Claude Code orchestrate multiple specialized agents to handle complex tasks in parallel.
# Multi-Agent Usage Guide
<p align="center">
<a href="#1-what-is-the-multi-agent-system">Multi-Agent System</a> · <a href="#2-six-built-in-agents">Six Built-in Agents</a> · <a href="#3-how-to-spawn-agents">Spawning Agents</a> · <a href="#4-background-task-management">Background Tasks</a> · <a href="#5-agent-teams--multi-agent-collaboration">Agent Teams</a> · <a href="#6-custom-agents">Custom Agents</a> · <a href="#7-permission-modes">Permission Modes</a> · <a href="#8-quick-reference">Quick Reference</a>
</p>
Let Claude Code orchestrate multiple specialized agents to handle complex tasks in parallel.
![Multi-Agent System Overview](./images/01-agent-overview.png)
---
## 1. What Is the Multi-Agent System?
## What Is the Multi-Agent System?
Claude Code's multi-agent system is an **intelligent task orchestration framework** that enables the primary agent to spawn multiple specialized subagents, each executing different tasks independently, then aggregating results for the user.
@@ -23,15 +24,13 @@ Core philosophy: **Break large tasks into specialized subtasks, execute them in
| Code review | Single-threaded, file by file | Multiple reviewers in parallel |
| Debug a complex bug | Try one hypothesis at a time | Multiple debuggers verify in parallel |
---
## 2. Six Built-in Agents
## Six Built-in Agents
![Six Built-in Agents](./images/02-agent-types.png)
Claude Code ships with 6 specialized agent types, each with a specific tool pool and intended use case:
### 2.1 general-purpose (General Agent)
### general-purpose (General Agent)
**Use case**: Complex multi-step research, code search, tasks requiring full tool access.
@@ -47,7 +46,7 @@ Agent({
- **Model**: Inherited from parent
- **Characteristics**: The all-rounder — choose this when you are unsure which agent type to use
### 2.2 Explore (Exploration Agent)
### Explore (Exploration Agent)
**Use case**: Quickly search files, find code patterns, answer questions about codebase structure.
@@ -59,11 +58,11 @@ Agent({
})
```
- **Tool pool**: Read-only tools (Glob, Grep, Read, Bash)
- **Tool pool**: Every tool except Agent, ExitPlanMode, Edit, Write, and NotebookEdit. **Bash is in that set** — it cannot change files, but it can run any command
- **Model**: Haiku (fast, low cost)
- **Characteristics**: Cannot modify files; fast; ideal for research
### 2.3 Plan (Planning Agent)
### Plan (Planning Agent)
**Use case**: Design implementation plans, analyze architectural trade-offs, generate step-by-step plans.
@@ -75,11 +74,11 @@ Agent({
})
```
- **Tool pool**: Read-only tools (same as Explore)
- **Tool pool**: Same denylist as Explore
- **Model**: Inherited from parent (requires strong reasoning)
- **Characteristics**: Outputs structured plans including key files and dependency analysis
### 2.4 verification (Verification Agent)
### verification (Verification Agent)
**Use case**: Independently verify that an implementation is correct, run tests, perform boundary checks.
@@ -91,11 +90,11 @@ Agent({
})
```
- **Tool pool**: Read-only tools
- **Tool pool**: Same as Explore; its system prompt additionally allows writing throwaway test scripts under tmp
- **Model**: Inherited from parent
- **Characteristics**: Always runs in the background; outputs PASS/FAIL/PARTIAL verdicts; displayed with a red badge
### 2.5 claude-code-guide (Guide Agent)
### claude-code-guide (Guide Agent)
**Use case**: Answer questions about Claude Code, Agent SDK, or the Claude API.
@@ -107,11 +106,11 @@ Agent({
})
```
- **Tool pool**: Bash, Read, WebFetch, WebSearch
- **Tool pool**: Glob, Grep, Read, WebFetch, WebSearch (internal builds swap in Bash for Glob/Grep)
- **Model**: Haiku
- **Characteristics**: Focused on documentation queries; uses the dontAsk permission mode
### 2.6 statusline-setup (Status Bar Configuration Agent)
### statusline-setup (Status Bar Configuration Agent)
**Use case**: Configure the Claude Code status bar display.
@@ -124,15 +123,13 @@ Agent({
| Agent | Access | Tool Pool | Model | Purpose |
|-------|--------|-----------|-------|---------|
| general-purpose | Read/Write | All | Inherited | General tasks |
| Explore | Read-only | Search + Read | Haiku | Quick exploration |
| Plan | Read-only | Search + Read | Inherited | Architecture planning |
| verification | Read-only | Search + Read | Inherited | Independent verification |
| Explore | No file edits | All tools minus the edit set | Haiku | Quick exploration |
| Plan | No file edits | Same as Explore | Inherited | Architecture planning |
| verification | No project-file edits | Same as Explore | Inherited | Independent verification |
| claude-code-guide | Read-only | Search + Web | Haiku | Documentation guide |
| statusline-setup | Read/Write | Read + Edit | Sonnet | Status bar config |
---
## 3. How to Spawn Agents
## How to Spawn Agents
### Parameters
@@ -208,9 +205,7 @@ Agent({
- If changes were made, returns the worktree path and branch name on completion
- If no changes were made, cleans up automatically
---
## 4. Background Task Management
## Background Task Management
![Agent Spawn Flow](./images/03-spawn-flow.png)
@@ -251,9 +246,7 @@ When a background agent finishes, the primary agent receives an XML-formatted no
When the `tengu_auto_background_agents` feature flag is enabled, foreground agents that run for more than **120 seconds** are automatically moved to background execution, freeing the primary agent to continue working.
---
## 5. Agent Teams — Multi-Agent Collaboration
## Agent Teams — Multi-Agent Collaboration
![Agent Teams Collaboration](./images/04-agent-teams.png)
@@ -344,9 +337,7 @@ Agent Teams supports two execution backends:
| **tmux** | Runs in a separate tmux pane | When an independent terminal view is needed |
| **iTerm2** | Runs in a separate iTerm2 window | For macOS iTerm2 users |
---
## 6. Custom Agents
## Custom Agents
In addition to built-in agents, you can create your own specialized agents.
@@ -447,9 +438,7 @@ When several sources define an Agent with the same name, cc-haha selects the act
The desktop app still shows overridden definitions and their sources, but spawning uses the highest-priority active definition.
---
## 7. Permission Modes
## Permission Modes
Each agent can be configured with a different permission mode:
@@ -463,9 +452,7 @@ Each agent can be configured with a different permission mode:
| `auto` | AI-driven permission classification (Anthropic internal only) |
| `bubble` | Permission prompts bubble up to the parent agent's terminal |
---
## 8. Quick Reference
## Quick Reference
| Action | Method |
|--------|--------|
@@ -1,16 +1,17 @@
# Claude Code Memory System — AutoDream Memory Consolidation
---
title: AutoDream Memory Consolidation
nav_title: AutoDream
description: The four-phase background pass that reviews recent sessions and prunes memories.
order: 11
---
> Claude "dreams" -- silently reviewing recent sessions in the background to consolidate, update, and prune memories, much like the human brain organizes memories during sleep.
# AutoDream Memory Consolidation
<p align="center">
<a href="#1-what-is-autodream">AutoDream</a> · <a href="#2-trigger-conditions">Trigger Conditions</a> · <a href="#3-four-phase-consolidation-process">Consolidation Process</a> · <a href="#4-security-restrictions">Security</a> · <a href="#5-ui-presentation">UI</a> · <a href="#6-configuration-and-toggles">Configuration</a> · <a href="#7-relationship-with-extractmemories">Comparison</a> · <a href="#8-source-code-navigation">Source Code</a>
</p>
Claude "dreams" -- silently reviewing recent sessions in the background to consolidate, update, and prune memories, much like the human brain organizes memories during sleep.
![AutoDream Overview](./images/11-autodream-overview.png)
---
## 1. What Is AutoDream?
## What Is AutoDream?
AutoDream is Claude Code's **background memory consolidation mechanism**, internally codenamed **"Dream: Memory Consolidation"**.
@@ -21,7 +22,7 @@ Core metaphor:
| Jotting down notes throughout the day | `extractMemories` -- extracts new memories after each conversation |
| Organizing the notebook while sleeping | `autoDream` -- periodically reviews multiple sessions to consolidate all memories |
When you're inactive (default interval: 24 hours with 5 accumulated sessions), Claude silently launches a **"dreaming" sub-agent** (forked subagent) in the background. It reviews all recent session transcripts and consolidates scattered memories into structured, deduplicated, up-to-date persistent knowledge.
When you're inactive (default interval: 24 hours with 5 accumulated sessions), Claude silently launches a **"dreaming" subagent** (forked subagent) in the background. It reviews all recent session transcripts and consolidates scattered memories into structured, deduplicated, up-to-date persistent knowledge.
**Key source**: `src/services/autoDream/autoDream.ts`
@@ -30,9 +31,7 @@ When you're inactive (default interval: 24 hours with 5 accumulated sessions), C
// subagent when time-gate passes AND enough sessions have accumulated.
```
---
## 2. Trigger Conditions
## Trigger Conditions
![AutoDream Trigger Flow](./images/12-autodream-trigger.png)
@@ -85,15 +84,13 @@ AutoDream will not trigger in the following situations:
- **Remote mode**: `getIsRemoteMode() === true`
- **autoMemory not enabled**
- **`--bare` / SIMPLE mode**
- **Inside a sub-agent**: Only the main agent triggers it
- **Inside a subagent**: Only the main agent triggers it
---
## 3. Four-Phase Consolidation Process
## Four-Phase Consolidation Process
![AutoDream Four Phases](./images/13-autodream-phases.png)
Once all gates pass, AutoDream launches a **forked sub-agent** that operates according to the 4-phase prompt defined in `consolidationPrompt.ts`:
Once all gates pass, AutoDream launches a **forked subagent** that operates according to the 4-phase prompt defined in `consolidationPrompt.ts`:
### Phase 1 -- Orient
@@ -132,11 +129,9 @@ Don't exhaustively read transcript files. Only look for content you already susp
**Key source**: `src/services/autoDream/consolidationPrompt.ts:10-64`
---
## Security Restrictions
## 4. Security Restrictions
The AutoDream sub-agent operates under strict tool permission constraints:
The AutoDream subagent operates under strict tool permission constraints:
### Bash: Read-Only Only
@@ -172,9 +167,7 @@ Uses a `.consolidate-lock` file for process-level mutual exclusion:
**Key source**: `src/services/autoDream/consolidationLock.ts`
---
## 5. UI Presentation
## UI Presentation
### Bottom Status Bar
@@ -213,9 +206,7 @@ appendSystemMessage({
- `src/tasks/DreamTask/DreamTask.ts` -- Task state management
- `src/components/tasks/DreamDetailDialog.tsx` -- UI component
---
## 6. Configuration and Toggles
## Configuration and Toggles
### settings.json
@@ -255,9 +246,7 @@ The `tengu_onyx_plover` feature flag can configure:
In addition to automatic triggering, users can manually trigger memory consolidation via the `/dream` command. Manual triggers call `recordConsolidation()` to update the lock file timestamp.
---
## 7. Relationship with extractMemories
## Relationship with extractMemories
| Dimension | extractMemories | autoDream |
|-----------|----------------|-----------|
@@ -286,9 +275,7 @@ Merge duplicates / Fix stale data / Delete conflicts / Compress index
MEMORY.md and topic files refreshed
```
---
## 8. Source Code Navigation
## Source Code Navigation
| File | Responsibility |
|------|----------------|
@@ -303,9 +290,7 @@ MEMORY.md and topic files refreshed
| `src/utils/backgroundHousekeeping.ts` | Initialization entry point: `initAutoDream()` |
| `src/services/extractMemories/extractMemories.ts` | Shared `createAutoMemCanUseTool` |
---
## 9. Analytics Events
## Analytics Events
AutoDream records its operational state through the following events:
@@ -1,25 +1,17 @@
# Channel System Architecture
---
title: Channel System
nav_title: Channel System
description: The message protocol, six-layer access control, and permission relay behind IM remote control.
order: 13
---
> A deep dive into how Claude Code enables remote Agent control via IM platforms
# Channel System
<p align="center">
<a href="#1-what-is-a-channel">Concepts</a> ·
<a href="#2-architecture-overview">Architecture</a> ·
<a href="#3-message-protocol">Protocol</a> ·
<a href="#4-six-layer-access-control">Access Control</a> ·
<a href="#5-permission-relay-system">Permission Relay</a> ·
<a href="#6-ui-components">UI</a> ·
<a href="#7-plugin-channel-architecture">Plugins</a> ·
<a href="#8-security-design">Security</a> ·
<a href="#9-command-line-interface">CLI</a> ·
<a href="#10-feature-flags-and-analytics">Feature Flags</a>
</p>
A deep dive into how Claude Code enables remote Agent control via IM platforms
![Channel System Overview](./images/01-channel-overview.png)
---
## 1. What is a Channel
## What is a Channel
A Channel is Claude Code's **IM integration system** that allows users to remotely control a running Claude Code Agent through instant messaging platforms such as Telegram, Feishu (Lark), Discord, and Slack.
@@ -45,9 +37,7 @@ type ChannelEntry =
**Plugin kind**: Verified plugins from a marketplace (e.g., `plugin:telegram@anthropic`)
**Server kind**: Directly specified MCP server names (always requires dev bypass)
---
## 2. Architecture Overview
## Architecture Overview
![Message Flow](./images/02-message-flow.png)
@@ -107,11 +97,9 @@ The Channel system follows a clear bidirectional message path:
| **Plugin Integration** | `mcpPluginIntegration.ts` | Plugin scoped naming |
| **State** | `bootstrap/state.ts` | Global channel allowlist state |
---
## Message Protocol
## 3. Message Protocol
### 3.1 Inbound Notification Schema
### Inbound Notification Schema
The notification format Channel Servers push to Claude Code:
@@ -128,7 +116,7 @@ const ChannelMessageNotificationSchema = z.object({
})
```
### 3.2 XML Wrapping
### XML Wrapping
After receiving the notification, the system wraps it in a `<channel>` XML tag:
@@ -158,7 +146,7 @@ Can you check what's wrong with main.ts?
When the model sees this tag, it knows the message came from Telegram user "alice" and will use Telegram's `reply` tool to respond.
### 3.3 Safe Metadata Filtering
### Safe Metadata Filtering
Meta keys become XML attribute names. A crafted key like `x="" injected="y` could break out of the attribute structure. The system uses strict regex filtering:
@@ -169,7 +157,7 @@ const SAFE_META_KEY = /^[a-zA-Z_][a-zA-Z0-9_]*$/
In practice, channel servers only send safe keys like `chat_id`, `user`, `thread_ts`, `message_id`.
### 3.4 Message Enqueuing
### Message Enqueuing
Wrapped messages are pushed into the message queue:
@@ -186,9 +174,7 @@ enqueue({
SleepTool polls `hasCommandsInQueue()` every ~1 second, waking the Agent when new messages arrive.
---
## 4. Six-Layer Access Control
## Six-Layer Access Control
![Access Control](./images/03-access-control.png)
@@ -205,7 +191,7 @@ function gateChannelServer(
): ChannelGateResult // { action: 'register' } | { action: 'skip', kind, reason }
```
### 4.1 Layer 1: Capability Declaration
### Layer 1: Capability Declaration
```typescript
if (!capabilities?.experimental?.['claude/channel']) {
@@ -216,7 +202,7 @@ if (!capabilities?.experimental?.['claude/channel']) {
The MCP Server must declare `experimental['claude/channel']: {}` during handshake. This is MCP's "presence signal" idiom (similar to `tools: {}`), separating Channel Servers from ordinary MCP Servers.
### 4.2 Layer 2: Runtime Gate
### Layer 2: Runtime Gate
```typescript
if (!isChannelsEnabled()) {
@@ -227,7 +213,7 @@ if (!isChannelsEnabled()) {
`isChannelsEnabled()` checks GrowthBook feature flag `tengu_harbor` (default false, 5-minute refresh). This is the global "emergency brake" — flipping this switch immediately disables all Channels without a release.
### 4.3 Layer 3: OAuth Authentication
### Layer 3: OAuth Authentication
```typescript
if (!getClaudeAIOAuthTokens()?.accessToken) {
@@ -238,7 +224,7 @@ if (!getClaudeAIOAuthTokens()?.accessToken) {
Channels are restricted to OAuth-authenticated users. API key users are blocked because Console doesn't have a `channelsEnabled` admin surface yet.
### 4.4 Layer 4: Organization Policy
### Layer 4: Organization Policy
```typescript
const sub = getSubscriptionType()
@@ -252,7 +238,7 @@ if (managed && policy?.channelsEnabled !== true) {
Teams/Enterprise organizations must explicitly enable `channelsEnabled: true` in managed settings. Default is OFF — even a team org with zero configured policy keys is still considered managed and does not fall through to the unmanaged path.
### 4.5 Layer 5: Session Allowlist
### Layer 5: Session Allowlist
```typescript
const entry = findChannelEntry(serverName, getAllowedChannels())
@@ -264,7 +250,7 @@ if (!entry) {
The MCP Server must be in the current session's `--channels` parameter list. Even if a trusted server dynamically adds the `claude/channel` capability, it cannot bypass this — the user must explicitly list it at startup.
### 4.6 Layer 6: Marketplace Verification + Allowlist
### Layer 6: Marketplace Verification + Allowlist
For plugin-kind channels:
@@ -303,19 +289,17 @@ type ChannelGateResult =
// kind enum: capability | disabled | auth | policy | session | marketplace | allowlist
```
---
## 5. Permission Relay System
## Permission Relay System
![Permission Relay](./images/04-permission-relay.png)
### 5.1 Why Permission Relay Exists
### Why Permission Relay Exists
When Claude Code needs to execute sensitive operations (like running a Bash command), it shows a permission confirmation dialog. But if the user is controlling the Agent remotely via Telegram, they can't see the local terminal dialog.
The permission relay system solves this: **forward permission prompts to the IM platform so users can approve or deny operations from their phone**.
### 5.2 Outbound: CC → Channel (Permission Request)
### Outbound: CC → Channel (Permission Request)
When the Agent triggers a permission dialog and a Channel has declared `experimental['claude/channel/permission']` capability:
@@ -334,7 +318,7 @@ type ChannelPermissionRequestParams = {
The Channel Server formats this for its platform (Telegram markdown, Discord embed, etc.) and sends it to the user.
### 5.3 Short Request ID Generation
### Short Request ID Generation
The 5-letter identifier design is thoughtfully crafted:
@@ -359,7 +343,7 @@ function shortRequestId(toolUseID: string): string {
- **Letters only**: Phone users don't need to switch keyboard modes (hex alternates between letters and digits)
- **Case insensitive**: Accommodates phone autocorrect
### 5.4 Inbound: Channel → CC (Permission Response)
### Inbound: Channel → CC (Permission Response)
Users reply in IM with format: `yes tbxkq` or `no tbxkq`
@@ -379,7 +363,7 @@ const ChannelPermissionNotificationSchema = z.object({
**Key design**: The Channel Server is responsible for parsing the user's reply and emitting a structured event — CC never does text regex matching. This means casual conversation text can never accidentally trigger permission approval.
### 5.5 Multi-Source Racing
### Multi-Source Racing
Permission responses come from four sources, first to resolve wins:
@@ -401,7 +385,7 @@ Permission responses come from four sources, first to resolve wins:
`createChannelPermissionCallbacks()` uses a closure to maintain a pending Map (not at module level, not in AppState — function references in state cause serialization issues), constructed once and stored in AppState.
### 5.6 Filtering Permission Relay Clients
### Filtering Permission Relay Clients
```typescript
// channelPermissions.ts:177-194
@@ -417,11 +401,9 @@ function filterPermissionRelayClients(clients, isInAllowlist) {
**All three conditions required**: connected + in allowlist + declares BOTH capabilities (`claude/channel` AND `claude/channel/permission`). The second capability is explicit opt-in — a relay-only Channel never accidentally becomes a permission approval surface.
---
## UI Components
## 6. UI Components
### 6.1 Terminal Message Rendering (UserChannelMessage)
### Terminal Message Rendering (UserChannelMessage)
```typescript
// UserChannelMessage.tsx
@@ -447,7 +429,7 @@ const TRUNCATE_AT = 60 // Message body truncation length
Where `◁` is the Channel arrow symbol (`CHANNEL_ARROW`), showing the server leaf name and optional username.
### 6.2 Status Notice (ChannelsNotice)
### Status Notice (ChannelsNotice)
Shows the status of `--channels` entries at startup, reporting blockers:
@@ -458,7 +440,7 @@ Shows the status of `--channels` entries at startup, reporting blockers:
| `policyBlocked` | Org policy hasn't enabled channels |
| `unmatched` | Not matched in `--channels` list |
### 6.3 Developer Confirmation Dialog (DevChannelsDialog)
### Developer Confirmation Dialog (DevChannelsDialog)
Warning dialog shown when using `--dangerously-load-development-channels`:
@@ -480,11 +462,9 @@ Warning dialog shown when using `--dangerously-load-development-channels`:
Selecting "Exit" calls `gracefulShutdownSync(1)` for immediate exit.
---
## Plugin Channel Architecture
## 7. Plugin Channel Architecture
### 7.1 Channel Declaration in Plugin Manifest
### Channel Declaration in Plugin Manifest
Plugins declare channels via the `channels` array in `plugin.json`:
@@ -539,7 +519,7 @@ const PluginManifestChannelsSchema = z.object({
}
```
### 7.2 Configuration Flow (PluginOptionsFlow)
### Configuration Flow (PluginOptionsFlow)
After enabling a plugin, if a Channel has unconfigured `userConfig` fields:
@@ -571,7 +551,7 @@ function getUnconfiguredChannels(plugin: LoadedPlugin): UnconfiguredChannel[] {
}
```
### 7.3 Scoped Naming
### Scoped Naming
Plugin-provided MCP Servers get a scope prefix to avoid naming conflicts:
@@ -601,7 +581,7 @@ function addPluginScopeToServers(
`pluginSource` is preserved on the config for later marketplace verification in the Channel Gate.
### 7.4 Effective Allowlist Source
### Effective Allowlist Source
```typescript
// channelNotification.ts:127-138
@@ -617,18 +597,16 @@ function getEffectiveChannelAllowlist(sub, orgList) {
Organization admins can fully control which Channel plugins are trusted via `allowedChannelPlugins`, independent of the global ledger.
---
## Security Design
## 8. Security Design
### 8.1 XML Injection Prevention
### XML Injection Prevention
Channel message metadata becomes XML attributes. Two lines of defense:
1. **Key name filtering**: `SAFE_META_KEY = /^[a-zA-Z_][a-zA-Z0-9_]*$/` — only plain identifiers
2. **Value escaping**: `escapeXmlAttr()` applies XML escaping to attribute values
### 8.2 Marketplace Verification
### Marketplace Verification
`--channels plugin:slack@anthropic` is merely the user's "intent declaration." The runtime name `plugin:slack:X` could come from `slack@anthropic` or `slack@evil`. The gate verifies they must match:
@@ -641,25 +619,23 @@ if (actual !== entry.marketplace) {
}
```
### 8.3 Trust Boundary of Permission Relay
### Trust Boundary of Permission Relay
From Kenneth's analysis in code comments (PR discussion #2956440848):
> "Would this let Claude self-approve?" Answer: the approving party is the human via the channel, not Claude. But the trust boundary isn't the terminal — it's the allowlist (tengu_harbor_ledger). A compromised channel server CAN fabricate "yes \<id\>" without the human seeing the prompt. Accepted risk: a compromised channel already has unlimited conversation-injection turns (social-engineer over time, wait for acceptEdits, etc.); inject-then-self-approve is faster, not more capable. The dialog slows a compromised channel; it doesn't stop one.
### 8.4 skipSlashCommands
### skipSlashCommands
Channel messages are enqueued with `skipSlashCommands: true`, ensuring text like `/help` sent by IM users is not interpreted as Claude Code slash commands.
### 8.5 Dev Bypass Granularity
### Dev Bypass Granularity
The `dev` flag from `--dangerously-load-development-channels` is **per-entry**, not global. After accepting the dev dialog, only explicitly dev-flagged entries bypass the allowlist — normal `--channels` entries still go through full allowlist verification.
---
## Command-Line Interface
## 9. Command-Line Interface
### 9.1 Startup Parameters
### Startup Parameters
```bash
# Use approved Channel plugins
@@ -673,7 +649,7 @@ claude --channels plugin:telegram@anthropic \
--dangerously-load-development-channels plugin:dev-channel@local
```
### 9.2 Argument Parsing
### Argument Parsing
```typescript
// main.tsx
@@ -687,7 +663,7 @@ const parseChannelEntries = (raw: string[], flag: string): ChannelEntry[] => {
setAllowedChannels(channelEntries)
```
### 9.3 Feature Gating
### Feature Gating
These CLI options are only available when feature flags are enabled:
@@ -701,11 +677,9 @@ if (feature('KAIROS') || feature('KAIROS_CHANNELS')) {
`hideHelp()` means these options don't appear in `--help` output — the Channel feature is currently in hidden feature stage.
---
## Feature Flags and Analytics
## 10. Feature Flags and Analytics
### 10.1 Feature Flags
### Feature Flags
| Flag | Source | Purpose |
|------|--------|---------|
@@ -714,7 +688,7 @@ if (feature('KAIROS') || feature('KAIROS_CHANNELS')) {
| `tengu_harbor_ledger` | GrowthBook runtime | Approved plugin allowlist |
| `tengu_harbor_permissions` | GrowthBook runtime | Permission relay feature switch |
### 10.2 Analytics Events
### Analytics Events
```typescript
// Channel Gate result
@@ -742,33 +716,7 @@ tengu_mcp_channel_flags: {
}
```
---
## 11. Summary
The Channel system is Claude Code's **IM integration framework**, and its design embodies several core principles:
### 1. Security First
Six access control layers ensure only approved, user-explicitly-chosen Channels can push messages. From build-time feature flags to runtime allowlists, each layer can independently interrupt the flow.
### 2. Protocol-Driven
Channels aren't special code — they're ordinary MCP Servers with an additional notification protocol. This means any language that can implement MCP can write Channel plugins.
### 3. Loose Coupling
Channel failures don't block local workflows. Permission relay is a multi-source racing mechanism where any source's response is valid.
### 4. Progressive Trust
From global switch → authentication → org policy → session allowlist → marketplace verification → allowlist, trust levels increase progressively, with each step serving a clear security purpose.
### 5. Plugin-Friendly
Through declarative Channel configuration in `plugin.json`, automatic user config prompting, and scoped naming, third-party developers can easily build their own Channel plugins.
### Source File Index
## Source File Index
| File | Lines | Responsibility |
|------|-------|---------------|
@@ -1,6 +1,13 @@
---
title: Computer Use Architecture
nav_title: Computer Use
description: Tool registration, coordinate mapping, authorization gates, the Python bridge, and safety limits.
order: 12
---
# Computer Use Architecture
> This page describes the Computer Use tools, authorization model, platform executors, and current safety boundaries. For setup, start with the [Computer Use guide](./computer-use.md).
Between the model asking for a click and the mouse actually moving sit four gates: tool definitions, authorization checks, the Python bridge, and a platform helper. This page takes each one apart. For setup and everyday usage, start with the [Computer Use guide](../desktop/computer-use.md).
## Layers
@@ -1,4 +1,11 @@
# Contributing and Local Quality Gates
---
title: Contributing and Quality Gates
nav_title: Contributing
description: Local setup, impact checks, quality gates, testing requirements, and the PR workflow.
order: 14
---
# Contributing and Quality Gates
This guide explains how to install, develop, test, and run the local quality gates before opening a PR. The goal is to help maintainers and contributors answer one question before review: did this change break the core Coding Agent workflow?
@@ -236,6 +243,67 @@ Release mode composes PR checks, baseline catalog validation, live baseline case
In release mode, live lanes are not allowed to be silently skipped. Missing providers, model quota, or external account access will fail the gate and must be recorded as a release blocker.
## Releases and Auto-Update
`desktop/package.json` is the single source of the desktop version number. A real release requires the version, the Git tag, and `release-notes/vX.Y.Z.md` to match exactly.
In-app updates are driven by `electron-updater`, with artifacts hosted on GitHub Releases:
| Platform | Install / update target | Metadata |
|---|---|---|
| macOS arm64 / x64 | `dmg` for first install, `zip` for Squirrel.Mac updates | `latest-mac.yml` |
| Windows x64 / ARM64 | NSIS `.exe` | `latest.yml` |
| Linux x64 | `.AppImage` for updates, `.deb` for manual install | `latest-linux.yml` |
| Linux arm64 | `.AppImage` for updates, `.deb` for manual install | `latest-linux-arm64.yml` |
The release workflow generates `latest*.yml` inside each platform matrix job, renames colliding metadata to `latest-<platform>.yml`, and finally lets `scripts/release-update-metadata.ts` merge them back into the standard filenames electron-updater expects. Do not change this so each matrix job publishes the GitHub Release directly — the metadata files would overwrite each other.
### Signing secrets
macOS signing and notarization depend on these GitHub Actions repository secrets:
```text
MACOS_CERTIFICATE
MACOS_CERTIFICATE_PASSWORD
APPLE_ID
APPLE_APP_SPECIFIC_PASSWORD
APPLE_TEAM_ID
```
`MACOS_CERTIFICATE` is the base64 content of a Developer ID Application `.p12`. The project does not ship a `.pkg`, so no Developer ID Installer certificate is needed.
Windows signing is optional:
```text
WINDOWS_CERTIFICATE
WINDOWS_CERTIFICATE_PASSWORD
```
Auto-update still works without Windows signing; users may just see a SmartScreen prompt.
### Pre-release checks
```bash
bun run scripts/release.ts <version> --dry
bun test scripts/pr/release-workflow.test.ts scripts/release-update-metadata.test.ts scripts/quality-gate/package-smoke/index.test.ts
bun run check:policy
```
Confirm `release-notes/v<version>.md` exists before running `bun run scripts/release.ts <version>` for real.
### Verify one real update path
Every release should be verified by upgrading from the previous stable build at least once:
1. Install the previous stable release from GitHub Releases.
2. Push the tag and let the `Release Desktop` workflow finish green.
3. Open the old build and wait for the startup check, or check for updates manually in settings.
4. Confirm the new version is offered, then install and restart.
5. After restart, confirm the version in About, and that providers, sessions, skills, agents, memories, custom pets, and a custom data directory all still work.
6. Confirm historical attachment context, subagent details, and task state restore correctly; open a pet window and check the overlay and current-session navigation.
Platforms differ in what matters: on macOS confirm the release job used the signed artifacts and the launch-policy check passed; on Windows confirm `latest.yml`, `.exe`, and `.exe.blockmap` are all in the release assets, and remember that a SmartScreen prompt on an unsigned build does not mean the updater failed; on Linux verify auto-update through the AppImage, since `.deb` ships as a manual installer only.
## PR Workflow
1. Create a product branch such as `fix/session-reconnect` or `feat/provider-quality-gate`.
+205
View File
@@ -0,0 +1,205 @@
---
title: Desktop Architecture
nav_title: Desktop
description: Process boundaries between the Electron main process, server sidecar, CLI subprocesses, and IM adapters.
order: 1
---
# Desktop Architecture
The desktop app does not embed the CLI inside React. It is four kinds of processes cooperating, and most maintenance mistakes come from crossing a boundary that was meant to stay closed.
## Process map
```text
Electron main
├── Chromium renderer (React UI)
├── claude-sidecar server
│ ├── Bun.serve HTTP API
│ ├── Bun WebSocket session gateway
│ └── CLI subprocess per session
└── claude-sidecar adapters (launched per platform)
├── Telegram
├── Feishu
├── WeChat
├── DingTalk
└── WhatsApp
```
- **Electron main** owns windows, native dialogs, updates, terminals, native preview, pet windows, and the sidecar lifecycle.
- **Renderer** owns the interface only. It reaches native capabilities through `window.desktopHost`, exposed by preload.
- **Server sidecar** is the local service shared by the desktop app and the H5 client: REST, WebSocket, provider proxy, and session management.
- **CLI subprocess** executes model requests, tool calls, and agent orchestration.
- **Adapter sidecar** bridges IM platform messages into the same server and CLI sessions.
`desktop/src-tauri/` now only holds packaging assets and historical code. It is not the desktop runtime; Electron is the current host.
## Current stack
| Layer | Main technologies | Purpose |
|---|---|---|
| Renderer | React 18, Vite 8, TypeScript, Zustand 5 | Desktop UI and state |
| Styling and content | Tailwind CSS 4, Marked, DOMPurify, Shiki 4, Mermaid, KaTeX | Layout, Markdown, code, diagrams |
| Diff | `react-diff-viewer-continued` | Diffs in chat and workspace |
| Electron host | Electron, electron-builder, electron-updater | Native windows, packaging, updates |
| Terminal | node-pty, xterm.js | Native PTY and terminal rendering |
| Local service | Bun, `Bun.serve` | HTTP API and WebSocket |
Versions follow `desktop/package.json`. This page lists only the majors that change how the architecture reads, not the full dependency list.
## Electron host
| Path | Responsibility |
|---|---|
| `desktop/electron/main.ts` | Electron main entry, windows, IPC registration |
| `desktop/electron/preload.ts` | Typed host API exposed to the main renderer |
| `desktop/electron/preview-preload.ts` | Isolated bridge for native web preview |
| `desktop/electron/pet-preload.ts` | Minimal bridge for pet windows |
| `desktop/electron/ipc/` | IPC channel registration and payload validation |
| `desktop/electron/services/` | Sidecar, update, terminal, preview, window, and proxy services |
The renderer must not import Electron directly, and must not assemble arbitrary IPC channel names. A new native capability means updating the host contract, the main-side validation, and the related tests together.
### Startup sequence
1. Electron resolves the default or portable storage mode and prepares the app config directory.
2. The host picks a free port and starts the unified `claude-sidecar server` binary.
3. The sidecar enters `startServer()` from `src/server/index.ts` and serves HTTP and WebSocket from one `Bun.serve`.
4. Once the health check passes, the renderer takes the loopback server URL and loads sessions and settings.
5. The host starts adapter sidecars for the platforms that are configured.
6. The server starts a CLI subprocess on demand when a session begins.
The server can bind a LAN-reachable address for H5, but the desktop renderer always talks to the loopback control address. H5, remote access, and pet windows each pass their own token and capability limits. A running local server does not make every caller trusted.
### Sidecar entry point
`desktop/sidecars/claude-sidecar.ts` is the single entry point:
```text
claude-sidecar server --app-root <path> --host <host> --port <port>
claude-sidecar cli --app-root <path> [CLI arguments]
claude-sidecar adapters --app-root <path> --telegram|--feishu|--wechat|--dingtalk|--whatsapp
```
The sidecar sets `CLAUDE_APP_ROOT`, `CALLER_DIR`, and the launch arguments before importing any business module, because the top-level modules of the server, CLI, and adapters read those values at import time.
## Server sidecar
The real server entry is `src/server/index.ts`, not a separate `server.ts` wrapper.
```text
src/server/
├── index.ts # Bun.serve, auth, CORS, upgrades, static H5
├── router.ts # REST resource routes
├── api/ # API boundary
├── services/ # Sessions, providers, indexing, diagnostics
├── ws/ # WebSocket protocol and session lifecycle
├── proxy/ # Provider protocol and streaming translation
├── middleware/ # Auth, CORS, error boundary
└── config/ # Provider presets
```
A single `Bun.serve` `fetch` boundary handles:
- `/api/*` REST requests
- `/ws/:sessionId` for desktop, H5, and pet clients
- `/sdk/:sessionId` for internal CLI connections
- OAuth callbacks
- Restricted preview and local file access
- Bundled H5 static assets
Auth rules differ by client capability. Before adding a route, decide whether it belongs to the local desktop, H5, pets, the internal SDK, or public static assets, then place it inside the matching auth and CORS boundary.
## WebSocket semantics
The renderer keeps one connection per session:
```text
ws://<server>/ws/<sessionId>
```
If the server has auth enabled, the client passes the token as a connection query parameter. After the connection opens, the server sends the current session identity, a snapshot of pending permission requests, and the running state.
### Heartbeat and reconnect
Current behavior in `desktop/src/api/websocket.ts`:
- Send `ping` every 30 seconds after connecting.
- If no `pong` arrives within 10 seconds, the client closes the connection and reconnects.
- Reconnect delay grows exponentially from 1 second, capped at 30 seconds.
- Automatic reconnect has no fixed retry ceiling; only closing the session stops it.
- Messages sent while offline are queued in memory.
- After reconnecting, queued messages are flushed first, then `sync_state` fetches the server's authoritative run state.
If you have read "at most 10 retries" somewhere, that wording is stale — the list above is current.
### A disconnected client does not stop the work
After the last client disconnects:
- Foreground turns and background tasks keep running.
- The idle grace period starts only after the work finishes.
- A client reconnecting inside the grace period cancels the cleanup.
- The server stops the CLI only once the grace period passes with no clients.
- Sessions waiting on a permission request have their own bounded cleanup policy so they cannot hold a process forever.
That is why locking a phone, refreshing the renderer, or a brief network switch does not interrupt a running task.
## CLI and provider proxy
The server starts a CLI per session and forwards output, permission requests, tool results, and background task state over an internal protocol. Provider, model, effort, and permission mode are maintained jointly by the server and the CLI; the renderer is not the single source of truth.
`src/server/proxy/` handles the supported provider protocols:
- Anthropic Messages
- OpenAI Chat Completions
- OpenAI Responses
Model mapping, authentication style, and context settings follow the actual provider configuration. Do not hardcode a vendor list in architecture docs.
## IM adapters
Each platform runs its own adapter sidecar, so one platform's bad credentials or failed startup cannot take the others down.
```text
IM platform
→ adapters/<platform>
→ adapters/common WebSocket bridge
→ server sidecar
→ CLI session
```
The shared layer in `adapters/common/` handles configuration, pairing, session mapping, message buffering, deduplication, attachments, and the server WebSocket bridge. Current platform directories:
- `adapters/telegram/`
- `adapters/feishu/`
- `adapters/wechat/`
- `adapters/dingtalk/`
- `adapters/whatsapp/`
## Persistence boundaries
Different data does not share one store:
| Data | Primary boundary |
|---|---|
| Renderer UI preferences and open tabs | Browser storage and its migrations |
| Sessions and messages | Local session data managed by the server and CLI |
| Provider, H5, Computer Use, and other settings | Config files managed by the server |
| Electron window and native state | Electron user data / app mode directory |
| IM configuration and pairing state | Adapter config and per-platform state directories |
Any change to a JSON shape, a `localStorage` key, or the app config layout needs a forward migration plus a regression test against old data. Never overwrite the user's shared Claude configuration, and never read the real user directory from tests.
## Build and verification
Desktop builds are orchestrated by scripts in `desktop/package.json`:
```bash
cd desktop
bun run build:sidecars
bun run build
bun run build:electron
```
`electron:package` runs electron-builder on top of that. Verifying a packaged artifact only proves the asset and installer structure is correct. User flows that touch windows, permissions, terminals, preview, updates, or Computer Use still need a real desktop smoke test.

Before

Width:  |  Height:  |  Size: 1.8 MiB

After

Width:  |  Height:  |  Size: 1.8 MiB

Before

Width:  |  Height:  |  Size: 556 KiB

After

Width:  |  Height:  |  Size: 556 KiB

Before

Width:  |  Height:  |  Size: 888 KiB

After

Width:  |  Height:  |  Size: 888 KiB

Before

Width:  |  Height:  |  Size: 490 KiB

After

Width:  |  Height:  |  Size: 490 KiB

Before

Width:  |  Height:  |  Size: 460 KiB

After

Width:  |  Height:  |  Size: 460 KiB

Before

Width:  |  Height:  |  Size: 1.8 MiB

After

Width:  |  Height:  |  Size: 1.8 MiB

Before

Width:  |  Height:  |  Size: 523 KiB

After

Width:  |  Height:  |  Size: 523 KiB

Before

Width:  |  Height:  |  Size: 445 KiB

After

Width:  |  Height:  |  Size: 445 KiB

Before

Width:  |  Height:  |  Size: 487 KiB

After

Width:  |  Height:  |  Size: 487 KiB

Before

Width:  |  Height:  |  Size: 655 KiB

After

Width:  |  Height:  |  Size: 655 KiB

Before

Width:  |  Height:  |  Size: 449 KiB

After

Width:  |  Height:  |  Size: 449 KiB

Before

Width:  |  Height:  |  Size: 912 KiB

After

Width:  |  Height:  |  Size: 912 KiB

Before

Width:  |  Height:  |  Size: 398 KiB

After

Width:  |  Height:  |  Size: 398 KiB

Before

Width:  |  Height:  |  Size: 562 KiB

After

Width:  |  Height:  |  Size: 562 KiB

Before

Width:  |  Height:  |  Size: 2.2 MiB

After

Width:  |  Height:  |  Size: 2.2 MiB

Before

Width:  |  Height:  |  Size: 103 KiB

After

Width:  |  Height:  |  Size: 103 KiB

Before

Width:  |  Height:  |  Size: 800 KiB

After

Width:  |  Height:  |  Size: 800 KiB

Before

Width:  |  Height:  |  Size: 508 KiB

After

Width:  |  Height:  |  Size: 508 KiB

Before

Width:  |  Height:  |  Size: 463 KiB

After

Width:  |  Height:  |  Size: 463 KiB

Before

Width:  |  Height:  |  Size: 492 KiB

After

Width:  |  Height:  |  Size: 492 KiB

Before

Width:  |  Height:  |  Size: 430 KiB

After

Width:  |  Height:  |  Size: 430 KiB

Before

Width:  |  Height:  |  Size: 1.9 MiB

After

Width:  |  Height:  |  Size: 1.9 MiB

Before

Width:  |  Height:  |  Size: 595 KiB

After

Width:  |  Height:  |  Size: 595 KiB

Before

Width:  |  Height:  |  Size: 116 KiB

After

Width:  |  Height:  |  Size: 116 KiB

Before

Width:  |  Height:  |  Size: 568 KiB

After

Width:  |  Height:  |  Size: 568 KiB

Before

Width:  |  Height:  |  Size: 785 KiB

After

Width:  |  Height:  |  Size: 785 KiB

Before

Width:  |  Height:  |  Size: 1.7 MiB

After

Width:  |  Height:  |  Size: 1.7 MiB

Before

Width:  |  Height:  |  Size: 576 KiB

After

Width:  |  Height:  |  Size: 576 KiB

Before

Width:  |  Height:  |  Size: 1.5 MiB

After

Width:  |  Height:  |  Size: 1.5 MiB

Before

Width:  |  Height:  |  Size: 2.2 MiB

After

Width:  |  Height:  |  Size: 2.2 MiB

Before

Width:  |  Height:  |  Size: 91 KiB

After

Width:  |  Height:  |  Size: 91 KiB

Before

Width:  |  Height:  |  Size: 668 KiB

After

Width:  |  Height:  |  Size: 668 KiB

Before

Width:  |  Height:  |  Size: 681 KiB

After

Width:  |  Height:  |  Size: 681 KiB

Before

Width:  |  Height:  |  Size: 513 KiB

After

Width:  |  Height:  |  Size: 513 KiB

Before

Width:  |  Height:  |  Size: 87 KiB

After

Width:  |  Height:  |  Size: 87 KiB

Before

Width:  |  Height:  |  Size: 517 KiB

After

Width:  |  Height:  |  Size: 517 KiB

Before

Width:  |  Height:  |  Size: 92 KiB

After

Width:  |  Height:  |  Size: 92 KiB

Before

Width:  |  Height:  |  Size: 5.0 MiB

After

Width:  |  Height:  |  Size: 5.0 MiB

Before

Width:  |  Height:  |  Size: 757 KiB

After

Width:  |  Height:  |  Size: 757 KiB

Before

Width:  |  Height:  |  Size: 646 KiB

After

Width:  |  Height:  |  Size: 646 KiB

Before

Width:  |  Height:  |  Size: 484 KiB

After

Width:  |  Height:  |  Size: 484 KiB

Before

Width:  |  Height:  |  Size: 521 KiB

After

Width:  |  Height:  |  Size: 521 KiB

Before

Width:  |  Height:  |  Size: 470 KiB

After

Width:  |  Height:  |  Size: 470 KiB

Before

Width:  |  Height:  |  Size: 665 KiB

After

Width:  |  Height:  |  Size: 665 KiB

+64
View File
@@ -0,0 +1,64 @@
---
title: Architecture Overview
nav_title: Architecture
description: How the five layers divide the work, how one message flows through them, and which page to read for the change you want to make.
order: 0
---
# Architecture Overview
Claude Code Haha looks like one desktop app, but it is five pieces of code that each run on their own. Before reading source or opening a PR, get the boundaries straight — every other page in this section lives inside one of them.
```text
desktop/src/ Frontend — React + Zustand. Draws the UI, touches no native capability.
desktop/electron/ Desktop shell — Electron main process: windows, updates, terminals, native preview.
src/server/ Local server — REST + WebSocket on Bun.serve, shared by desktop and phone.
src/ CLI core — agent loop, tool system, permissions, memory, skills.
adapters/ IM bridges — one sidecar per platform, all wired back to the same sessions.
```
Three facts are worth holding onto:
- **The CLI core is the only thing that executes.** Every desktop session makes the server spawn a CLI subprocess; every button in the frontend eventually becomes a message sent to it.
- **The local server is the only entry point.** Desktop, mobile H5, and IM adapters share one REST and WebSocket surface — they differ only in how much they are trusted.
- **Electron is the current desktop path.** `desktop/src-tauri/` keeps packaging assets and historical code as a rollback option; it is not the runtime.
## How the CLI core is layered
![The layered structure of the CLI core](../../images/01-overall-architecture.png)
After bootstrap, the entry layer splits in two. On the left is the path that carries one request to completion: terminal UI, query engine, tool system, subagents. On the right are the cross-cutting capabilities that path calls into: state management, skills and plugins, and the service layer holding MCP, OAuth, and memory. The desktop app replaces the terminal UI box and reuses everything else as-is.
## What happens to one message
![The lifecycle of a single request](../../images/02-request-lifecycle.png)
Input is parsed, context is assembled, and the request goes to the model. The moment a tool call appears in the stream, it passes a permission check before it runs; the result is folded back into the context and the next turn starts, until the model stops asking for tools. The tool cards and permission prompts you see in the desktop app are two nodes of this path, rendered.
## How tools and permissions relate
![The tool system and its permission gate](../../images/03-tool-system.png)
Every tool registers in one registry, grouped by capability: files, shell, system, subagents, external integrations, and communication. The fixed pipeline underneath is what actually defines the safety boundary — any call passes argument validation and the permission gate before it reaches the sandbox. Adding a tool to the registry is easy; deciding which side of that gate it belongs on is the real work.
## I want to change X — where do I start
| What you want to touch | Start here |
|---|---|
| Windows, tray, auto-update, embedded terminal, native preview | [Desktop architecture](./desktop.md) |
| REST endpoints, WebSocket events, auth, provider proxy | [Local server and API](./server.md) |
| You cannot find which directory a feature lives in | [Project structure](./structure.md) |
| Subagent behavior, built-in agents, Agent Teams | [Multi-agent usage guide](./agent.md) |
| Spawn paths, tool pool filtering, the background task engine | [Multi-agent internals](./agent-internals.md) |
| The agent loop, system prompts, context compression | [Agent framework deep dive](./agent-framework.md) |
| Writing skills, source precedence, activation | [Skills usage guide](./skills.md) |
| Skill discovery, injection, fork execution | [Skills internals](./skills-internals.md) |
| Where memories live and when they are written | [Memory system usage guide](./memory.md) |
| Memory injection, extraction, retrieval, team sync | [Memory system internals](./memory-internals.md) |
| Background memory consolidation | [AutoDream memory consolidation](./autodream.md) |
| Screen control tools, authorization, coordinate mapping | [Computer Use architecture](./computer-use.md) |
| IM message protocol, access control, permission relay | [Channel system](./channel.md) |
| What to run before a PR, and the release process | [Contributing and quality gates](./contributing.md) |
| Running the CLI in a terminal or from a script | [CLI install and run](../cli/index.md) |
Product features are covered elsewhere — start from [Get started](../start/index.md) and [Desktop features](../desktop/index.md).
@@ -1,16 +1,17 @@
# Claude Code Memory System — Implementation Details
---
title: Memory System Internals
nav_title: Memory Internals
description: Path resolution, prompt injection, auto-extraction, retrieval, and team memory sync.
order: 10
---
> From system prompt injection to background auto-extraction, dissecting every technical detail of the memory system.
# Memory System Internals
<p align="center">
<a href="#1-overall-architecture">Architecture</a> · <a href="#2-path-resolution-system">Path Resolution</a> · <a href="#3-system-prompt-injection">Prompt Injection</a> · <a href="#4-automatic-memory-extraction">Auto-Extraction</a> · <a href="#5-intelligent-memory-retrieval">Intelligent Retrieval</a> · <a href="#6-memory-scanning-in-detail">Memory Scanning</a> · <a href="#7-agent-memory">Agent Memory</a> · <a href="#8-team-memory-sync">Team Sync</a> · <a href="#9-key-constants-quick-reference">Constants</a> · <a href="#10-data-flow-overview">Data Flow</a>
</p>
From system prompt injection to background auto-extraction, dissecting every technical detail of the memory system.
![Implementation Architecture Overview](./images/05-architecture-overview.png)
---
## 1. Overall Architecture
## Overall Architecture
The memory system is powered by 5 core modules working in concert:
@@ -26,16 +27,14 @@ Auxiliary modules:
| Module | Source Location | Responsibility |
|--------|----------------|----------------|
| **AutoDream** | `src/services/autoDream/` | Background memory consolidation ("dreaming"); see [03-autodream.md](./03-autodream.md) |
| **AutoDream** | `src/services/autoDream/` | Background memory consolidation ("dreaming"); see [AutoDream](./autodream.md) |
| **Type Definitions** | `src/memdir/memoryTypes.ts` | Taxonomy and prompt templates for the four memory types |
| **Freshness** | `src/memdir/memoryAge.ts` | Calculates memory age, generates stale warnings |
| **File Detection** | `src/utils/memoryFileDetection.ts` | Determines whether a path belongs to the memory system |
| **Agent Memory** | `src/tools/AgentTool/agentMemory.ts` | Three-level memory directories exclusive to sub-agents |
| **Agent Memory** | `src/tools/AgentTool/agentMemory.ts` | Three-level memory directories exclusive to subagents |
| **Team Sync** | `src/services/teamMemorySync/` | Remote upload/download of memories |
---
## 2. Path Resolution System
## Path Resolution System
![Path Resolution Flow](./images/06-path-resolution.png)
@@ -89,9 +88,7 @@ settings.json autoMemoryEnabled -> Follows setting
Default -> Enabled
```
---
## 3. System Prompt Injection
## System Prompt Injection
![Prompt Injection Flow](./images/07-prompt-injection.png)
@@ -148,9 +145,7 @@ if (bytes > 25,000) -> Truncate at last newline character
The prompt explicitly tells the model the directory already exists, avoiding wasted turns on `ls` or `mkdir`.
---
## 4. Automatic Memory Extraction
## Automatic Memory Extraction
![Auto-Extraction Flow](./images/08-auto-extraction.png)
@@ -236,9 +231,7 @@ If a previous extraction is still running:
2. After the old extraction completes, a **tail extraction** is immediately launched
3. The tail extraction only processes messages added between the two calls
---
## 5. Intelligent Memory Retrieval
## Intelligent Memory Retrieval
![Memory Retrieval Flow](./images/09-memory-retrieval.png)
@@ -293,9 +286,7 @@ function memoryFreshnessText(mtimeMs: number): string {
}
```
---
## 6. Memory Scanning in Detail
## Memory Scanning in Detail
### `scanMemoryFiles()`
@@ -330,13 +321,11 @@ Generates a manifest consumed by Sonnet or the extraction agent:
- [project] freeze.md (2026-03-10T15:00:00.000Z): Merge freeze starting 3/5
```
---
## 7. Agent Memory
## Agent Memory
![Agent Memory Three-Level Scoping](./images/10-agent-memory.png)
Sub-agents (launched via the Agent tool) have an independent three-level memory system:
Subagents (launched via the Agent tool) have an independent three-level memory system:
| Scope | Path | Description |
|-------|------|-------------|
@@ -349,9 +338,7 @@ Differences from main memory:
- Files can be written directly without the two-step process
- Each agent type is isolated (explorer, planner, etc. each have their own directory)
---
## 8. Team Memory Sync
## Team Memory Sync
When the `TEAMMEM` feature flag is enabled:
@@ -391,9 +378,7 @@ In `memoryTypes.ts`, each type has a `<scope>` directive:
| project | Leans toward team |
| reference | Usually team |
---
## 9. Key Constants Quick Reference
## Key Constants Quick Reference
```typescript
// Index file
@@ -415,9 +400,7 @@ maxTurns = 5 // Forked agent max 5 turns
Max 5 relevant memories returned
```
---
## 10. Data Flow Overview
## Data Flow Overview
```
┌─────────────────────────────────────────────────────┐
@@ -475,6 +458,6 @@ Max 5 relevant memories returned
│ -> Merge duplicates / fix stale / compress index │
│ -> Notify user: "Improved N memories" │
│ │
│ See 03-autodream.md for details │
│ See autodream.md for details │
└─────────────────────────────────────────────────────┘
```
@@ -1,16 +1,17 @@
# Claude Code Memory System — Usage Guide
---
title: Memory System Usage Guide
nav_title: Memory Usage
description: The four memory types, how saving is triggered, where memories live, and how to manage them.
order: 9
---
> Let Claude Code remember who you are, what you prefer, and what's happening in your project across sessions.
# Memory System Usage Guide
<p align="center">
<a href="#1-what-is-the-memory-system">Memory System</a> · <a href="#2-four-memory-types">Four Memory Types</a> · <a href="#3-how-to-trigger-memory-saving">Trigger Saving</a> · <a href="#4-where-are-memories-stored">Storage Location</a> · <a href="#5-how-to-manage-memories">Manage Memories</a> · <a href="#6-memory-lifecycle">Lifecycle</a> · <a href="#7-quick-reference">Quick Reference</a>
</p>
Let Claude Code remember who you are, what you prefer, and what's happening in your project across sessions.
![Memory System Overview](./images/01-memory-overview.png)
---
## 1. What Is the Memory System?
## What Is the Memory System?
Claude Code's memory system is a **file-based persistent knowledge store** that allows Claude to continuously build understanding of you and your project across multiple conversations.
@@ -23,15 +24,13 @@ Core principle: **Only remember things that cannot be inferred from the code its
| Non-critical merges frozen after Thursday | Existing CLAUDE.md content |
| Bug tracking is in Linear's INGEST project | Debugging solutions (fixes are already in the code) |
---
## 2. Four Memory Types
## Four Memory Types
![Four Memory Types](./images/02-memory-types.png)
Claude Code strictly categorizes memories into four types:
### 2.1 User (User Profile)
### User (User Profile)
Records your role, goals, skill level, and preferences to help Claude tailor its collaboration approach.
@@ -40,7 +39,7 @@ User says: I've written Go for ten years, but this is my first time touching the
Claude saves: Deep Go experience, React newcomer — explain frontend concepts using backend analogies
```
### 2.2 Feedback (Behavioral Feedback)
### Feedback (Behavioral Feedback)
Your corrections or affirmations about how Claude works. These memories prevent Claude from repeating the same mistakes.
@@ -51,7 +50,7 @@ Claude saves: User prefers concise replies, no trailing summaries
**Important**: Not only corrections are recorded -- affirmations are too. If Claude makes a non-obvious choice and you approve, that gets remembered as well.
### 2.3 Project (Project Context)
### Project (Project Context)
Project context that cannot be derived from the code or Git history: who's doing what, why, and deadlines.
@@ -62,7 +61,7 @@ Claude saves: Merge freeze starting 2026-03-05, flag non-critical PR work after
**Note**: Claude converts relative dates ("Thursday") to absolute dates ("2026-03-05") to ensure memories don't become ambiguous over time.
### 2.4 Reference (External References)
### Reference (External References)
Pointers to information in external systems: dashboards, issue trackers, Slack channels.
@@ -71,9 +70,7 @@ User says: On-call monitors the grafana.internal/d/api-latency dashboard
Claude saves: grafana.internal/d/api-latency is the on-call latency dashboard — check when editing request path code
```
---
## 3. How to Trigger Memory Saving
## How to Trigger Memory Saving
![Memory Trigger Flow](./images/03-memory-trigger.png)
@@ -84,8 +81,8 @@ This is the primary method. **You don't need to do anything** -- Claude automati
Workflow:
1. You have a normal conversation with Claude
2. Claude finishes its response (no tool calls pending)
3. A **memory extraction sub-agent** starts in the background
4. The sub-agent analyzes the recent conversation content
3. A **memory extraction subagent** starts in the background
4. The subagent analyzes the recent conversation content
5. It identifies memories worth saving
6. Writes memory files + updates the MEMORY.md index
@@ -121,9 +118,7 @@ Type `/remember` to trigger the memory review skill, which will:
- Detect duplicate, outdated, and conflicting memories
- **Does not modify anything directly** -- all changes require your approval
---
## 4. Where Are Memories Stored?
## Where Are Memories Stored?
### Directory Structure
@@ -173,9 +168,7 @@ MEMORY.md is an index, not content. It is **always loaded into context**, with o
**Limit**: Maximum 200 lines or 25KB; content beyond this is truncated.
---
## 5. How to Manage Memories
## How to Manage Memories
### Ask Claude to Forget
@@ -215,9 +208,7 @@ Set in `~/.claude/settings.json`:
Supports `~/` expansion. For security reasons, the project-level `.claude/settings.json` is **not allowed** to set this option.
---
## 6. Memory Lifecycle
## Memory Lifecycle
![Memory Lifecycle](./images/04-memory-lifecycle.png)
@@ -245,14 +236,14 @@ New information learned during conversation
### AutoDream -- "Dreaming" to Organize Memories
Claude Code has a hidden **AutoDream** feature, analogous to how the human brain organizes memories during sleep. When the following conditions are met, Claude silently launches a "dreaming" sub-agent in the background:
Claude Code has a hidden **AutoDream** feature, analogous to how the human brain organizes memories during sleep. When the following conditions are met, Claude silently launches a "dreaming" subagent in the background:
- At least **>= 24 hours** since the last consolidation
- At least **>= 5 sessions** accumulated in the interim
The dreaming process has four phases: Orient -> Gather -> Consolidate -> Prune. The bottom status bar shows **"dreaming"**, and you can press `Shift+Down` to view progress or `x` to terminate.
For a detailed technical analysis, see [AutoDream Memory Consolidation](./03-autodream.md).
For a detailed technical analysis, see [AutoDream Memory Consolidation](./autodream.md).
### Freshness Management
@@ -260,9 +251,7 @@ For a detailed technical analysis, see [AutoDream Memory Consolidation](./03-aut
- **Memories older than 1 day**: Accompanied by a stale warning, reminding Claude to verify before citing
- **Memories referencing file paths/function names**: Confirmed via grep before use to ensure they still exist
---
## 7. Quick Reference
## Quick Reference
| Action | Method |
|--------|--------|
@@ -1,4 +1,11 @@
# Local Server
---
title: Local Server and API
nav_title: Local Server
description: Startup flags, access control, REST surface, and chat WebSocket of the local server.
order: 2
---
# Local Server and API
The local server is the runtime boundary between Desktop, the H5 browser client, and the Claude CLI. It provides REST APIs, the chat WebSocket, provider protocol translation, and Desktop web assets. The packaged Desktop app manages it automatically. Start it manually only for source development, headless deployment, or a custom client.
@@ -1,16 +1,17 @@
# Claude Code Skills System -- Implementation Details
---
title: Skills Internals
nav_title: Skills Internals
description: How skills are discovered, parsed, injected into conversations, and run in fork subagents.
order: 8
---
> A deep dive into how Skills are discovered, loaded, injected, executed, and managed.
# Skills Internals
<p align="center">
<a href="#1-overall-architecture">Architecture</a> · <a href="#2-skill-discovery-and-loading">Discovery & Loading</a> · <a href="#3-frontmatter-parsing">Frontmatter Parsing</a> · <a href="#4-skill-injection-into-conversations">Injection</a> · <a href="#5-skilltool-execution-engine">Execution Engine</a> · <a href="#6-fork-sub-agent-execution">Fork Execution</a> · <a href="#7-conditional-activation-and-dynamic-discovery">Conditional Activation</a> · <a href="#8-hook-integration">Hook Integration</a> · <a href="#9-permission-system">Permission System</a> · <a href="#10-complete-lifecycle">Complete Lifecycle</a> · <a href="#11-source-code-index">Source Index</a>
</p>
A deep dive into how Skills are discovered, loaded, injected, executed, and managed.
![Skills Architecture Overview](./images/04-skills-architecture.png)
---
## 1. Overall Architecture
## Overall Architecture
The Skills system consists of 5 core modules working together:
@@ -40,9 +41,7 @@ The Skills system consists of 5 core modules working together:
| Activation | `loadSkillsDir.ts` (second half) | Conditional activation and dynamic discovery |
| Context | `forkedAgent.ts` | Context preparation and modification |
---
## 2. Skill Discovery and Loading
## Skill Discovery and Loading
### Loading Entry Point
@@ -225,9 +224,7 @@ async function fetchCommandsForClient(client) {
**Feature gate:** `feature('MCP_SKILLS')` controls whether MCP Skills are available.
---
## 3. Frontmatter Parsing
## Frontmatter Parsing
### Parsing Flow
@@ -309,9 +306,7 @@ async getPromptForCommand(args, toolUseContext) {
}
```
---
## 4. Skill Injection into Conversations
## Skill Injection into Conversations
### Injection Flow
@@ -394,9 +389,7 @@ Important:
})
```
---
## 5. SkillTool Execution Engine
## SkillTool Execution Engine
### Tool Definition
@@ -498,9 +491,7 @@ async function getAllCommands(context: ToolUseContext): Promise<Command[]> {
}
```
---
## 6. Fork Sub-Agent Execution
## Fork Sub-Agent Execution
### Execution Flow
@@ -603,9 +594,7 @@ export async function prepareForkedCommandContext(
}
```
---
## 7. Conditional Activation and Dynamic Discovery
## Conditional Activation and Dynamic Discovery
### Conditional Skills
@@ -679,9 +668,7 @@ clearCommandMemoizationCaches()
Next conversation turn loads the updated Skill list
```
---
## 8. Hook Integration
## Hook Integration
### Hook Registration
@@ -732,9 +719,7 @@ During session
└─ Hooks with once: true are automatically removed after first execution
```
---
## 9. Permission System
## Permission System
### Check Flow
@@ -775,9 +760,7 @@ SAFE_SKILL_PROPERTIES = {
**Core principle:** Newly added frontmatter fields require permission by default, unless explicitly added to the whitelist.
---
## 10. Complete Lifecycle
## Complete Lifecycle
### Data Flow Overview
@@ -841,9 +824,7 @@ On compression → buildPostCompactMessages()
Restored scoped by agentId (prevents cross-agent leakage)
```
---
## 11. Source Code Index
## Source Code Index
### Core Files
@@ -1,16 +1,17 @@
# Claude Code Skills System -- Usage Guide
---
title: Skills Usage Guide
nav_title: Skills Usage
description: Skill sources, definition format, invocation, execution context, and permission control.
order: 7
---
> Skills are the extensible capability engine of Claude Code, allowing you to define custom automated workflows using Markdown files.
# Skills Usage Guide
<p align="center">
<a href="#1-what-are-skills">What Are Skills</a> · <a href="#2-six-skill-sources">Six Sources</a> · <a href="#3-skill-definition-format">Definition Format</a> · <a href="#4-invocation-methods">Invocation</a> · <a href="#5-execution-context">Execution Context</a> · <a href="#6-conditional-activation">Conditional Activation</a> · <a href="#7-permission-control">Permissions</a> · <a href="#8-quick-reference">Quick Reference</a>
</p>
Skills are the extensible capability engine of Claude Code, allowing you to define custom automated workflows using Markdown files.
![Skills System Overview](./images/01-skills-overview.png)
---
## 1. What Are Skills?
## What Are Skills?
Skills are Claude Code's **extensible capability plugin system**. Each Skill is a Markdown file (with YAML frontmatter) that defines a specialized prompt and behavioral configuration, enabling Claude to execute professional workflows in specific scenarios.
@@ -21,13 +22,11 @@ Core capabilities:
| Specialized Workflows | Define standard processes for code review, TDD, debugging, etc. |
| Tool Permission Control | Restrict a Skill to only use specified tools |
| Model Switching | Assign different models to different Skills |
| Execution Isolation | Fork mode runs in an isolated sub-agent |
| Execution Isolation | Fork mode runs in an isolated subagent |
| Conditional Activation | Activate only when specific files are being operated on |
| Hook Injection | Automatically register lifecycle hooks when a Skill is invoked |
---
## 2. Six Skill Sources
## Six Skill Sources
![Skill Source Types](./images/02-skill-sources.png)
@@ -128,9 +127,7 @@ Provided by connected MCP servers, naming format: `mcp__server-name__prompt-name
**Security restriction**: MCP Skills are from remote untrusted sources and are **prohibited from executing** `!`...`` inline shell commands.
---
## 3. Skill Definition Format
## Skill Definition Format
### Directory Structure
@@ -206,9 +203,7 @@ Supported special syntax:
| `argument-hint` | string | -- | Argument hint text |
| `version` | string | -- | Version number |
---
## 4. Invocation Methods
## Invocation Methods
![Skill Invocation Flow](./images/03-skill-invocation.png)
@@ -259,9 +254,7 @@ When Skills with the same name exist in multiple sources, they are resolved in t
7. Built-in Commands ← Lowest priority
```
---
## 5. Execution Context
## Execution Context
### Inline Mode (Default)
@@ -279,7 +272,7 @@ context: inline # Default value, can be omitted
### Fork Mode (Sub-Agent)
The Skill runs in an **isolated sub-agent** with its own independent token budget and context.
The Skill runs in an **isolated subagent** with its own independent token budget and context.
```yaml
context: fork
@@ -303,9 +296,7 @@ agent: general-purpose # Optional, specifies agent type
| Use Cases | Brief guidance, extended context | Long tasks, independent computation |
| Tool Restrictions | contextModifier modification | modifiedGetAppState |
---
## 6. Conditional Activation
## Conditional Activation
Skills can use the `paths` frontmatter to implement **on-demand activation**, becoming visible to the model only when matching files are operated on.
@@ -342,9 +333,7 @@ In addition to conditional activation, Skills also support **runtime discovery**
5. New directory found → addSkillDirectories() → load and register
```
---
## 7. Permission Control
## Permission Control
### Auto-Allow
@@ -370,9 +359,7 @@ Allow? (y)es / (n)o / (a)lways allow / (d)eny
**Processing order:** Deny rules → Allow rules → Safe property check → Ask user
---
## 8. Quick Reference
## Quick Reference
### Creating a Skill
@@ -1,3 +1,10 @@
---
title: Project Structure
nav_title: Structure
description: Responsibility boundaries across the CLI, server, desktop app, adapters, and docs site.
order: 3
---
# Project Structure
The repository contains the CLI/TUI, local Server, Electron desktop app, IM adapters, and documentation site. This map lists stable responsibility boundaries rather than every file.
-121
View File
@@ -1,121 +0,0 @@
# Claude Code Memory System Documentation
> Complete usage guide and technical implementation documentation for the memory system
---
## Documentation Index
### [01-usage-guide.md](./01-usage-guide.md) — Usage Guide
A comprehensive user-facing manual covering:
- **Four memory types**: User (user profile), Feedback (behavioral feedback), Project (project context), Reference (external references)
- **Four trigger methods**: Automatic extraction, explicit requests, `/memory` command, `/remember` command
- **Storage format**: YAML frontmatter + Markdown content
- **Management operations**: Forgetting, ignoring, manual editing, disabling, custom directories
- **Lifecycle**: From learning to injection, freshness management
**Target audience**: All Claude Code users
---
### [02-implementation.md](./02-implementation.md) — Implementation Details
A deep technical analysis for developers, covering:
- **5 core modules**: Path resolution, prompt construction, memory scanning, intelligent retrieval, automatic extraction
- **Path resolution system**: Priority chain, security validation, enable conditions
- **System prompt injection**: `loadMemoryPrompt()` -> `buildMemoryLines()`, MEMORY.md truncation strategy
- **Automatic memory extraction**: Forked agent, mutual exclusion mechanism, tool permissions, merge mechanism
- **Intelligent retrieval**: `scanMemoryFiles()` -> Sonnet selection -> freshness warnings
- **Agent memory**: Three-level scoping (user/project/local)
- **Team memory sync**: Pull/Push API, merge semantics
- **Complete data flow**: From session startup to context injection
**Target audience**: Contributors, architects, developers who want a deep understanding of the implementation
---
### [03-autodream.md](./03-autodream.md) — AutoDream Memory Consolidation
Claude's "dreaming" mechanism -- a deep dive into background silent memory consolidation, covering:
- **Core concept**: Like how the human brain organizes memories during sleep, periodically reviewing multiple sessions to consolidate knowledge
- **Five-gate check**: Feature toggle -> Time gate (24h) -> Scan throttle (10min) -> Session gate (5 sessions) -> Lock gate
- **Four-phase process**: Orient -> Gather -> Consolidate -> Prune
- **Security restrictions**: Read-only Bash, write operations limited to memory directory, PID lock file mutual exclusion
- **UI presentation**: Bottom "dreaming" label, Shift+Down detail dialog, completion notification
- **Configuration control**: settings.json local toggle + GrowthBook remote feature flag
- **Comparison with extractMemories**: Taking notes during the day vs. organizing the notebook while sleeping
**Target audience**: Contributors, architects, developers interested in Claude's automated memory management
---
## Illustrations
All illustrations use a dark background (#1a1a2e) + Claude Code Haha orange-blue accent (#FF7A00) style, consistent with the official Claude Code documentation.
| Image | Description | Size |
|-------|-------------|------|
| `01-memory-overview.png` | Memory system overview -- four-layer architecture (trigger/type/storage/retrieval) | 632 KB |
| `02-memory-types.png` | Four memory types -- 2x2 grid showing User/Feedback/Project/Reference | 507 KB |
| `03-memory-trigger.png` | Memory trigger flow -- four paths from conversation to storage | 474 KB |
| `04-memory-lifecycle.png` | Memory lifecycle -- complete cycle flow + freshness checks | 1.0 MB |
| `05-architecture-overview.png` | Implementation architecture overview -- 5 core modules + auxiliary modules | 3.5 MB |
| `06-path-resolution.png` | Path resolution flow -- three-level priority + security validation | 1.0 MB |
| `07-prompt-injection.png` | Prompt injection flow -- loadMemoryPrompt dispatch logic | 1.1 MB |
| `08-auto-extraction.png` | Auto-extraction flow -- forked agent complete process | 1.2 MB |
| `09-memory-retrieval.png` | Intelligent retrieval flow -- Sonnet selection + freshness management | 816 KB |
| `10-agent-memory.png` | Agent memory scoping -- three-level nested structure | 523 KB |
| `11-autodream-overview.png` | AutoDream overview -- dreaming mechanism core architecture and human sleep analogy | 777 KB |
| `12-autodream-trigger.png` | AutoDream trigger flow -- five-gate check chain | 493 KB |
| `13-autodream-phases.png` | AutoDream four phases -- Orient/Gather/Consolidate/Prune | 602 KB |
---
## Quick Start
### For Users
1. Read the [Usage Guide](./01-usage-guide.md)
2. Learn about the four memory types and trigger methods
3. Try the `/memory` and `/remember` commands
### For Developers
1. Read the [Implementation Details](./02-implementation.md)
2. Explore the source code:
- `src/memdir/paths.ts` -- Path resolution
- `src/memdir/memdir.ts` -- Prompt construction
- `src/memdir/memoryScan.ts` -- Memory scanning
- `src/memdir/findRelevantMemories.ts` -- Intelligent retrieval
- `src/services/extractMemories/` -- Automatic extraction
3. Understand the data flow and module interactions
---
## Core Concepts Quick Reference
| Concept | Description |
|---------|-------------|
| **MEMORY.md** | Index file, always loaded into context (max 200 lines / 25KB) |
| **Topic files** | `*.md` files containing frontmatter + content |
| **Auto-extraction** | Runs in the background after each response; a forked agent analyzes the conversation |
| **AutoDream** | Triggers after 24h + 5 sessions; consolidates/deduplicates/prunes all memories in the background |
| **Intelligent retrieval** | Sonnet model selects up to 5 relevant memories from the entire collection |
| **Freshness** | <=1 day: no warning; >1 day: stale warning attached |
| **Forked agent** | Shares prompt cache, restricted tool permissions, max 5 turns |
| **Three-level scoping** | Agent memory: user (global) > project (project-level) > local (local-level) |
---
## Related Resources
- [Claude Code Haha Home](/en/)
- [Memory system source code](https://github.com/NanmiCoder/cc-haha/tree/main/src/memdir/)
- [Auto-extraction service](https://github.com/NanmiCoder/cc-haha/tree/main/src/services/extractMemories/)
- [AutoDream service](https://github.com/NanmiCoder/cc-haha/tree/main/src/services/autoDream/)
- [DreamTask](https://github.com/NanmiCoder/cc-haha/tree/main/src/tasks/DreamTask/)
- [GitHub Issues](https://github.com/NanmiCoder/cc-haha/issues)
-13
View File
@@ -1,13 +0,0 @@
# Fixes Compared with the Original Leaked Source
The leaked source could not run directly. This repository mainly fixes the following issues:
| Issue | Root cause | Fix |
|------|------|------|
| TUI does not start | The entry script routed no-argument startup to the recovery CLI | Restored the full `cli.tsx` entry |
| Startup hangs | The `verify` skill imports a missing `.md` file, causing Bun's text loader to hang indefinitely | Added stub `.md` files |
| `--print` hangs | `filePersistence/types.ts` was missing | Added type stub files |
| `--print` hangs | `ultraplan/prompt.txt` was missing | Added resource stub files |
| **Enter key does nothing** | The `modifiers-napi` native package was missing, `isModifierPressed()` threw, `handleEnter` was interrupted, and `onSubmit` never ran | Added try/catch fault tolerance |
| Setup was skipped | `preload.ts` automatically set `LOCAL_RECOVERY=1`, skipping all initialization | Removed the default setting |
+103
View File
@@ -0,0 +1,103 @@
---
title: Run your first session
nav_title: First session
description: Pick a folder, set permissions, state a goal, watch it edit files, review the diff.
order: 3
---
# Run your first session
Your model is connected. Now let it actually do something. Find a project you don't mind touching — ideally one under Git, so anything it breaks is one command away from being undone.
## 1. New session, pick the project folder
Click "New session" in the sidebar, or press `Cmd/Ctrl + N`.
![Empty session, with permission mode, launch location, and model controls in the composer](../../images/app/session-new.webp)
Look at the row along the bottom of the composer: `+` for attachments, then **permission mode**, **launch location**, **model and effort**, and finally "Run".
Start with the **launch location** pill in the middle (it reads `task-board / main` in the screenshot) and pick a project folder. This sets the agent's boundary: reading files, searching, running commands, checking Git status — all of it happens inside this directory and nowhere else.
If the folder is a Git repository, the same pill also lets you choose a branch and whether to use an isolated worktree. Skip the worktree for now — working directly in the folder makes the changes easiest to follow.
## 2. Permissions: start with "Ask permissions"
Click the permission mode button and all five levels open up.
![The five permission modes](../../images/app/permission-modes.webp)
| Mode | What the app says it does |
|---|---|
| **Ask permissions** | Confirm file edits and higher-risk commands when CLI asks |
| **Auto accept edits** | Claude writes to disk without asking |
| **Auto mode** (enable once) | Claude reviews tool calls and runs actions it considers safe |
| **Plan mode** | Architecture & reasoning only, no files |
| **Bypass permissions** (high risk) | Full tool access for shell and file system |
**Leave it on "Ask permissions" for your first run.** Every file write and every risky command stops and asks, so you can see exactly what it intends to do. Loosen it later, once you know how it behaves.
What the other four are for:
- **Auto accept edits** opens up file writes only; commands still prompt. Good once you're confident about the scope of the change.
- **Auto mode** has to be enabled once by hand. After that Claude reviews each tool call itself, running what it judges safe and blocking what it judges risky. **It reduces prompts; it does not guarantee safety.** Use it in isolated environments only.
- **Plan mode** never touches a file — it only produces a plan. Use it for read-only investigation, or to make it explain its approach before it starts.
- **Bypass permissions** removes every check. Shell and filesystem are wide open, it can delete anything, and it won't ask. **Never use it in a directory holding production credentials, personal data, or anything not under version control.**
:::danger
"Bypass permissions" is not just a slightly looser "Auto mode". Auto mode at least keeps Claude's own review in the loop; bypass removes even that.
:::
Permission modes are locked while a turn is running — what the UI shows and what's actually enforced have to match. Wait for the turn to finish, or press `Cmd/Ctrl + .` to stop it.
## 3. Say what you want
Plain language is fine; you don't need to write a spec. What matters is that the goal is **verifiable** — that you'll know how to check whether it got there.
For example:
```text
Read this project, then add a "sort by due date" toggle to the task list.
Put it to the right of the heading; clicking it switches between manual
order and due date, with undated items last. Tell me which files you changed.
```
If you'd rather ease in, ask for something read-only first: "Read this project and tell me what it does and how to start it. Don't change anything."
Press `Enter` to send (`Shift + Enter` for a newline).
## 4. Watch it work
![A full turn: collapsed tool calls, thinking, a permission prompt, an inline diff](../../images/app/session-main.webp)
It orients itself before it starts editing, and you'll see several kinds of card go by:
- **Tool call cards** ("Searched files, ran one command, read 4 files") — collapsed by default; expand to see exactly what it read and ran.
- **Thinking blocks** — its reasoning about the next step.
- **Permission prompts** ("Allow Claude to Edit index.html?") — these stop and wait for you. The card includes a preview of the change: green for added lines, red for removed. **Read the diff before you decide.**
Three buttons:
- **Allow** — this one time only.
- **Allow for session** — stop asking about this kind of operation for the rest of the session. Once the scope is clear, this saves a long string of repeated clicks.
- **Deny** — send it back. It will try a different approach, or ask you for more information.
To stop mid-run, use the stop button at the bottom right of the composer or press `Cmd/Ctrl + .`.
## 5. Review file by file
After each edit lands, the session shows an **inline diff**: the file path, an `+8 / -3` line count, and old and new lines side by side with syntax highlighting. That's the right view for following a single change as it happens.
Once a turn finishes and the changes pile up, open the **workspace** panel on the right ("Show Workspace" at the top right of the tab bar). It collects everything changed this turn into a list you can open for a full review, comment on specific lines, and send those comments — with their code location attached — straight back into the composer for another round.
The full workspace walkthrough is in [Workspace](../desktop/workspace.md).
:::warning
Denied edits never reach the disk, but the disk and `git diff` are the only source of truth. Run `git status` and `git diff` yourself before you ship anything — don't take the UI's word for it.
:::
## Next
- See what else it can do — [desktop feature map](../desktop/index.md)
- Keep going from your phone — [Phone and IM handoff](../desktop/remote.md)
- Something got stuck — [Won't install, won't open, won't connect](./troubleshooting.md)
+56
View File
@@ -0,0 +1,56 @@
---
title: What Claude Code Haha is
nav_title: What it is
description: An AI coding workbench that runs on your own machine. You pick the model; you approve every change.
order: 0
---
# What Claude Code Haha is
It's an app on your computer. You hand it a project folder, describe what you want in plain language, and it goes off to read the code, edit files, and run commands — with every change laid out in front of you, waiting for your approval.
![A full session: prompt, tool calls, file edits, inline diff](../../images/app/session-main.webp)
That's a real session. Projects and history on the left, the conversation in the middle, and when Claude edits a file the diff appears right underneath, line by line.
## How this relates to the Claude Code CLI
Claude Code is Anthropic's command-line coding agent. The engine inside Claude Code Haha is a CLI built from repaired Claude Code sources (it's called `claude-haha` in this repo), and the desktop app is the graphical shell wrapped around it.
Two practical consequences:
- **You don't install Claude Code first.** The CLI engine ships inside the installer. No Node.js, no npm, no global commands.
- **Nothing is missing.** Permission prompts, subagents, Skills, MCP, memory — it's the same machinery, just shown as an interface instead of scrolling terminal output.
If you'd rather stay in the terminal, the CLI is still there: see [Command line](../cli/index.md).
## What it does for you
**Writes code.** Describe a goal — add a feature, fix a bug, restyle a page — and it finds the files, reads the surrounding context, makes the edits, and tells you which files it touched.
**Shows its work.** Every edit comes with an inline diff, and the workspace panel on the right collects everything changed this turn into a list you can open file by file. Don't like it? Send it back.
**Delegates.** Big tasks can be split across subagents running in parallel, with their progress visible in the activity panel. You can also give each agent its own model, tools, and system prompt.
**Runs on a schedule.** Tidy up logs every morning, audit dependencies every week — set a job on a schedule and it clocks in on its own, leaving a record of each run.
**Follows you to your phone.** Turn on H5 access, scan a QR code, and pick the conversation back up on your phone. Or connect Telegram, Feishu, or WeChat and drive it from a chat window.
## Three commitments
**Local first.** Sessions, settings, memory, and skills live on your machine (under `~/.claude` by default). No accounts, no cloud sync, no uploading your code. The only outbound traffic goes to the model service you configured yourself.
**Your choice of model.** Nothing is locked to one vendor. Sign in with a Claude, ChatGPT, or Grok account; use a built-in preset for DeepSeek, Kimi, or Zhipu GLM; or point it at a local model running in LM Studio or Ollama and pay nothing at all.
**You approve the changes.** The default is "Ask permissions" — it stops and asks before writing a file or running a risky command. Loosen it to auto-accept edits, or tighten it so it can only plan and never touch a file. Five levels, switchable any time.
## Start here
Work through these in order; about twenty minutes gets you to a working first session.
1. [Download and install](./install.md) — installers for all three platforms, and what to do when the OS blocks them.
2. [Connect a model](./models.md) — official accounts, third-party APIs, or local models. Pick one.
3. [Run your first session](./first-session.md) — pick a folder, set permissions, state a goal, watch it work, review the diff.
4. [Desktop feature map](../desktop/index.md) — once it's running, see what else is in the box.
Stuck along the way? [Won't install, won't open, won't connect](./troubleshooting.md) is organized by symptom.
+121
View File
@@ -0,0 +1,121 @@
---
title: Download and install
nav_title: Install
description: Installers for macOS, Windows, and Linux, plus what to do when the OS blocks them.
order: 1
---
# Download and install
Install and go. You don't need Node.js, Python, or Claude Code — the CLI engine and the ripgrep binary used for file search are both bundled inside the installer.
## Pick the right package
Everything lives on [GitHub Releases](https://github.com/NanmiCoder/cc-haha/releases/latest). Choose by operating system and CPU architecture:
| Your system | Download |
|---|---|
| macOS, Apple Silicon | `Claude-Code-Haha-<version>-mac-arm64.dmg` |
| macOS, Intel | `Claude-Code-Haha-<version>-mac-x64.dmg` |
| Windows x64 | `Claude-Code-Haha-<version>-win-x64.exe` |
| Windows ARM64 | `Claude-Code-Haha-<version>-win-arm64.exe` |
| Linux x64 | `Claude-Code-Haha-<version>-linux-x86_64.AppImage` or `-linux-amd64.deb` |
| Linux ARM64 | `Claude-Code-Haha-<version>-linux-arm64.AppImage` or `-linux-arm64.deb` |
Not sure which architecture you have? On macOS check the chip listed in "About This Mac"; on Windows check the system type under Settings → System → About. Don't guess from the brand of the machine.
The `.blockmap` and `latest*.yml` files are used by the app's own updater. You don't need to download them.
## macOS
1. Open the DMG.
2. Drag Claude Code Haha into Applications.
3. Launch it from Applications.
### If macOS says the app is damaged
Nothing is actually damaged. macOS quarantines anything downloaded from the web and refuses to launch it without an Apple signature — but it words the error as "damaged", which is thoroughly misleading. Signed and notarized releases skip this entirely; if you have an unsigned build, use one of the two routes below.
**Option 1: the official script (recommended)**
Download `install-macos-unsigned.sh` from the same Release into **the same folder as the DMG** (Downloads, for example), then run:
```bash
cd ~/Downloads
bash install-macos-unsigned.sh
```
The script picks the DMG matching your architecture, mounts it, installs the app into `/Applications`, strips the quarantine attribute, and launches it. Any existing install is moved to the Trash first rather than overwritten in place.
**Option 2: clear the quarantine flag yourself**
If the app is already in Applications:
```bash
xattr -dr com.apple.quarantine "/Applications/Claude Code Haha.app"
```
Only do this for packages you have confirmed came from this repository's Releases. Never bypass Gatekeeper for software of unknown origin.
## Windows
1. Fully quit any running copy of the old version, including the system tray icon.
2. Double-click the `.exe`.
3. **Don't** right-click and choose "Run as administrator" — the installer is per-user, and running it elevated puts your data directory in the wrong place.
Unsigned packages trigger a SmartScreen warning. Once you've confirmed the file came from this repository's Releases, click "More info" → "Run anyway".
When upgrading in place, the installer inspects user data in the old install directory. If it reports that the program is still running, quit the main window and the tray icon, give the background sidecar, terminal, and IM adapter processes a few seconds to exit, then run the installer again. Don't delete the old install directory by hand first.
## Linux
**AppImage** (no installation, just run it):
```bash
chmod +x Claude-Code-Haha-<version>-linux-x86_64.AppImage
./Claude-Code-Haha-<version>-linux-x86_64.AppImage
```
If it fails with a FUSE-related error, install the runtime: `sudo apt install libfuse2` on Ubuntu 22.04 and earlier, `libfuse2t64` on 24.04 and later.
**deb** (installs into your application menu):
```bash
sudo apt install ./Claude-Code-Haha-<version>-linux-amd64.deb
```
On ARM64 machines, use the corresponding `linux-arm64` file.
## Running from source
If you want to modify the code, debug the engine, or just use the CLI in a terminal:
```bash
git clone https://github.com/NanmiCoder/cc-haha.git
cd cc-haha
bun install
cp .env.example .env
./bin/claude-haha
```
Requires [Bun](https://bun.sh) and Git. This runs the CLI only; for building the desktop app and configuring the local server, see [Command line](../cli/index.md).
## Updating
**In-app updates (recommended).** Open Settings → About → App Updates and click "Check now". It compares your installed version against the latest GitHub Release, downloads the new build, and offers "Install and restart".
Before updating, stop any running sessions and save uncommitted work.
If the download stalls, you probably can't reach GitHub. The same panel has an "Advanced update proxy" setting where you can switch to the system proxy or enter a local HTTP proxy address (`http://127.0.0.1:7890`, for instance). This proxy only affects the app's own update downloads — it has no effect on model requests.
**Manual replacement.** Download the new installer from Releases and repeat the steps for your platform. Sessions, provider configuration, skills, agents, and memory live under `~/.claude`, not in the application directory, so installing over the top doesn't touch them.
:::warning
The installer's data protection is not a backup. Keep your own copy of anything you can't afford to lose.
:::
## Next
Go to [Connect a model](./models.md). Until a model is connected, the app opens but can't send a single message.
If it won't install or won't open, see [Won't install, won't open, won't connect](./troubleshooting.md).
+124
View File
@@ -0,0 +1,124 @@
---
title: Connect a model
nav_title: Connect a model
description: Official accounts, third-party APIs, or a local model — pick one and be running in ten minutes.
order: 2
---
# Connect a model
Claude Code Haha ships without a model. It's the shell that does the work; you have to give it a brain first.
Click "Settings" at the bottom of the sidebar, then pick the first tab, "Providers". From there you have three routes:
- **You have an official account** — Claude, ChatGPT, and Grok each have a built-in card. One click opens a browser sign-in. No API key to type.
- **You have a third-party API key** — DeepSeek, Kimi, Zhipu GLM and others come as presets. Paste the key and you're done.
- **You want it free** — run LM Studio or Ollama on your own machine. The model runs on your GPU, costs nothing, and works offline.
You can configure all three and switch between them in the provider list.
## Sign in with an official account
Three cards sit at the top of Settings → Providers:
| Card | What you need |
|---|---|
| **Claude Official** | A Claude.ai account (Pro / Max subscription, or an API account with credit) |
| **ChatGPT Official** | A ChatGPT account, via OpenAI OAuth |
| **Grok Official** | An xAI account, via official Grok OAuth |
Click the sign-in button on the card ("Sign in to Claude" / "Sign in with ChatGPT" / "Sign in with Grok"). Your system browser opens the provider's authorization page; complete it with the same account and the browser hands you back to the app, where the card now reads as signed in.
Three things to watch for:
- Leave Claude Code Haha running for the whole flow — the callback has to land in the running app.
- If the browser doesn't open by itself, click "Copy authorization link" and paste it in manually.
- Proxies and blocking browser extensions can intercept either the authorization page or the local callback. Turn them off before retrying.
Which models you get afterwards depends on your account tier and entitlements, not on the app. A model name appearing in documentation doesn't mean your account can call it.
## Third-party API providers
With an API key in hand, this is the fastest route. Click "Add Provider", pick something under "Preset", and the base URL and default models are filled in for you — all you supply is the key.
The built-in presets, as they appear in the dialog:
- **DeepSeek** · **Zhipu GLM** · **Kimi** · **MiniMax** — major Chinese model vendors; the base URLs point at each one's Anthropic-compatible endpoint.
- **接口AI (JiekouAI)** · **胜算云 (Shengsuanyun)** · **TeamoRouter** — routing services that give you access to official Claude models through their own gateway.
- **LM Studio** · **Ollama** — local models; see the next section.
- **Custom** — anything not listed above.
When a preset has a signup page for API keys, a "Get API Key" button appears under the key field.
## Local models
To spend nothing and stay offline, run a model server on your machine and point the app at it.
**LM Studio**: load a model, start the local server, then choose the `LM Studio` preset in "Add Provider" with base URL `http://localhost:1234`.
**Ollama**: run `ollama serve`, choose the `Ollama` preset, and use base URL `http://localhost:11434`.
Two hard requirements:
1. **Do not append `/v1` to the base URL.** Both of these expose an Anthropic-compatible protocol and the app uses that path. Adding `/v1` gives you a straight 404.
2. **Raise the context window — at least 200K.** Claude Code's system prompt, tool definitions, and Skills consume a substantial amount of context before your first message. A default 4K or 8K window can't even hold the opening. Change this in LM Studio or Ollama's own model settings, not in this app.
Whether a local model can actually sustain an agent workflow comes down to its tool-calling ability. Small models often talk endlessly without ever calling a tool — that's the model, not your configuration.
## The Add Provider dialog, field by field
![Add Provider dialog: preset, base URL, auth variable, API key, model mapping](../../images/app/settings-provider-add.webp)
**Name** (required) — how this provider appears in the list. A preset fills it in; rename it to something you'll recognize, like "DeepSeek — work account".
**Notes** — a line for yourself, e.g. "expires end of month". Doesn't affect any request.
**Base URL** (required) — the provider's API root, not its marketing site. Presets get this right. When entering it yourself, watch for duplicated path segments: if the URL already ends in `/anthropic`, don't also append `/v1/messages`.
**Auth Variable** — decides how your key is transmitted. Five options:
| Option | When to use it |
|---|---|
| API Key (`ANTHROPIC_API_KEY`) | Direct Anthropic API access, sends an `x-api-key` header |
| Bearer Token (`ANTHROPIC_AUTH_TOKEN`) | Nearly all third-party Anthropic-compatible services, sends `Authorization: Bearer` |
| Bearer + Empty API_KEY | OpenRouter, Ollama and similar, which must not fall back to an Anthropic key |
| Write token to both variables | Services like Hugging Face Router that check both |
| Write dummy to both variables | Local vLLM-style services that only need placeholder auth |
Presets pick the right one. If you're guessing, start with Bearer Token and switch to API Key if you get a 401.
**Enable Tool Search** (on by default) — loads MCP tools and part of the tool definitions on demand, saving a large chunk of schema tokens on the first turn. It depends on the model supporting `tool_reference`. **Turn it off for weaker models, or for providers that reject this request shape.**
**Disable experimental beta headers** — sets `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` for this provider. Some third-party gateways error out on beta-shaped API requests; checking this stops the headers being sent. **Try this first when you see a 400 or an "unsupported parameter" error.**
**API Key** (required except for local models) — copied from the provider's console. Watch for leading or trailing whitespace when pasting. The key stays on your machine.
**Model Mapping** — four slots, each taking a **model ID**, not a product name:
- **Main Model** (required) — the model used for conversation. It must support tool calling.
- **Haiku Model** — leave blank to follow the main model. This slot handles lightweight work like title generation and file summaries; mapping it to a cheap small model saves real money.
- **Sonnet Model / Opus Model** — leave blank to follow the main model. Fill these in only when you actually want tiering.
Each slot has a `1M` checkbox. Tick it only if that model genuinely supports a one-million-token context window.
**API Format** — this field only appears when you pick the `Custom` preset or edit an existing provider. Three options: Anthropic Messages (native), OpenAI Chat Completions (proxied), and OpenAI Responses API (proxied). The proxied formats are translated by a loopback proxy the app starts locally — you don't deploy anything. Every built-in preset uses Anthropic native, which is why the field is hidden for them.
Click "Add" when you're done.
## After saving
Back in the provider list, on the entry you just created:
1. Click "Test". Anthropic-native providers run one step, "① Connectivity"; OpenAI formats add "② Proxy pipeline". Both must pass.
2. Click "Set default" so new sessions use it.
3. Multiple providers can be dragged to reorder. Order only affects how the list is displayed.
Then start a new session and **pick the specific model from the model selector at the bottom right of the composer** — that list reflects what the active provider actually offers. The control next to it sets reasoning effort; leave it at the default if you're unsure.
:::tip
A passing test isn't a guarantee. It proves the endpoint is reachable and the credentials work — not that the model can sustain tool calls and long context. The real check is asking for a task that edits a file, and seeing whether it actually does.
:::
Getting a 401, a connection failure, or an empty model list? See the model section of [Won't install, won't open, won't connect](./troubleshooting.md).
With a model connected, go [run your first session](./first-session.md).
+211
View File
@@ -0,0 +1,211 @@
---
title: Won't install, won't open, won't connect
nav_title: Troubleshooting
description: Organized by symptom — install failures, blank window, model 401s, stuck sessions, port conflicts, phone access.
order: 4
---
# Won't install, won't open, won't connect
Find your symptom below. Each entry is "what you see → why → what to do".
First, one check: make sure you're on the latest stable build from [GitHub Releases](https://github.com/NanmiCoder/cc-haha/releases/latest). A lot of problems on older versions are already fixed.
## Won't install
### macOS says the app is damaged and can't be opened
**Why** — The file isn't damaged. macOS quarantines downloads and refuses to launch anything without an Apple signature, but words the error as "damaged", which sends everyone down the wrong path.
**What to do** — Download `install-macos-unsigned.sh` from the same Release, put it in the same folder as the DMG, and run `bash install-macos-unsigned.sh`. Or, if the app is already in Applications, run `xattr -dr com.apple.quarantine "/Applications/Claude Code Haha.app"`. Full details in [Download and install](./install.md).
### Windows shows a SmartScreen warning
**Why** — Unsigned installers get flagged by SmartScreen.
**What to do** — Confirm the file came from this repository's Releases, then click "More info" → "Run anyway". If the filename or origin doesn't match, don't bypass it.
### The Windows installer says the program is still running
**Why** — The old version's main process, sidecar, embedded terminal, or IM adapter hasn't fully exited, so the installer won't overwrite it.
**What to do**
1. Quit the main window, and quit the tray icon too.
2. Give background processes a few seconds to exit.
3. Still stuck? End any remaining Claude Code Haha processes in Task Manager.
4. Run the installer again. **Don't** use "Run as administrator", and **don't** manually delete data from the old install directory.
### The Linux AppImage does nothing when I run it
**Why** — Usually a missing execute bit, or missing FUSE.
**What to do** — Run `chmod +x <name>.AppImage` first. If you still get a FUSE error, install `libfuse2` on Ubuntu 22.04 and earlier, or `libfuse2t64` on 24.04 and later.
## Won't open
### Nothing happens when I launch it
**Why** — Most often the wrong CPU architecture: `mac-x64` on Apple Silicon, or `win-x64` on ARM64 Windows.
**What to do** — Re-check your architecture against [Download and install](./install.md) and reinstall the right package. An old version still running in the background can also block startup, so quit everything first.
### The window opens but stays blank
**Why** — The interface assets didn't load, usually after an interrupted upgrade or because a graphics driver issue prevented the renderer process from starting.
**What to do**
1. Fully quit the app (not just close the window) and reopen it.
2. Still blank? Reinstall the same version over the top. Sessions and configuration live under `~/.claude`, not in the application directory, so nothing is lost.
3. Still blank after that, it's failing during startup. File an issue at [GitHub Issues](https://github.com/NanmiCoder/cc-haha/issues) with your OS version, CPU architecture, and installer filename.
:::warning
Never delete `~/.claude` while troubleshooting. Your sessions, provider configuration, skills, agents, and memory are all in there, and they don't come back.
:::
### My session list is empty after updating
**Why** — Almost certainly not data loss. The sidebar list is served by a local SQLite index, which is derived data that can be rebuilt; the original sessions are still stored as JSON / JSONL files. While the index rebuilds, the list looks empty.
**What to do** — Open Settings → Diagnostics → Local index and check whether it's still building; let it finish. If there are degraded sources or an error code, click "Rebuild local index" — that only rebuilds the index and leaves conversations and settings untouched.
## Model 401s and connection failures
### 401 / 403 / "API Key invalid"
**Why** — The auth method doesn't match the provider, or the key itself is wrong.
**What to do** — Open Settings → Providers, edit the entry, and check in order:
1. Is **Base URL** the API root rather than the marketing site?
2. Is **Auth Variable** correct? Third-party Anthropic-compatible services almost always want `Bearer Token (ANTHROPIC_AUTH_TOKEN)`; only direct Anthropic access uses `API Key (ANTHROPIC_API_KEY)`. If unsure, try both.
3. Did the **API Key** pick up stray whitespace, or has it expired or run out of credit?
4. Save, click "Test", and verify with a fresh session once it passes.
Don't hand-edit `settings.json` for any of this — changes made in the app sync automatically.
### 404, or "model not found"
**Why** — A duplicated path segment in the URL, or a model ID borrowed from another platform.
**What to do** — Check the base URL for repetition (if it already ends in `/anthropic`, don't append `/v1/messages`). Use the exact model ID string from the provider's console, not the product name.
For local models (LM Studio / Ollama), **do not append `/v1` to the base URL** — that's the single most common source of 404s.
### 400, or "unsupported parameter"
**Why** — A third-party gateway is rejecting beta-shaped API requests, or the model doesn't support `tool_reference`.
**What to do** — Edit the provider and check "Disable experimental beta headers". If that isn't enough, turn off "Enable Tool Search". Both switches exist for exactly these gateways.
### Signed in, but the model I want isn't in the selector
**Why** — Which models you see is determined by your account entitlements, region, and plan — not by the app.
**What to do** — Refresh the provider status and reopen the model selector. If it's still missing, trust the provider's own console over anything else.
### It talks endlessly but never edits a file
**Why** — The model's tool-calling isn't strong enough to sustain an agent workflow. Small local models do this a lot.
**What to do** — Try a model with solid function-calling support. Also confirm you aren't in Plan mode, which forbids touching files by design.
### OAuth sign-in hangs
**Why** — The authorization callback isn't reaching the running app.
**What to do** — Keep the app running through the whole flow; complete the authorization in your system browser with the same account; disable proxies and blocking extensions; and check that your system clock is correct, since a skewed clock breaks the handshake. If the browser never opens, click "Copy authorization link" and paste it manually.
## Sessions that hang
### It spins forever without producing anything
**Why** — Could be a network hiccup, or an upstream that never properly closed the response.
**What to do** — In this order; don't jump straight to restarting:
1. Wait ten or fifteen seconds in case it's transient.
2. Click stop, or press `Cmd/Ctrl + .`.
3. Open the activity panel and check the real state of background tasks and subagents — a quiet main conversation doesn't mean the background is idle.
4. Switch to another session and back to rule out a stale view.
5. If none of that helps, fully quit and reopen the app.
6. Copy the error summary from Settings → Diagnostics.
### I can't change permission mode mid-session
**Why** — This is deliberate. Changing permissions mid-turn would leave the UI showing something different from what's enforced.
**What to do** — Wait for the turn to end, or stop it first, then switch.
### I denied an edit and I'm not sure whether the file changed
**Why** — Denied writes never reach the disk, but what the UI displays and what's on disk are separate things.
**What to do** — Run `git status` and `git diff` before shipping. Disk and Git are the source of truth.
## Port conflicts
**What you see** — H5 won't load, or the local service doesn't come up after launch.
**Why** — The local server defaults to port `3456`. If something else already holds that port, the service moves elsewhere or fails to start.
**What to do** — Open Settings → H5 Access and read "Current port" — it may have already moved, leaving your old QR code pointing at the wrong place. If you need a stable address (for a bookmark or a reverse proxy), set "Fixed port" on the same page, anywhere in 1024–65535. Port changes apply after restarting the app.
## Phone won't connect
**What you see** — Scanning the QR code opens nothing, or you get an unauthorized error.
**Why** — Address, port, token, and network all have to line up. Any one of them being wrong breaks the connection.
**What to do** — Go through Settings → H5 Access:
1. Is H5 access actually switched on?
2. Does "Access host / IP" still match your computer's current network interface? It changes when you switch Wi-Fi.
3. Does the port in the QR code match "Current port"?
4. Are the phone and computer on the same local network?
5. Does your firewall allow that port?
6. Has the token been regenerated? **The moment you regenerate it, every old QR code is dead.**
If you changed the fixed port, restart the app. Full deployment guidance and security boundaries are in [Phone and IM handoff](../desktop/remote.md).
### Does locking my phone kill a running task?
No. A brief disconnect doesn't stop work in progress — it finishes in the background and you'll see the result when you reconnect. Only when a task is already idle *and* no client is connected does the disconnect grace timer stop the corresponding CLI (30 seconds by default).
That's not a promise it will never drop, though. System sleep, process exit, proxy failures, and service restarts all still end the connection.
### I scanned the IM QR code but my contacts still can't talk to it
Scanning only binds the platform account; it doesn't authorize everyone who can message you. Each person still has to send the one-time pairing code generated in the desktop app, or be added to the allowlist. With both empty, access is denied by default. Per-platform differences are in [IM integrations](../im/index.md).
## Computer Use does nothing
**What you see** — You ask it to click something or control another app, and it says it can't, or simply doesn't move.
**Why** — Computer Use has a chain of prerequisites and won't run unless every link passes. It supports macOS and Windows only — **there is no Linux support**.
**What to do** — Open Settings → Computer Use and find the first item that isn't green:
1. Is the toggle at the top on? (With it off, new sessions never get these tools at all.)
2. Did the Python 3 check pass? If not, install it, or point "Python Interpreter Path" at one you already have — conda and pyenv both work.
3. Are the virtual environment and dependencies ready and installed? If not, click "Install Environment".
4. On macOS, both "Accessibility Permission" and "Screen Recording Permission" must show as granted. Grant them under System Settings → Privacy & Security.
5. **Restart Claude Code Haha after granting them.** System permissions don't apply to an already-running process.
6. Is the app you want to control listed under "Authorized Apps"?
Full details in [Computer Use](../desktop/computer-use.md).
## Still stuck
Go to Settings → Diagnostics:
1. Click "Copy issue report" for a structured snapshot of the current state.
2. Search [GitHub Issues](https://github.com/NanmiCoder/cc-haha/issues) for the same problem before opening a new one.
3. If the report alone isn't enough to diagnose it, click "Export Bundle" and attach that too.
Including these makes a fix much faster: app version, OS and CPU architecture, installer filename, which kind of provider you're using (**never paste an API key**), the shortest reproduction steps, the full error text, and whether the problem is in the desktop app, on the phone, or in the CLI.
:::warning
Diagnostic reports make a real effort to omit chat content, file contents, full environment variables, and API keys — but they can still include local paths and provider hostnames. **Read one before you share it.**
:::