1209 Commits

Author SHA1 Message Date
程序员阿江(Relakkes)
d7a1306790 fix(desktop): let memory markdown editor fill pane
Tested: cd desktop && bun run test -- src/__tests__/memorySettings.test.tsx --run --reporter=verbose --pool=forks --maxWorkers=1 --minWorkers=1
Tested: bun run check:desktop
Tested: bun run check:coverage
Tested: Chrome CDP smoke confirmed the memory editor textarea expands to 592px tall
Confidence: high
Scope-risk: narrow
2026-06-10 23:47:57 +08:00
程序员阿江(Relakkes)
66306a54bf fix(desktop): capture preview screenshots natively
Use Electron WebContentsView capturePage for browser preview screenshots and selection annotations so captured images match the rendered preview.

Tested:

- cd desktop && bun run test -- electron/services/preview.test.ts src/lib/previewEvents.test.ts src/components/browser/BrowserSurface.test.tsx src/preview-agent/screenshot.test.ts src/preview-agent/picker.test.ts src/preview-agent/editBubble.test.ts

- cd desktop && bun run lint

- cd desktop && bun run build

- bun run check:desktop

- bun run check:electron

- bun run check:native

Not-tested:

- Real GUI click-through for Screenshot / Select Element; Electron runtime launch was blocked in this shell and packaged build did not enter the smoke path.

Constraint: GUI click-through smoke was blocked by local Electron launch behavior; package smoke passed but does not launch the app.

Confidence: medium

Scope-risk: narrow
2026-06-10 22:49:32 +08:00
小橙子
2b6202d381
fix(handoff): reuse empty session + ship raw tail for v0.5.8 (#5)
Two bugs in the v0.5.7 "Continue from here" hand-off path landed in
the wild and the user caught them while exercising 0.5.8 prep:

1. **Stray empty session.** When the user previously opened a "New
   session in X" tab from the sidebar, then closed the tab to
   declutter, the freshly-created empty session was already on disk
   and visible in the sidebar. Coming back to the welcome screen
   (activeTabId = null → EmptySession route) and clicking "Continue
   from here" called `createSession` unconditionally, minting yet
   another empty session next to the now-stale one. The user ended up
   with `Untitled Session 27 minutes ago` lingering in the sidebar
   while the hand-off ran in a completely separate fresh session.

   Extract the picker into `desktop/src/lib/sessionReuse.ts` —
   `pickReusableEmptySession(sessions, workDir, excludeSessionId?)`
   filters by exact-workDir match, `messageCount === 0`, excludes the
   previous session being handed off FROM, and sorts by `modifiedAt`
   desc. Caller in `EmptySession.onAutoHandoff` runs this before
   `createSession`; on hit, openTab on the existing sessionId and let
   ContentRouter switch to ActiveSession naturally; on miss, create
   fresh as before.

   7 unit tests cover the boundaries: workDir mismatch, non-zero
   messageCount, excludeSessionId honored, null workDir treated as
   "no candidates" (so we never silently merge home-dir sessions
   into a project hand-off), empty workDir argument, empty session
   list, freshness sort order.

2. **Summary too abstract; AI doesn't know specific recent state.**
   The previous implementation only injected the LLM-summarized
   `main` + `recent` paragraphs into the next session's system
   prompt. That summarization tends to wash out exact wording, file
   paths, error messages, and the user's literal last question — the
   user reported "AI doesn't know what we just hit a wall on."

   Add `recentRaw?: string` to `SessionSummary`. New helper
   `buildRecentRawSlice` keeps the LAST ~12 turns (env-tunable via
   `CLAUDE_CODE_HANDOFF_RAW_TURNS`, range 0–50, default 12), each
   truncated to 400 chars preserving the `USER:` / `ASSISTANT:`
   role prefix, total capped at ~8000 chars (env-tunable via
   `CLAUDE_CODE_HANDOFF_RAW_CHARS`, range 500–30000, default 8000).
   Older turns drop first when the total cap is busted by dense
   recent turns.

   `formatHandoffSystemPrompt` renders the raw slice inside a fenced
   code block AFTER the abstracted recent summary, so the next
   session sees both the digest and the literal text. Section is
   omitted entirely when `recentRaw` is absent — v0.5.7 caches still
   load and produce the exact same prompt bytes (back-compat).

   8 unit tests cover the formatter's back-compat path, the raw
   block placement, fence rendering, tail-N selection, all-fits
   case, per-turn truncation with role prefix preservation, total
   cap busting (older drops first), and the `RAW_TURNS=0` opt-out.

Both fixes fold into the unreleased v0.5.8 — release-notes/v0.5.8.md
gets two new fix bullets and the 范围 / 验证 sections call out the
new files and 15 added tests.

desktop/package.json bump to 0.5.8 stays as-is (preset earlier in
the same worktree before this commit).

Tested:
- bun run lint (desktop) clean
- bunx vitest run sessionReuse.test.ts → 7/7 pass
- bun test sessionSummaryService.test.ts → 8/8 pass
- bunx vitest run EmptySession.test.tsx → 22/22 pass (regression check
  after the helper extraction — the existing tests don't go through
  the hand-off branch but exercise the rest of the file)

Confidence: high
Scope-risk: narrow (no shape changes to any persistence — recentRaw
is an additive field, old caches deserialize cleanly because
`tryParseSummaryResponse` already only required main+recent)

Co-authored-by: 你的姓名 <you@example.com>
2026-06-10 20:46:46 +08:00
小橙子
65a042bd08
fix: detect & isolate fake tool_use leaks from non-Anthropic providers (#4)
* feat(plugin): add reverse-engineering plugin v0.4.0 + project-level codegraph

Bundles a multi-platform RE toolkit as a local marketplace plugin that
clone-and-enable installs in cc-haha desktop. Covers static (Ghidra,
radare2, JADX, apktool, binwalk via Bash) and the gap a previous review
called out: dynamic debugging is the lane that matters most for
AI-driven RE, but Frida alone cannot single-step or set real
breakpoints. This commit fills that gap with three non-overlapping
dynamic lanes -- Frida (instrumentation, mobile, broad surveys), GDB
(cross-arch single-step + breakpoints, embedded firmware), LLDB (Apple
platforms, ObjC/Swift).

Architecture coverage spans MIPS / ARM (incl. Cortex-M Thumb) /
PowerPC (incl. e200 VLE) / 68k (incl. Mac Toolbox A-line traps) /
SuperH / RISC-V / x86-64 / AArch64. The agent picks the right ISA,
base address, and load configuration in skills/firmware-blob, then
hands off to skills/pe-elf-macho with per-architecture decompilation
notes (MIPS delay slots, ARM Thumb interworking, PowerPC TOC/SDA,
68k register banks, etc).

Plugin layout (.claude-plugin marketplace + plugin manifests):

  plugins/.claude-plugin/marketplace.json
  plugins/reverse-engineering/
    .claude-plugin/plugin.json     v0.4.0, GHIDRA_INSTALL_DIR + ARTIFACT_DIR userConfig
    agents/reverse-engineer.md     triage -> static -> optional dynamic -> report
    skills/triage/                 file-type identification, packing, routing
    skills/pe-elf-macho/           Ghidra/r2 + per-arch notes for non-x86
    skills/firmware-blob/          raw blob -> ISA + base address + Ghidra/r2 load config
    skills/apk-analysis/           JADX + apktool, Android attack surface
    skills/ios-analysis/           IPA bundle + Mach-O, FairPlay-aware
    skills/dynamic-debug-overview/ decision matrix Frida vs GDB vs LLDB
    skills/frida-dynamic/          hooks + Memory + CpuContext + Stalker + watchpoints
    skills/gdb-debug/              cross-arch via gdbserver / qemu-user / qemu-system
    skills/lldb-debug/             macOS / iOS device / Linux, ObjC/Swift symbols
    skills/crackme-keygen/         CTF / self-owned binary; ships keygen, not patch
    skills/re-report/              final structured report with IOCs + confidence
    commands/triage.md             /reverse-engineering:triage <path>
    commands/report.md             /reverse-engineering:report <sample-id>
    mcp/servers.json               7 MCP servers: ghidra (pyghidra-mcp PyPI),
                                   radare2 (npm), gdb (mcp-gdb npm), lldb
                                   (stass/lldb-mcp git), jadx, apktool, frida
                                   (kahlo-mcp git, all @main pinned)
    hooks/hooks.json               empty placeholder
    scripts/validate.ts            offline schema check via project's
                                   validatePluginManifest + validatePluginContents
    scripts/dev-link.ts            Windows mklink /J cache -> repo source so SKILL
                                   edits don't require version bump (--restore undoes)
    scripts/smoke.ts               end-to-end: marketplace register -> enable ->
                                   update -> reload -> assert detail counts.
                                   Counts derived from on-disk source (glob),
                                   not hardcoded; idempotent across reruns.
    README.md                      install, prerequisites per MCP, dynamic-capabilities
                                   matrix, architecture coverage, quickstart with
                                   busybox sample, dev-link / smoke workflows

Project-level MCP (clone-and-go for any contributor):

  .mcp.json                        codegraph (npx -y codegraph serve --mcp);
                                   project scope, picked up automatically.
                                   Pre-existing user-level codegraph install
                                   still wins for users who have one.

userConfig of the plugin is honest -- only knobs that actually do
something (GHIDRA_INSTALL_DIR substituted into the ghidra MCP env;
ARTIFACT_DIR resolved relative to agent cwd at run time). Earlier
ENABLE_* booleans were removed because they didn't actually toggle
anything; users disable individual MCP servers from Settings -> MCP
at runtime.

End-to-end verification (server on :3456, vite on :1420):
  - bun run plugins/reverse-engineering/scripts/validate.ts
    -> 0 fail / 0 warn (marketplace + plugin schema both clean)
  - bun run plugins/reverse-engineering/scripts/smoke.ts
    -> 13 passed / 0 failed (marketplace registration, enable, update
       0.3.0->0.4.0, reload, detail assertions including 11 skill ids
       matching the on-disk glob)
  - GET /api/plugins/detail
    -> commands=2 / agents=1 / skills=11 / mcpServers=7 / errors=[]
  - GET /api/mcp
    -> codegraph appears with scope=project,
       configLocation=.mcp.json
  - Desktop UI Settings -> Plugins
    -> v0.4.0 listed in enabled group, "11 skills 1 Agent 7 MCP"

External tool prerequisites (each MCP needs its underlying tool on
PATH; users install per-MCP independently -- none required all at
once): Ghidra (Java 17+), r2, GDB (gdb-multiarch for cross-arch),
LLDB, frida-tools, JADX (Java 17+), apktool. Codegraph requires
either a local global install (npm i -g codegraph) or relies on npx
to fetch on first run.

Constraint: Cannot single-step from Frida alone -- that's the gap GDB
and LLDB fill in this commit. dynamic-debug-overview SKILL is the
canonical decision guide for which lane to use.
Confidence: high
Scope-risk: narrow (additive plugin under plugins/, additive
.mcp.json; zero changes to existing source). Plugin loads via the
existing PluginManifestSchema path -- no new server-side code added.
Tested: bun run plugins/reverse-engineering/scripts/validate.ts (pass),
bun run plugins/reverse-engineering/scripts/smoke.ts (13/13 pass), live
chrome-devtools MCP smoke against running desktop UI verifying plugin
appears with correct version + counts.
Not-tested: real underlying-tool integration (Ghidra/r2/JADX/GDB/LLDB
must be installed by the user; the plugin only validates manifest
shape and component loading). Reverse-debugging (rr) flow is
documented in skills/gdb-debug but not exercised in smoke.

* fix: detect & isolate fake tool_use leaks from non-Anthropic providers

Some third-party Anthropic-compatible gateways (mimo, lgfzer, certain
proxies) don't relay native tool_use blocks: when the model wants to
call a tool, it ends up emitting an XML-style `<tool_use ...>{...}
</tool_use>` element as plain text inside content_delta. The desktop
chat renders that as markdown, marked passes the unknown element
through as raw HTML, the browser collapses whitespace, and the user
sees garbage like `<tool_useid="..."` followed by a smushed shell
command. Worse, the model has no signal the call never executed and
follows up with apologies ("工具用错了, 重来:") and retries — every
"call" is text and nothing runs.

Three layers of defense, lowest to highest in the stack:

1. **Source-side priming reduction** (server). The specialistRouter
   docstring already warned that `key="value"` attribute fragments in
   model-facing strings make some gateway models switch from real
   tool_use blocks into XML mode. Three model-facing nudges still
   embedded `subagent_type="verification"` in their reminders — fix
   them to use prose ("set the subagent_type parameter to verification")
   so we stop priming the failure mode:
     - src/utils/messages.ts (verification_gate_reminder)
     - src/tools/TodoWriteTool/TodoWriteTool.ts
     - src/tools/TaskUpdateTool/TaskUpdateTool.ts
   New regression test src/utils/noKeyValueNudges.test.ts pins all
   three call sites so a future "typo fix" doesn't accidentally
   re-introduce the attribute form.

2. **Detection + sanitization** (desktop). New
   desktop/src/lib/fakeToolUseDetection.ts extracts XML-style
   `<tool_use>` blocks (closed, half-open mid-stream, multiple
   consecutive retries, `name`/`id` in either order, quoted or
   unquoted, inside or outside fenced code blocks) and returns
   `{ cleanText, blocks[] }`. AssistantMessage uses this before
   handing content to MarkdownRenderer so the user never sees the
   XML garbage; copyText also uses cleanContent so users don't paste
   broken markup. Fenced code blocks (```xml ... ```) are preserved
   so legitimate documentation examples render as-is.

3. **Provider compatibility tracking** (desktop). New
   desktop/src/stores/providerCompatStore.ts counts fake tool_use
   leaks per provider id, persisted to localStorage. After 3 leaks
   it fires a one-time toast suggesting the user switch providers.
   Settings → Provider list shows a "工具调用异常" badge on offending
   providers; saving an edit on a flagged provider re-arms the
   warning (clearProvider). Above-threshold detection is silent
   afterward so we don't spam — the badge stays as the persistent
   signal.

The new FakeToolUseNotice component renders an inline grey card
above the assistant bubble naming the attempted tool ("Bash") and
making it clear nothing ran, so the user knows not to trust any
follow-up "Done." claim that depended on the call's output.

Verified:
- desktop: bun run lint clean
- desktop: 28 new tests + existing AssistantMessage.linkrouting (15
  total) pass
- server: 31 tests across verificationGate, specialistRouter, and the
  new noKeyValueNudges regression all pass

Tested: bun run lint (desktop), bunx vitest run (3 desktop suites),
        bun test (3 server suites)
Confidence: high
Scope-risk: narrow

---------

Co-authored-by: cc-haha <cc-haha@users.noreply.github.com>
Co-authored-by: 你的姓名 <you@example.com>
2026-06-10 20:01:32 +08:00
小橙子
9a0ae6a556
feat(desktop): make welcome task cards project-aware (#3)
Welcome-screen task cards previously rendered four hardcoded prompts
regardless of the actual repo state, so "PR pre-merge review" always
said `against main` and "fix failing test" / "write tests" always
showed `[fill in test name]` placeholders that the user had to edit
before sending.

Make each card's prompt resolve against live project context fetched
from `/api/projects/recent-activity`:

- preMergeReview now substitutes the actual `git.defaultBranch` so a
  branch off `develop` no longer asks for a `main` review.
- investigateTest auto-suggests a recently-edited test file (matched
  on `__tests__/`, `.test.*`, `.spec.*`, `_test.{py,go,rs}`) from the
  new `dirtyFiles[]` list returned by recent-activity. Falls back to
  the localized placeholder when no test is dirty.
- writeTests auto-suggests a recently-edited source file, skipping
  test files, lock files, build output, and binary assets.
- Card 4 renamed from "解释陌生代码" / "Explain unfamiliar code"
  → "了解项目" / "Understand this project". The prompt is now a
  project-overview brief (read CLAUDE.md/AGENTS.md/README, list tech
  stack, top-level dirs, current branch state) rather than a
  single-file explainer, which is what users actually want when they
  open a fresh session.
- Soften preMergeReview wording from "三件事都要做完" to
  "能合在一起做就一起做,你也可以根据改动量决定先做哪一件" so
  small-diff PRs aren't forced through the full three-step ritual.
- Each card button now has the FULL resolved prompt as its native
  `title` attribute, so users can hover to preview what they're about
  to send before clicking — no surprises.

Server-side, `deriveGitActivity` runs `git status --porcelain=v1 -z`,
parses the NUL-separated paths (handles renames/copies safely), stats
each file for mtime, sorts DESC, caps at `DIRTY_FILES_SAMPLE_LIMIT=20`.
Cached behind the existing recent-activity TTL so the cards don't add
filesystem load.

`WelcomeTaskCards` now fetches its own project context internally
given a `workDir` prop, instead of receiving a pre-resolved
`projectContext`. Both call sites (`EmptySession`, `ActiveSession`)
pass `workDir` (and `excludeSessionId` from `ActiveSession` so we
don't accidentally pick the just-created empty session as "latest").

Verified end-to-end via Chrome MCP smoke on `feat/welcome-cards-polish`:
all four cards rendered with real `feat/welcome-cards-polish` branch
substituted in, dirty test/source files auto-filled, and clicking a
card prefilled the composer with the resolved prompt.

Tested: bun run lint, EmptySession.test.tsx (22/22 pass), Chrome MCP
smoke against running server + vite.
Confidence: high
Scope-risk: narrow

Co-authored-by: 你的姓名 <you@example.com>
2026-06-10 19:16:05 +08:00
小橙子
a7a48d15f3
feat: zero-prompt session handoff via two-layer LLM summary (#2)
* feat(server): two-layer session summary service for cross-session handoff

Backs the desktop "Continue from here" auto-handoff path. When a user
returns to a project to start a new chat session, the next-session AI
needs to know what the previous session was about; the existing
zero-token textarea-prefill carries 80 tokens of metadata that's good
for a human to read but useless for an AI ("title" alone tells it nothing
about the actual conversation). This service generates a compact
two-layer summary (project-level + recent-detailed) of the previous
session, caches it on disk, and lets the WS handler stage it as
`--append-system-prompt` on the next CLI launch so the receiving AI
starts with context but without paying for the whole prior transcript.

Token economics:
  - One LLM call per session, capped at 1500 output tokens
    (CLAUDE_CODE_HANDOFF_MAX_TOKENS, env-overridable).
  - Input transcript is tail-sliced at 80k chars
    (CLAUDE_CODE_HANDOFF_INPUT_CHARS) — recent context dominates the
    hand-off, older turns drop first; tool uses are condensed to
    one-liners so the input budget isn't spent on tool mechanics.
  - Cached on disk as `<sessionId>.summary.json` next to the JSONL.
    First call after a session has new turns (baseMessageCount drift)
    triggers regeneration; otherwise it's an instant disk hit.
  - Generation reuses the user's active provider via the same
    resolution pattern as titleService.ts: prefer the configured
    haiku/cheap model on the active provider, fall back to its main
    model. OpenAI Codex (ChatGPT OAuth) path uses the existing
    anthropicToOpenaiResponses + streaming response path.
  - All failures degrade silently to `null` so the welcome-screen
    falls back to the existing zero-token textarea prefill.

Wiring:
  - sessionSummaryService.ts (NEW) exports getSessionSummary,
    invalidateSessionSummary, and formatHandoffSystemPrompt for the
    WS handler.
  - sessions API gains GET/POST/DELETE /api/sessions/:id/summary.
  - conversationService SessionStartOptions gets handoffSystemPrompt;
    getRuntimeArgs appends it via --append-system-prompt independently
    of orchestration so both can be active at once.
  - WS handler adds `set_handoff_summary` ClientMessage type, an
    in-memory handoffSummarySessions map (one-shot consumed at next
    CLI start), getRuntimeSettings reads-and-deletes the staged
    summary so unrelated restarts don't re-attach a stale prompt,
    and session-end cleanup drops it.

Tested live via Chrome DevTools MCP (commit 2 covers the frontend
side):
  - POST /api/sessions/<id>/summary on a 9-message transcript:
    489 input + 335 output tokens, mimo-v2.5 produced clean two-layer
    JSON, written to <sessionId>.summary.json (~3KB).
  - Subsequent GET hits cache in ~40ms.
  - End-to-end: clicked "Continue from here" → server resolved cached
    summary → WS staged → CLI relaunched with --append-system-prompt
    carrying the formatted hand-off → mimo received it and listed
    exactly the modified files in the previous session it never saw,
    correctly identifying `feat/session-handoff-summary` branch state.

Confidence: high
Scope-risk: moderate (new endpoint, new service, new WS message type,
new SessionStartOptions field; all read-only / append-only; failures
fall back to existing zero-token path).
Tested: live MCP smoke + targeted curl on this branch.
Not-tested: server-side unit tests for sessionSummaryService (planned
follow-up — covers JSONL parsing, transcript tail-slice, summary
parsing, cache staleness detection, formatHandoffSystemPrompt
formatting). Live integration validated the happy path end-to-end.

* feat(desktop): auto-handoff flow on "Continue from here"

When a user clicks the Recent activity panel's "Continue from here"
button, the desktop now resolves the previous session's two-layer
summary server-side, stages it as the next session's system prompt
addendum (via WS set_handoff_summary), and auto-sends a short trigger
message so the receiving AI starts with full context — no textarea
preview, no manual confirm. Falls back to the original zero-token
textarea-prefill path on any failure.

User flow:
  1. User clicks "从这里继续" → button shows spinner + label
     "正在准备上下文..." while the summary resolves.
  2. projectsApi.resolveSessionSummaryForHandoff tries cache first
     (instant) and only falls back to LLM generation on miss.
  3. EmptySession path: createSession → connectToSession → WS
     set_handoff_summary → sendMessage(continueTriggerMessage). Server
     stages the summary in handoffSummarySessions and the next CLI
     launch picks it up via --append-system-prompt.
  4. ActiveSession's empty welcome state: same WS staging + auto-send
     against the live session — server schedules a runtime restart so
     the CLI relaunches with the hand-off context.
  5. On any error (provider down / generation failed / network), card
     falls back to dispatching the existing composer-prefill event with
     the static hand-off paragraph; user can edit and send manually.

Files:
  - desktop/src/api/projects.ts: SessionSummary type + getSessionSummary
    + generateSessionSummary (90s timeout — first generation can take
    30-60s on long transcripts) + resolveSessionSummaryForHandoff helper
    that prefers cache.
  - desktop/src/components/welcome/RecentActivityCard.tsx: replaced
    onApplyHandoff (sync, just writes to textarea) with onAutoHandoff
    (async, host owns the full flow). Button shows progress_activity
    spinner + i18n'd "preparing context..." label while the host's
    promise is in flight; disabled during pending.
  - desktop/src/pages/EmptySession.tsx, ActiveSession.tsx: implement
    the host-side handoff flow — resolve summary, then either create
    new session + WS stage + auto-send, or fall back.
  - desktop/src/types/chat.ts: ClientMessage union gains
    set_handoff_summary.
  - i18n: 2 new keys per locale (handoffGenerating, continueTriggerMessage)
    across en/zh/zh-TW/jp/kr.

Tested live (companion to ce53f503 server-side commit):
  - Cache-hit path: card click → spinner ~50ms → new session → CLI
    launched with --append-system-prompt → mimo-v2.5-pro received the
    formatted hand-off and listed exactly the previously-modified files
    on feat/session-handoff-summary, identified the feature state
    correctly, and asked to verify TS — proving it had ground-truth
    context from the previous session it never directly saw.
  - Cache-miss path: ~10-15s LLM generation, button stays in loading
    state, then proceeds.
  - Failure path: simulated by manually breaking the provider URL —
    button completes silently, textarea gets the fallback static
    paragraph (existing zero-token path).
  - Screenshot: artifacts/desktop-handoff-auto-resume.png

bun run lint passes (tsc --noEmit clean).

Confidence: high
Scope-risk: moderate (new auto-send-on-click behavior; mitigated by
silent fallback to existing textarea path and a generous timeout).
Tested: live MCP smoke for cache-hit happy path + manual failure path.
Not-tested: vitest unit tests for the new RecentActivityCard handoff
states (planned follow-up).

* fix(handoff): cache-only WS handler so summary generation can't double-run

Acted on a code review against the auto-handoff feature. Two findings,
one real one false-positive (left a defensive note for the false one
to keep future maintainers from "fixing" what isn't broken):

REAL ISSUE: handleSetHandoffSummary used getSessionSummary(), which
falls back to LLM generation on cache miss. The frontend's "Continue
from here" path always calls POST /api/sessions/:id/summary first
(which performs any needed LLM call), and only dispatches the WS
set_handoff_summary message AFTER the HTTP returned a successful
summary. So the WS handler should always find the cached summary on
disk. If it somehow doesn't, the OLD behavior would silently re-invoke
the LLM — blocking the WS handler for up to 60s and double-charging
the user. New behavior: read cache only, fail-fast with a clear
warning log if the cache is missing, and let the new session start
without hand-off context (the trigger message reads as a normal
"continue" prompt). Better than hanging or surprise-billing.

  - sessionSummaryService.ts: new exported getCachedSessionSummary
    helper — pure disk read, never calls the LLM.
  - ws/handler.ts: handleSetHandoffSummary swapped over with a
    detailed comment explaining the contract.

FALSE POSITIVE (kept as inline doc): the review claimed
EmptySession's flow has a race because wsManager.send happens before
the WS open. Actually wsManager.send tolerates a not-yet-OPEN socket
by queueing into pendingMessages and flushing on ws.onopen (see
desktop/src/api/websocket.ts). The current code is correct. Added a
defense-in-depth comment at the call site so a future "fix" doesn't
introduce an isConnected() gate that would break the contract.

Confidence: high
Scope-risk: narrow (one WS handler swap + one new export, no
behavioral change on the happy path; only the LLM-double-call edge
case is now fail-fast).
Tested: bun run lint passes (tsc --noEmit clean). Live MCP smoke
already covered the happy path on the previous commit; this commit
narrows a failure mode the smoke didn't actively exercise.

* feat(desktop): handoff progress stages + preview panel

Two UX additions on top of the auto-handoff flow:

1. Stage-aware progress label inside the "Continue from here" button.
   Replaces the opaque "Preparing context..." spinner with a localized
   walk-through:
     preparing            (default while host kicks off)
     reading-cache        (during the GET /api/sessions/:id/summary)
     generating-summary   (during the POST → LLM round-trip on cache miss)
     starting-session     (during createSession + WS staging + auto-send)
   The card stays dumb — host calls a setStage(...) callback handed to
   it as the third arg of onAutoHandoff. Card maps the stage to the
   corresponding i18n key; defaults to "preparing" when the host hasn't
   advanced. Falls back gracefully if a host doesn't call setStage at
   all (the spinner just shows "preparing" the whole time).

2. "Preview summary" toggle next to "Continue from here". Click expands
   a 280px-max scrollable inline panel showing the cached summary's
   main + recent layers, plus a one-liner with the model used and token
   counts. If no summary is cached yet (first time the user is about to
   click Continue), the panel shows "尚未生成摘要…" so the user knows
   what to expect.

   Card fetches the cached summary on mount via projectsApi
   .getSessionSummary (cache-only GET, never triggers an LLM call). The
   fetch is opportunistic: if it fails or returns null, the preview
   panel just shows the "no preview yet" hint.

i18n: 9 new keys per locale across en/zh/zh-TW/jp/kr — three stage
labels (handoffStage.readingCache, generatingSummary, startingSession),
preview toggle/close labels, three preview-panel section labels, and
the "no preview yet" placeholder.

Tested live in browser via Chrome DevTools MCP:
  - Empty cache state: preview button visible, panel shows "no preview
    yet" placeholder.
  - Primed cache (mimo-v2.5 summary, 17.5s LLM call): reload → preview
    panel renders with full main + recent content (292px high), token
    metadata line at the bottom.
  - Stage progression visible to the eye on the next "Continue from
    here" click during the cache-miss path; cache-hit path skips the
    "generating-summary" stage and goes straight to "starting-session".
  - Screenshot: artifacts/desktop-handoff-preview-panel.png

bun run lint passes (tsc --noEmit clean).

Confidence: high
Scope-risk: narrow (UI-only additions; existing onAutoHandoff signature
gained a third arg, both hosts updated; default value pathway preserved
when callers don't use setStage).

* feat(desktop): handoff chip in chat header + tests

The previous commits in this PR plumbed in the auto-handoff machinery
(server summary service, frontend resolution + auto-send, progress
stages, preview panel). What's missing is a *persistent visual signal*
in the active chat that the AI was bootstrapped with prior context.
Without it, a returning user sees an ordinary chat with no hint that
they pre-loaded a summary, and may distrust or forget that context is
present.

This adds a small chip in the ActiveSession chat header next to the
existing token / message-count chips:

    ↗ 接续上次 (639 t)

with a hover tooltip:

    接续自"<previous session title>"。AI 启动时已带上了 639 tokens
    的上次会话摘要作为系统提示。

The chip is purely informational — clicking does nothing. Token count
is a frontend-side estimate (chars / 4) of the staged summary's
main+recent text size, so it tells the user "roughly how big is the
hand-off addendum in my system prompt right now". The exact tokens are
also surfaced server-side in ContextUsageIndicator's "System prompt"
category, so this chip doesn't double-count.

Plumbing:
  - sessionRuntimeStore gains a new handoffInfo: Record<sessionId,
    SessionHandoffInfo> field, persisted to localStorage under
    cc-haha-session-handoff. Mirrors the existing coordinatorModes
    pattern: setHandoffInfo / clearHandoffInfo actions, automatic
    cleanup in clearSelection and migration in moveSelection so the
    drafts → real-session-id transition during session creation
    doesn't lose the handoff record.
  - SessionHandoffInfo carries previousSessionId, snapshotted previous
    title, approxTokens, and generatedAt for staleness display.
  - RecentActivityCard's onAutoHandoff signature gains a third arg
    previousSessionTitle so the host can stash it without re-fetching
    the recent-activity payload.
  - EmptySession + ActiveSession hosts call setHandoffInfo on the
    target session right before WS staging, so the chip lights up the
    moment the new session opens (or the live session is restarted).
  - i18n: 2 new keys per locale (session.handoffChip,
    session.handoffChipTooltip).

Tests: new desktop/src/stores/sessionRuntimeStore.test.ts with 6
focused tests:
  - setHandoffInfo persists to store + localStorage
  - clearHandoffInfo removes selectively
  - clearHandoffInfo on missing key is a no-op
  - clearSelection cleans up handoffInfo for the same key
  - moveSelection migrates handoffInfo from draft → real key
  - handoffInfo and coordinatorModes are independent

Verified live in the browser via Chrome DevTools MCP:
  - Click "Continue from here" → cache-hit summary path → chip appears
    in the new session's header showing exact previous title and
    639-token approx count.
  - localStorage persists across reload (chip survives F5).
  - Multiple sessions: the chip is per-session, cleared on session
    close (via clearSelection wiring).
  - Screenshot: artifacts/desktop-handoff-chip-in-header.png

bun run lint passes (tsc --noEmit clean).
bun run vitest passes for chatStore + EmptySession (124/124) +
sessionRuntimeStore (6/6 new).

Confidence: high
Scope-risk: narrow (one new persisted store field, additive UI in the
header, no behavioral change on the handoff path itself).

---------

Co-authored-by: 你的姓名 <you@example.com>
2026-06-10 18:35:02 +08:00
程序员阿江(Relakkes)
0d45439fbe feat(trace): add session trace monitoring (#606)
Tested:
- cd desktop && bun run test -- --run src/pages/TraceList.test.tsx
- bun test src/server/__tests__/trace-capture.test.ts
- bun run check:desktop
- bun run check:server

Scope-risk: broad
2026-06-10 17:13:27 +08:00
程序员阿江(Relakkes)
56661c7967 fix: support desktop output style config (#652)
Fixes #652.

Add desktop/server output style settings APIs and a General Settings picker
that mirrors the Claude CLI outputStyle sources.

Persist active-project choices to .claude/settings.local.json and global
choices to user settings, and route /config to the local settings UI.

Tested:
- bun test src/server/__tests__/settings.test.ts
- cd desktop && bun run test -- src/components/chat/composerUtils.test.ts src/stores/settingsStoreOutputStyle.test.ts src/pages/SettingsOutputStyle.test.tsx
- cd desktop && bun run test -- src/__tests__/generalSettings.test.tsx
- bun run check:desktop
- bun run check:server
- bun run verify

Confidence: high
Scope-risk: moderate
2026-06-10 15:24:37 +08:00
你的姓名
f41d88673d release: v0.5.7 2026-06-10 07:12:53 +08:00
你的姓名
ae533041de fix(thinking): default any third-party Anthropic-compatible proxy to thinking-capable
Two fixes that together unblock visible thinking on third-party Anthropic
proxies (mimo, lgfzer, kiro-account-manager loopback, etc.):

1. src/utils/thinking.ts: when the API provider is firstParty but the base
   URL is not Anthropic's own, treat the gateway as thinking-capable rather
   than relying on a per-host whitelist. Drops the dead
   isMiniMaxAnthropicEndpoint helper. SUPPORTED_CAPABILITIES=none still
   opts out for any endpoint that genuinely cannot return reasoning.

2. desktop ThinkingBlock now defaults to expanded so reasoning is visible
   the moment thinking deltas start arriving, matching the streaming
   experience users expect from reasoning models. Click still toggles.

Verified end-to-end against mimo-v2.5-pro through the desktop UI: thinking
content streams in 40-110 char increments over a few seconds with the
"思考中" cursor visible, then collapses to "已思考" on completion.

Tested:
- src/utils/__tests__/thinking.test.ts (113 utils tests pass)
- desktop ThinkingBlock.test.tsx + chatBlocks.test.tsx (16 tests pass)
- Live mimo SSE probe confirms many small thinking_delta events upstream
- Browser MCP smoke: complex prompts show real-time thinking growth

Confidence: high
Scope-risk: narrow
2026-06-10 07:06:16 +08:00
小橙子
1401107509
feat: harden agent routing, ship welcome surfaces, and unblock browser MCP smoke (#1)
* feat(agent-tool): rewrite contradictory review verdicts at the harness layer

The code-reviewer and security-reviewer subagents end with a parseable
verdict (`REVIEW: APPROVE` / `SECURITY: PASS`) and their prompts forbid
emitting the positive form alongside `[CRITICAL]` or `[HIGH]` findings.
That rule is enforced only in prose, so when the model slips, the
parent agent reads "APPROVE" and ships.

Add a pure post-processing pass over the subagent's final text in
`finalizeAgentTool`: detect the verdict line, look for severity tags
above it, rewrite to `CHANGES_NEEDED` when they conflict, and prepend
a one-line audit notice so the parent (and a human reading the
transcript) can see the harness intervened. Logs a
`tengu_agent_sentinel_mismatch` analytics event for fleet visibility.

Debugger's `ROOT CAUSE: FOUND` is intentionally out of scope — its
"evidence from a real command" rule isn't mechanically verifiable.

Tested: bun test src/tools/AgentTool/sentinelCheck.test.ts (11 cases)
Confidence: high
Scope-risk: narrow

* feat(agent-tool): show subagent_type in tool-use transcript line

Each `Agent` tool call previously rendered as just its description in
the transcript, which made it invisible whether the main agent picked
a specialist or fell back to the default. Diagnosing "why didn't it
use code-reviewer?" required reading analytics or guessing.

Append `→ <subagent_type>` to the rendered description when the type
is explicit. Omitted (general-purpose default or fork) keeps the old
format — absence of the marker is itself the signal.

Tested: visual smoke only — pure UI render change with no behaviour
Confidence: high
Scope-risk: narrow

* feat(agent-tool): nudge verification subagent after edits accumulate

Verification is the only built-in agent that's *required* by the
session-specific guidance for non-trivial implementations, but the
trigger condition ("3+ file edits / backend / infra change") is
evaluated by the main agent from its own work — a classic LLM
self-audit failure mode. The model can convince itself that 4 edits
"are mostly one core change with small adjustments" and skip
verification entirely.

Add a `verification_gate_reminder` attachment that fires once when
the main thread accumulates `editCount >= threshold` (default 3)
file-mutating tool uses (Edit / Write / NotebookEdit) without
invoking the verification subagent. The reminder renders as a
`<system-reminder>` block on the next turn telling the model to
spawn `subagent_type="verification"` before reporting completion,
along with a permission to skip if the change is genuinely trivial.

A verification subagent invocation resets the counter; a previously
fired reminder suppresses re-firing until the next reset, so the
nudge isn't spammed every turn.

Subagents don't see the reminder (it's main-thread only). The gate
also no-ops when the verification agent isn't loaded in the current
session (e.g. SDK build with built-ins disabled). Disabled by
`CLAUDE_CODE_VERIFICATION_GATE_OFF=1`; threshold tuned via
`CLAUDE_CODE_VERIFICATION_GATE_THRESHOLD=N`.

Tested: bun test src/utils/verificationGate.test.ts (10 cases)
Confidence: high
Scope-risk: moderate

* feat(agent-tool): tighten subagent routing and ship real coordinator mode

The flat 15-agent routing has three known weak spots that this commit
addresses together because they share AgentTool.tsx call-path edits:

1. **General-purpose default is a vacuum cleaner.** Omitting
   `subagent_type` falls back to general-purpose, which has full tool
   access and cheerfully takes on review/security/perf/migration work
   that a specialist would have done better and cheaper. Add a regex
   heuristic over the prompt: when it strongly signals a specialist
   (e.g. "code review", "security audit", "root cause", "migrate from
   X to Y") and that specialist is available, refuse the default and
   throw an error pointing the model at the right type. Disabled with
   `CLAUDE_CODE_GP_DEFAULT_STRICT=0`. Skipped in fork and coordinator
   modes (they have their own routing).

2. **No bound on verifier-style loops.** The verification prompt says
   "On FAIL: fix, resume, repeat until PASS" — if the verifier
   incorrectly keeps flagging the same change, the loop only stops
   when the token budget runs out. Add a per-session counter for each
   built-in agent type. Cap defaults are 5 for verification (loop
   target) and 8 for everything else. When exceeded, throw with a
   message asking the model to consult the user. Caps tunable per
   type via `CLAUDE_CODE_AGENT_LIMIT_<TYPE>=N`; off via
   `CLAUDE_CODE_AGENT_LIMITER_OFF=1`. Counts main-thread invocations
   only.

3. **Coordinator mode was a stub.** The `coordinator/workerAgent.ts`
   that `getBuiltInAgents()` requires when `COORDINATOR_MODE` is on
   was an auto-generated Proxy returning nothing useful. Replace with
   a real registry: a new `WORKER_AGENT` (full-tool-access generic
   delegate matching the existing coordinator prompt's
   `subagent_type: "worker"` references) plus the 11 specialists
   minus general-purpose, claude-code-guide, statusline-setup. Also
   harden the coordinator's own prompt with a "you orchestrate, you
   do not implement" rule (no direct Edit/Write/NotebookEdit) and a
   specialist roster table so the coordinator picks the right
   delegate. Improve the AgentTool error message when coordinator
   mode receives a request without `subagent_type`.

Constraint: cannot restrict the coordinator's own tool pool at the
tool layer without invasive changes to main-thread tool assembly —
enforcement is prompt-level, consistent with the rest of the
architecture.

Rejected: a tighter heuristic that requires a domain-specific phrase
(e.g. "review THIS PR") was considered for #1 but rejected — too
narrow to catch the routing problems we actually see in fleet logs.

Tested: bun test src/tools/AgentTool/specialistRouter.test.ts
        bun test src/tools/AgentTool/invocationLimiter.test.ts
        bun test src/coordinator/workerAgent.test.ts
        (29 cases total, all pass)
Confidence: medium
Scope-risk: moderate
Directive: if specialistRouter false-positives in production, prefer
  tightening the regexes over removing the gate — the failure mode
  it catches (general-purpose taking specialist work) is worse than
  the failure mode of an extra retry.

* docs(desktop): add local Chrome DevTools MCP testing guide

Document the workflow for driving the desktop UI end-to-end via
Chrome DevTools MCP without going through Electron. The pure-browser
dev path normally fails — CSP blocks non-loopback origins, the
H5-disabled CORS gate blocks cross-origin browsers, and the desktop's
loopback exemption clears the H5 token so even the loopback path
401s. The stable bypass is a Vite proxy that same-origins API
traffic to the renderer.

This commit covers the docs side only:
- `docs/desktop/10-local-mcp-testing.md`: full how-to with Quick
  Start, MCP control snippets, can/can't test matrix, the exact
  reproduction recipes for each `feat/agent-routing-hardening`
  change (sentinel, gp-strict, limiter, verification gate, routing
  observability), a Windows-specific `electron:dev` workaround, and
  a fault dictionary mapping the symptoms each blocker produces back
  to its cause.
- `docs/desktop/index.md`: link the new section in the docs index.
- `.kiro/steering/desktop-mcp-testing.md`: concise rule file with
  `inclusion: manual` so it loads via #-mention only when relevant,
  not as default-included context.

The Vite proxy itself is intentionally NOT committed here. The docs
explain it as a dev-only change the developer adds locally and
either reverts or commits to a separate `feat/desktop-browser-dev-proxy`
branch — the production renderer relies on Electron IPC for the
server URL and must not see the dev proxy in main.

Tested: docs render via VitePress local preview; steering file
loads via Kiro `#desktop-mcp-testing` mention; instructions
verified against the live setup that produced
artifacts/desktop-bootstrapped.png and artifacts/desktop-agents-page.png.
Confidence: high
Scope-risk: narrow

* feat(desktop): unblock pure-browser dev MCP smoke with proxy + forceH5

The standalone-browser dev workflow (vite + Chrome DevTools MCP)
fails three ways without these escape hatches:

1. CSP `connect-src` only allows loopback origins, so the renderer
   can't talk to the API server cross-origin even on a LAN.
2. With H5 access disabled the server returns 403 to all
   cross-origin browser requests.
3. With H5 access enabled the desktop's loopback exemption
   (`requiresH5AuthForServerUrl` returns false for `localhost:1420`)
   calls `setAuthToken(null)`, so `buildSessionWebSocketUrl` omits
   `?token=`, and the session WS handshake 401s into a 1006 close
   reconnect loop. The UI shows "处理中..." but the message never
   reaches the server — `wsManager.connections` stays `{}`, server
   stdout is silent, and no `POST /api/sessions/.../message` ever
   fires.

Two dev-only changes break that deadlock without affecting
production:

- `desktop/vite.config.ts`: add `server.proxy` for `/health`,
  `/api`, `/ws`, `/local-file`, `/preview-fs` to `127.0.0.1:3456`.
  Same-origins all backend traffic so CSP and CORS pass. Vite dev
  only — Electron production path is unaffected because Electron
  resolves the server URL via IPC at runtime.

- `desktop/src/lib/desktopRuntime.ts`: add a `?forceH5=1` query
  param check in `initializeBrowserServerUrl`. When present, treat
  the URL as needing H5 auth even if it's loopback, so the H5 token
  survives bootstrap and gets attached to the session WS URL. In
  production the param is never set, so behavior is identical.

Workflow with both in place:

  http://localhost:1420/?serverUrl=http%3A%2F%2Flocalhost%3A1420
  &forceH5=1&h5Token=<TOKEN>

Verified end-to-end via WS sniffer + server log:
- ws://localhost:1420/ws/<sid>?token=... opens
- recv {"type":"connected","sessionId":"..."}
- server logs "[WS] Client connected" + "Starting CLI for ..."
- Sending a "use Task with subagent_type=Explore" prompt produced
  a `tool_use_complete` for the Agent tool with
  `input.subagent_type="Explore"` exactly as routed.

Also rewrites `docs/desktop/10-local-mcp-testing.md` to record:
- both patches as required steps (was missing the forceH5 piece)
- exact PowerShell commands to mint a one-shot H5 token
- the WebSocket sniffer script and three-signal diagnosis recipe
  for "is the backend actually processing"
- the Item 5 (routing observability) gap: the change in
  `src/tools/AgentTool/UI.tsx` only affects CLI/Ink rendering;
  the desktop's separate `ToolCallGroup.tsx` doesn't surface
  `subagent_type` and would need its own follow-up patch.

Tested: ran `bun run src/server/index.ts` + `bun run dev` in the
desktop/, opened the documented URL via Chrome DevTools MCP,
sent a Task-with-explicit-subagent_type chat message, observed
the full WS frame sequence and server log lines listed above.
Confidence: high
Scope-risk: narrow
Directive: do not generalize the forceH5 escape hatch to a
  permanent feature — when it's needed, the dev-mode
  loopback-clears-token branch is doing the right thing for
  Electron and should not be removed.

* feat(desktop): show subagent_type badge next to Agent tool header

The CLI/Ink change in src/tools/AgentTool/UI.tsx surfaces routing
in the terminal transcript as `<description> → <subagent_type>`,
but the desktop renderer is independent (`ToolCallGroup.tsx` /
`AgentCallCard`) and was dropping that signal entirely. Diagnosing
"why didn't it use code-reviewer?" required reading analytics or
guessing.

Add a small `→ <subagent_type>` chip next to the Agent header in
AgentCallCard. The badge mirrors the CLI behavior:
- Renders only when `input.subagent_type` is a non-empty string;
  general-purpose / fork / omitted defaults stay quiet.
- Carries `title="subagent_type: <name>"` for hover detail.
- Uses the same border-pill treatment as other secondary
  transcript chips so it doesn't compete with the description.

Verified end-to-end via the local Chrome DevTools MCP path: a
real chat with `subagent_type=Explore` renders `Agent → Explore
查找specialistRouter相关文件` in the live transcript, screenshot
saved to artifacts/desktop-routing-arrow-highlighted.png.

Tested: cd desktop && bun run test -- --run -t "subagent_type"
        2 tests pass:
        - renders subagent_type as → Explore badge next to Agent header
        - omits subagent_type badge when input has no subagent_type
Confidence: high
Scope-risk: narrow

* feat(desktop): MCP tool visibility, per-tool toggle, marketplace install

Surface MCP capabilities so the agent never under-uses a configured server,
and add a one-click install path for new servers.

Tool visibility & control:
- GET /api/mcp/:name/tools returns the live tool catalog (handles disabled /
  needs-auth / host-preflight-failed / connected states)
- POST /api/mcp/:name/tools/:tool/toggle hides a single tool from the agent,
  persisted to ~/.claude.json disabledMcpTools (user scope, all projects);
  fetchToolsForClient filters on it and the cache is invalidated on toggle
- Desktop MCP detail page gains an Overview/Tools tab strip (both edit and
  read-only views) listing name, description, annotations, JSON schema, plus
  an inline enable/disable switch with a "hidden from agent" marker

Marketplace:
- Curated builtin catalog (14 servers across 8 categories) plus opt-in remote
  catalog sources cached under ~/.claude/cc-haha/mcp-marketplace.json
- Endpoints: list / refresh / add-source / remove-source
- Desktop marketplace page with category grouping, install dialog (with
  required-env prompts), and source management

Plugins:
- Plugin list rows get inline enable/disable + uninstall, matching the MCP
  settings UX, instead of requiring users to open the detail panel
- Extract shared ToggleSwitch component

Tested: bun test src/server/__tests__/mcp.test.ts (tools list + per-tool
toggle + marketplace), cd desktop && bun run lint, mcpSettings vitest, and a
browser smoke pass via the local Chrome DevTools MCP flow (tool toggle
persistence, marketplace render, plugin enable/disable + uninstall dialog).
Not-tested: full bun run verify gate
Scope-risk: moderate

* fix(agent-tool): rephrase gp-strict error to stop poisoning sessions

Live testing surfaced an unintended side effect of the gp-strict
redirect message. The original wording embedded attribute-style
fragments like `subagent_type="code-reviewer"`. When that error
came back to the model as a tool result, the model copied the
shape into a TEXTUAL `<tool_use name="Agent">{...}` block instead
of issuing a real tool call — and every subsequent Agent call in
that session degraded to text. Confirmed by inspecting the
session JSONL via /api/sessions: messages 4..N in the affected
session were `assistant: text(...)` carrying simulated tool_use
XML, not `tool_use` content blocks. Reproduced with two unrelated
models (claude-haiku-4.5 and mimo-v2.5-pro) on the same poisoned
session, and crucially DID NOT reproduce in a fresh session with
either model — so the failure is the message text inducing
text-form mimicry, not the model itself.

Fix:
- Extract the redirect text into `formatSpecialistRedirectMessage`
  in specialistRouter.ts so it has its own test surface.
- Rephrase as prose that names the parameter and value without
  any `key="value"` assignment-and-quotes form. New shape:

    "This task looks like a job for the code-reviewer specialist
     rather than the general-purpose default. Re-call Agent with
     the subagent_type parameter set to code-reviewer. ..."

- Wire AgentTool.tsx through the helper so the throw site is one
  line.

Add a regression assertion to specialistRouter.test.ts that the
formatted message contains NO `key="value"` shapes, no XML-ish
open tags, and no `"subagent_type":` JSON fragments — across all
ten specialist names. This is the property the live failure
violated; lock it.

The invocation-limiter message uses shell-style `KEY=N` syntax
(not `key="value"`), did not exhibit the same behavior in
testing, and is left untouched.

Tested: bun test src/tools/AgentTool/specialistRouter.test.ts
        17 pass (13 existing + 4 new), 61 expect() calls
Confidence: high
Scope-risk: narrow
Directive: any future change to redirect/error text returned
  to the model should preserve the no-key="value" property.

* feat(orchestration): require copying project tool rules into sub-agent prompts

Live test on the desktop `+` orchestration mode showed mimo-v2.5-pro fan
out three sub-agents in parallel for a multi-task PR review prompt — but
none of the dispatched prompts mentioned the project's mandated codegraph
MCP tool. The sub-agents fell back to plain Bash + git diff + grep, so the
project's preferred code-exploration tooling went unused.

Root cause: the orchestrator's `Propagate project conventions` rule was a
soft "for example, restate codegraph rules" bullet. The model read it as
optional. Built-in specialist system prompts (code-reviewer, security-
reviewer, debugger, etc.) explicitly enumerate Bash/grep/git as inspection
tools and never mention codegraph; CLAUDE.md is loaded but the system
prompt outweighs it. The orchestrator was the only place the gap could be
closed, and it skipped it.

This commit upgrades the rule from optional to REQUIRED in the
orchestration system prompt, names the concrete check ("scan CLAUDE.md,
AGENTS.md, .claude/rules"), and tells the orchestrator to copy matching
rules verbatim. Read-only research agents (Explore, Plan) are called out
explicitly because they run with omitClaudeMd:true and absolutely depend
on the orchestrator to forward project tool conventions.

A new ORCHESTRATION_PROPAGATE_RULES_MARKER export plus a regression test
in conversations.test.ts lock the imperative phrasing so it can't be
quietly downgraded back to a soft suggestion.

Wiring is unchanged: built-in code-reviewer/security-reviewer/commit-pr
already inherit CLAUDE.md and have MCP tools available. This is purely a
prompt-strength change, only active when the user enables the desktop
`+` orchestration toggle.

Tested: bun test src/server/__tests__/conversations.test.ts (new test
plus the two existing orchestration-prompt tests pass).
Confidence: high
Scope-risk: narrow

* fix(verification-gate): preserve reminder flag across verification reset

Caught by an orchestrator-mode code-reviewer dispatch on this branch.

When a verification subagent invocation reset the edit counter, the early
return in getVerificationGateState() hardcoded reminderAlreadyFired:false,
discarding any reminderFired flag we'd accumulated walking back from the
tail. Concretely: if a verification_gate_reminder fired AFTER the verify
(i.e. newer than verify in the transcript), the reverse walk would set
reminderFired=true on its way to the verify boundary, then throw it away
on the early return. The gate would then inject a fresh reminder every
turn while the post-verify edit count stayed above the threshold,
ballooning the prompt.

Fix: carry the accumulated reminderFired through the early return. The
reminder we saw is necessarily newer than this verify (we're walking
backwards from the tail and only got here after seeing it), so it's the
valid signal that the model was already nudged for the current edit batch.

Tests: extend verificationGate.test.ts with a "preserves a reminder fired
AFTER a verification reset (no spam loop)" case that fails on the old
hardcoded false and passes with the fix. The pre-existing "ignores a
reminder that fired BEFORE a verification reset" test still passes because
in that ordering the reminder is older than the verify, so reminderFired
stays false on the way back.

Tested: bun test src/utils/verificationGate.test.ts (11/11 pass).
Confidence: high
Scope-risk: narrow

* chat(i18n): rename `协调模式` to `编排模式` (zh / zh-TW)

Brings the simplified- and traditional-Chinese label for the desktop `+`
menu's orchestration toggle in line with the English (Orchestration mode),
Japanese (オーケストレーションモード), and Korean (오케스트레이션 모드)
labels — all five locales now use the same `orchestration` word family.
`协调` was accurate but vague; `编排` matches the term Chinese
developers already see in LangChain, CrewAI, AutoGen, and Anthropic docs,
and tells a user reading just the label what the toggle actually does
(plan → fan out → synthesize) rather than a generic `coordinate`.

UI-text only. No behavior, store, WS message, or i18n-key changes.

* feat(desktop): welcome-screen task cards with one-click orchestration

Adds four quick-start task cards to the new-session welcome screens — both
EmptySession (the no-session-yet entry) and ActiveSession's empty state
(when a fresh session tab has no messages yet). Each card pre-fills the
composer with a starter prompt for a common workflow:

  - Pre-merge code review        (auto-enables Orchestration mode)
  - Investigate a failing test   (auto-enables Orchestration mode)
  - Write unit tests             (no orchestration)
  - Explain unfamiliar code      (no orchestration)

Cards flagged `orchestrate: true` automatically turn on Orchestration
mode for the session that gets created (or for the live session in
ActiveSession's empty state), so a new user sees fan-out behavior on
first contact instead of having to discover the `+` menu first. This is
the intended discovery surface for the orchestration directive shipped on
this branch.

Plumbing:
  - WelcomeTaskCards.tsx is a shared component used by both welcome
    surfaces, so card list, copy, and styles stay in one place.
  - EmptySession owns its own composer textarea, so card click flips
    component state directly (setInput + draftOrchestrate flag) and the
    flag is applied via setSessionCoordinatorMode after createSession()
    resolves and before connectToSession() — chatStore replays the
    persisted toggle on connect, so the CLI launches with the right
    --append-system-prompt.
  - ActiveSession's session is already live, so a card click dispatches a
    `cc-haha:composer-prefill` window event that ChatInput listens for
    and applies to its composer state when the sessionId matches the
    active tab. Orchestration cards call setSessionCoordinatorMode
    directly, which sends `set_coordinator_mode` over WS and triggers a
    CLI restart with the orchestration directive — verified live: the
    server logs `[WS] Restarted CLI for <id> with runtime override`
    immediately after the click.
  - Cards never disable an existing Orchestration toggle. Non-orchestration
    cards simply leave the toggle alone, so users who pre-enabled
    Orchestration via the `+` menu don't have it silently turned off.
  - Hidden on phone-sized H5 viewports (composer is already dense there).

i18n: ten new keys under `empty.tasks.*` with localizations for en,
zh, zh-TW, jp, kr (the same five locales the orchestration label
already covers). The orchestration hint reads `将通过编排模式分派给
专家子代理处理` in zh — aligning with the
`协调模式`→`编排模式` rename shipped in the previous commit.

Tests: five new tests in EmptySession.test.tsx covering: cards render on
desktop, hidden on mobile, click pre-fills the composer, orchestration
cards persist coordinator mode after submit, non-orchestration cards
leave coordinator mode untouched. `bun run vitest` 22/22 pass.
`bun run lint` (tsc --noEmit) clean.

Tested live in browser: clicked the Pre-merge review card → composer
filled with the 78-char zh prompt → localStorage carries
`coordinator-modes[<sessionId>]: true` → server log shows the CLI
restarted with `--append-system-prompt`. Screenshot:
artifacts/desktop-welcome-task-cards.png.

Confidence: high
Scope-risk: narrow

* chore(scripts): one-click MCP smoke setup for Windows

Wraps the recurring 4-step setup from docs/desktop/10-local-mcp-testing.md
into a single idempotent script. Same workflow we'd otherwise type by hand
every time we want to drive the desktop UI from Chrome DevTools MCP.

scripts/dev-mcp-test.ps1 does:
  1. Health-check the API server on :3456 — fail fast with the start
     command if it's down.
  2. Health-check Vite on :1420 — fail fast with the start command if
     it's down.
  3. Confirm Vite's /health proxy is wired through to the server (catches
     stale vite.config.ts).
  4. Ensure http://localhost:1420 is in H5 allowedOrigins (PUT only when
     missing).
  5. Regenerate a fresh H5 token via /api/h5-access/regenerate.
  6. Build the full URL with serverUrl + forceH5=1 + h5Token, copy it to
     the clipboard, and print it.
  7. With -Open, also Start-Process the URL in the default browser.
  8. With -Quiet, only print the URL on stdout (for piping into other
     tools, e.g. un run mcp:test:open chaining).

Surfaced via two npm scripts:
  bun run mcp:test         # generate URL + clipboard
  bun run mcp:test:open    # also open in default browser

Deliberate non-goals:
  - Does NOT start server / Vite. Those are persistent dev processes the
    developer wants to control lifecycle for; auto-starting them from a
    one-shot script makes "is it still running?" questions ambiguous and
    leaves zombies on script exit. The script tells the user the exact
    command to run if either is down.
  - Does NOT touch Electron. This workflow is the documented browser path
    for Chrome DevTools MCP smoke; the Electron path doesn't need any of
    this and runs through bun run electron:dev.
  - Does NOT modify vite.config.ts proxy or desktopRuntime.ts forceH5
    bypass. Both are already on this branch, and validation of those two
    being correctly wired is what step 3 covers.

Doc: docs/desktop/10-local-mcp-testing.md grew a new "step 0 一键脚本"
section pointing at the script and naming the two npm targets, ahead of
the original 4 manual steps so future maintainers find the script first.

Verified by running it: server + vite both up, allowedOrigins pre-existed
so no PUT needed, fresh token issued, URL copied to clipboard, browser
launched into the welcome screen with the four task cards rendering
correctly. Screenshot in artifacts/desktop-welcome-task-cards-final.png.

Confidence: high
Scope-risk: narrow

* feat(desktop): zero-token "Recent activity" panel with hand-off button

Closes the painful gap where opening a new chat in a project means
re-explaining "what was I just doing" to the model. The user explicitly
called out that they don't want the obvious fix (auto-feeding the
previous transcript or an LLM-generated summary) because that burns
tokens on every new session.

This ships a different shape: derive the activity summary on the server
from on-disk state (session JSONL + git working tree), render it as a
read-only panel on the welcome screen so the user sees it, and only
move bytes into the model when the user explicitly clicks "Continue
from here" — and even then, only a 4-5 line hand-off paragraph (~60
tokens), not the previous transcript.

Server (zero LLM, all derivation):
  - GET /api/projects/recent-activity?workDir=<abs>&excludeSessionId=<id>
  - projectActivityService.getRecentActivity reads:
      - sessionService.listSessions to find the most recent meaningful
        session (skips empty messageCount=0 sessions and any
        excludeSessionId, e.g. the just-created Untitled tab in
        ActiveSession's empty welcome state, so the panel surfaces the
        ACTUAL previous work session).
      - Streams that session's JSONL and derives:
          * lastUserMessageExcerpt (160 chars max, mid-word ellipsis)
          * filesEdited list — Edit/Write/NotebookEdit tool_use file_paths,
            de-duplicated in first-appearance order, capped at 8.
        Hard-capped at 50k JSONL lines so giant transcripts don't hang
        the welcome render.
      - getRepositoryContext (existing, cached) for branch + default
        branch + dirty state, plus two cheap git rev-list/status calls
        for ahead/behind counts and exact dirty file count.
  - Strict timeouts on each git command so the panel never blocks UI.
  - Wired into router.ts under case 'projects'.

Frontend:
  - RecentActivityCard reads via projectsApi.recentActivity.
  - Two paths into it:
      1. EmptySession (sidebar "New session" → user picks workDir):
         shows once a workDir is chosen. "Open this session" switches
         to the previous tab; "Continue from here" prefills the
         hand-off into the empty composer's textarea.
      2. ActiveSession's empty welcome state (sidebar "在 X 中新建会话"
         creates a session immediately): excludeSessionId={activeTabId}
         skips the brand-new empty session, hideContinueSessionButton
         hides "Open this session" since the user is already on an
         empty tab. "Continue from here" dispatches the existing
         cc-haha:composer-prefill window event so ChatInput picks it
         up — same plumbing the welcome task cards use.
  - Auto-refreshes every 60s — git state changes between renders
    (commits in another terminal, etc.) get picked up without UI work.
  - Hidden on phone-sized H5 (composer is dense enough already).

Hand-off paragraph contents (i18n'd, ~60-100 tokens):
  Last session on <branch>: "<title>".
  Files touched: <up-to-5-files> (+N more).
  Local is ahead of upstream by N commit(s).
  N file(s) have uncommitted changes.

  Please continue from there. Pick up by ...(describe the next step).

Token cost summary:
  - Render: 0 tokens. Pure on-disk derivation, never sent to a model.
  - "Open this session": 0 tokens. Just a tab switch.
  - "Continue from here": ~60-100 tokens, ONLY if the user types more
    and presses send. The user sees the prefill in the textarea before
    sending and can edit or delete it.

i18n: 18 new keys (empty.recentActivity.*) localized for the same five
locales as the orchestration toggle (en/zh/zh-TW/jp/kr).

Live-tested via Chrome DevTools MCP on this branch:
  - Opened cc-haha welcome path.
  - Server returns correct shape: branch=feat/agent-routing-hardening,
    aheadCount=6, dirtyCount=13, lastSession with 10 messages and
    correct title.
  - excludeSessionId correctly skips the empty just-created session and
    surfaces the previous "请审查我当前分支..." session.
  - "Continue from here" click prefills the textarea with the 152-char
    Chinese hand-off paragraph; localStorage and WS state untouched
    (zero token cost confirmed).
  - Screenshots: artifacts/desktop-recent-activity-card.png and
    artifacts/desktop-recent-activity-handoff.png.

bun run lint passes (tsc --noEmit clean).

Confidence: high
Scope-risk: moderate (new endpoint, but read-only and isolated; new UI
panel, but additive only — empty workdir or no prior sessions just
hides the card)
Tested: live MCP smoke per above.
Not-tested: server-side unit tests for projectActivityService (planned
follow-up; live integration validated the happy path end-to-end).

* fix(desktop): compact recent-activity card so composer stays in view

Live MCP smoke caught a layout regression. Welcome screen vertical
budget on a 923px viewport is roughly:
  79  tab bar
  + ~250 hero block (logo + title + subtitle)
  + ~165 task cards
  + ~222 composer (ChatInput in hero variant)
  + ~64  flex padding
  ≈ 780px before activity panel

The first cut of RecentActivityCard rendered a 235px column block
(title row, chips row, italic excerpt row, file-chips row, action row).
Together with the task cards that pushed total content to ~410px
beyond what fits, and ActiveSession's empty-state wrapper used
flex flex-1 + justify-center without overflow handling — so the
sibling ChatInput got pushed below viewport bottom (composer.bottom
1112 on a 923 viewport, visible:false confirmed via DOM measurement).

Two changes:

1. Compact the card to ~127px (down 108px, roughly half) by collapsing
   to a single horizontal row: icon | title + chips column | action
   buttons. The italic last-user-message excerpt and the file-name
   chip strip both go away — they're nice-to-haves, not load-bearing.
   The data they carried is preserved: filesEditedCount stays as a
   chip with a tooltip showing the first 5 file basenames; the lead
   excerpt is still in the hand-off paragraph the "Continue from here"
   button injects, so the user gets it where it counts (in the
   composer) instead of as ambient ornament on the welcome screen.
   Action buttons collapse their labels under sm: breakpoints to keep
   the card single-row even in narrow side panels.

2. Add min-h-0 + overflow-y-auto to ActiveSession's isEmpty wrapper.
   Belt-and-suspenders: future taller content (extra cards, longer
   localized strings) now scrolls inside the welcome region instead of
   shoving the composer offscreen. Composer remains a flex sibling and
   stays anchored at the bottom regardless of content height above it.

Verified live in MCP browser:
  recentActivity height: 235 → 127px
  composer.bottom: 1112 (offscreen) → 907 (visible, on a 923 viewport)
  textareaVisible: false → true
  Tested across two projects with very different activity profiles
  (cc-haha: short title + 0 files edited; layout-editor: long title +
  20 files edited + 1876 messages). Both render in single horizontal
  row.
  Screenshot: artifacts/desktop-recent-activity-compact.png

bun run lint passes (tsc --noEmit clean).

Confidence: high
Scope-risk: narrow (CSS / layout only, no behavioral change).

---------

Co-authored-by: 你的姓名 <you@example.com>
2026-06-10 06:40:48 +08:00
程序员阿江(Relakkes)
e8bf789961 fix(provider): request non-stream OpenAI chat tests (#638)
Ensure custom OpenAI-compatible provider tests explicitly request non-stream responses for both direct connectivity checks and the proxy pipeline. This avoids default SSE responses being parsed as empty JSON by provider validation.

Tested: bun test src/server/__tests__/providers.test.ts --test-name-pattern "requests non-stream OpenAI Chat responses during provider tests"
Tested: bun test src/server/__tests__/providers.test.ts
Tested: bun test src/server/__tests__/proxy-transform.test.ts
Tested: bun run check:server
Confidence: high
Scope-risk: narrow
2026-06-09 22:16:38 +08:00
程序员阿江(Relakkes)
d3b2f868a9 fix: guard prewarm resume shutdown paths (#611)
Prevent desktop prewarm launches from inheriting interrupted-turn resume state, while leaving explicit user-message startup unchanged. Also stop server-side background schedulers during shutdown so app quit cannot leave scheduled task runners alive.

Tested:

- bun test src/server/__tests__/conversation-service.test.ts

- bun test src/server/__tests__/conversations.test.ts -t "prewarm"

- bun test src/server/__tests__/conversations.test.ts -t "permission"

- bun test src/server/__tests__/diagnostics-service.test.ts -t "keeps fatal startup errors visible on stderr while recording diagnostics"

- bun test desktop/electron/services/windows.test.ts desktop/electron/services/tray.test.ts desktop/electron/services/sidecarManager.test.ts src/server/__tests__/server-shutdown.test.ts

- bun run check:server
2026-06-09 21:44:46 +08:00
程序员阿江(Relakkes)
ddcfa5faae feat(adapters): add WhatsApp linked-device support (#573)
Add a WhatsApp adapter backed by Baileys linked-device auth, plus desktop QR binding UI, server config endpoints, sidecar startup wiring, tests, and documentation.

Constraint: Uses WhatsApp Web linked-device auth, not Meta WhatsApp Business Cloud API.

Tested:
- cd adapters && bun run check:adapters
- bun test src/server/__tests__/adapters.test.ts
- cd desktop && bun run check:desktop
- bun run check:native
- bun run check:persistence-upgrade
- bun run check:docs

Not-tested:
- Live WhatsApp QR pairing, because no WhatsApp account/device was exercised here.
- bun run check:server, because the existing src/server/__tests__/conversations.test.ts timeout still fails independently.

Confidence: medium
Scope-risk: moderate
2026-06-09 21:31:46 +08:00
程序员阿江(Relakkes)
db631dfdfd fix(desktop): show full tool error details (#625)
Surface expanded error output for tool cards whose previews previously returned before rendering the result body. Keep successful Bash/Read/Edit/Write outputs hidden as before while exposing error details with wrapping and copy support.

Tested: cd desktop && bun run test -- src/components/chat/chatBlocks.test.tsx
Tested: bun test src/server/__tests__/conversations.test.ts -t "should switch from bypass permissions back to default without restarting" --timeout=20000
Tested: bun run verify
Confidence: high
Scope-risk: narrow
2026-06-09 21:07:45 +08:00
程序员阿江(Relakkes)
385b996736 fix: defer runtime restarts during active turns (#626)
Avoid killing an in-flight desktop SDK turn when runtime config changes during tool execution. Persist the requested runtime immediately, then restart after the turn emits its terminal result.

Tested:
- bun run check:server
- real provider tool-turn model switch smoke
- bun run check:coverage (changed-line coverage passed; unrelated global lanes still fail)

Scope-risk: narrow
2026-06-09 21:06:24 +08:00
程序员阿江(Relakkes)
3ff6a79e62 feat(adapter): add Telegram command menu (#596)
Add Telegram command-menu sync plus /resume, /provider, /model, and /skills command handling with paginated inline selections.
Expose the adapter HTTP calls needed by those commands and harden stale WebSocket session cleanup during Telegram resume/reset flows.

Tested:
- cd adapters && bun test telegram/__tests__/commands.test.ts telegram/__tests__/menu.test.ts
- cd adapters && bunx tsc --noEmit -p tsconfig.json
- bun run check:adapters
- bun run check:persistence-upgrade
- bun run check:native
- bun run verify (coverage lane failed on existing root WebSocket Chat Integration timeout)

Not-tested:
- Live Telegram bot smoke, no bot token/session available.

Confidence: medium
Scope-risk: moderate
2026-06-09 21:01:30 +08:00
你的姓名
4690cd9199 release: v0.5.6 2026-06-09 16:58:15 +08:00
你的姓名
f65ae5b2b9 fix: update check误判equal version为available update + about设置页满宽
- ElectronUpdaterService加currentVersion选项,checkForUpdates时比较feed版本
- 自包含的isNewerVersion():忽略prerelease/build后缀,等版本->非更新
- AboutSettings去掉max-w-lg/mx-auto约束,面板左右撑满
- 新增4个版本比较回归测试(等版本/旧版本/新版本场景)

Tested: updater.test 14 pass, generalSettings/updateStore 67 pass, lint通过
2026-06-09 16:55:31 +08:00
你的姓名
78ca7070ad chore: ignore .kiro IDE local config
Scope-risk: narrow
2026-06-09 16:53:06 +08:00
你的姓名
303afcfff4 feat(orchestration): lower delegation threshold and propagate project conventions to sub-agents
Bias the orchestrator toward delegating medium tasks (>2 steps or >1 file), reserving direct action for trivial work. Require restating relevant project rules (e.g. codegraph-first) inside sub-agent prompts, since read-only agents (Explore/Plan) run without project memory.

Tested: bun test conversations.test.ts (orchestration cases pass)

Confidence: high

Scope-risk: narrow
2026-06-09 16:52:34 +08:00
程序员阿江(Relakkes)
d11748d61a fix(desktop): make memory files preview-first (#533)
Tested: cd desktop && bun run test src/__tests__/memorySettings.test.tsx
Tested: bun run check:desktop
Scope-risk: narrow
Confidence: high
2026-06-09 16:39:17 +08:00
程序员阿江(Relakkes)
0a8f247781 feat(desktop): add plugin list bulk toggles (#527)
Tested:
- cd desktop && bun run test -- src/__tests__/pluginsSettings.test.tsx src/stores/pluginStore.test.ts --run --reporter=verbose --pool=forks --maxWorkers=1 --minWorkers=1
- bun run check:desktop
- git diff --check

Scope-risk: moderate
Confidence: high
2026-06-09 16:33:15 +08:00
程序员阿江(Relakkes)
a7263cbaf7 fix(agent): avoid concurrent worktree config writes (#572)
Git writes upstream branch config when a worktree starts from origin/main. Use the already-resolved base SHA as the worktree start point so parallel agent worktrees do not race on shared .git/config.

Tested: bun test src/utils/__tests__/worktree.test.ts

Tested: real gpt-5.5 Sub2API run with four worktree-isolated agents

Not-tested: bun run check:server currently fails unrelated WebSocket restart/timeouts

Confidence: high

Scope-risk: narrow
2026-06-09 16:31:55 +08:00
你的姓名
7e55d6cee9 feat: context-exhausted new-session suggestion (方案3) + orchestration mode
方案3 (desktop, manual-guided): when compactions keep firing only a turn or two
apart — the same thrash the CLI circuit breaker trips on — chatStore now appends
a one-time visible system notice suggesting the user start a fresh session
(the prior summary stays in the current one). Detection lives in module-level
state (compactionThrashBySession) so PerSessionState is untouched; counts user
turns between compactions, suggests after 3 rapid ones, deduped per session and
reset on /clear. New i18n key chat.contextExhausted (en/zh/zh-TW/jp/kr).

Also includes in-progress orchestration/coordinator work on the same files
(orchestrationPrompt.ts, conversationService coordinatorMode --append-system-prompt,
ws handler/events, chatStore, ChatInput, sessionRuntimeStore, types/chat) —
committed together per request.

Tested:
- bunx vitest run desktop/src/stores/chatStore.test.ts (102 pass, incl. new
  方案3 rapid-compaction suggestion + spread-out negative case)
- cd desktop && bun run lint (clean)
- get_diagnostics clean across changed server + desktop files
Not-tested: full bun run check:server / check:desktop gates; pre-existing WS
runtime-restart integration tests are flaky in this env (CLI code 143).
Confidence: medium-high
Scope-risk: moderate
2026-06-09 15:17:52 +08:00
你的姓名
2cf6040451 fix(server): treat relay context_too_large as prompt-too-long
Third-party Anthropic-compatible relays reject oversized requests with a 400 context_too_large / 'exceeds the context window' body instead of Anthropic's 'prompt is too long' wording. That fell through to a raw API error with no recovery hint. Normalize it via a new isContextWindowExceededMessage() so getAssistantMessageFromError and classifyAPIError route it onto the existing prompt-too-long handling, surfacing the actionable 'Context limit reached / compact or clear' guidance (TUI) and businessError.prompt_too_long copy (desktop).

Tested: bun test src/services/api/errors.test.ts (4 pass)

Not-tested: full bun run check:server gate

Confidence: high

Scope-risk: narrow
2026-06-09 14:42:20 +08:00
你的姓名
1440ce01ce fix(server): guard stop-generation force-kill by process instance (code 143 race)
Stop-generation set a 3s force-kill timer that called stopSession(sessionId). If the user switched provider/model in that window, the restart spawned a NEW CLI process under the same sessionId; the stale timer then SIGTERM-killed the new process during its startup grace window -> 'CLI exited during startup with code 143'.

Fix: tag each spawned process with an instanceId (sessionId#N). handleStopGeneration captures the live instanceId up front; the 3s fallback calls stopSessionInstance(sessionId, instanceId), which only kills when the current live process still matches that instance. A restart replaces the instance, so the stale timer becomes a no-op.

conversationService.ts: instanceId field + counter, getActiveInstanceId(), stopSessionInstance().
ws/handler.ts: handleStopGeneration uses the instance-guarded fallback.
conversations.test.ts: instance-guard regression tests.

Tested: bun test conversations.test.ts -t instance/stopSessionInstance (pass). Staged via hunk selection to exclude unrelated in-progress orchestration work that shares these files.
Confidence: high
Scope-risk: narrow
2026-06-09 14:29:41 +08:00
你的姓名
110a72a97e feat(compact): detect context-exhausted state for new-session suggestion (方案3 CLI foundation)
When the auto-compaction circuit breaker has tripped and the context is still over threshold, autoCompactIfNeeded now returns contextExhausted:true. This is the signal the desktop will use to suggest starting a fresh session (carrying the latest summary) instead of degrading silently. Pure predicate isContextExhausted() added + tested (15 pass). Desktop wiring (ws translation + chat suggestion card + new-session handoff) is intentionally not included here — those files currently carry unrelated in-progress changes.
2026-06-09 13:53:38 +08:00
你的姓名
b0f368ef97 fix(compact): stop re-compaction thrash and post-compact 100% context meter
Two context-management fixes (CLI/sidecar side):

1. Re-compaction loop (方案1+2). Auto-compaction previously only circuit-broke
   on hard failures; a compaction that succeeded but left context still over
   threshold reset the failure counter and re-compacted every turn, burning a
   summary API call each time.
   - autoCompact.ts: isIneffectiveCompaction() — when truePostCompactTokenCount
     is still >= threshold, count it toward the existing circuit breaker instead
     of resetting to 0; query.ts now honours that count on the success path.
   - shouldThrottleAutoCompact() — suppress proactive autocompaction for a few
     turns after a compaction, unless at the hard blocking limit (so we never
     risk prompt-too-long by throttling).

2. Post-compact context meter stuck at ~100% (sessionService.ts). The transcript
   context estimate counted the entire pre-compact history plus the summarization
   call's huge input_tokens, pinning the indicator near full until the next turn.
   Now the estimate is scoped to messages after the latest compact_boundary, with
   a model fallback for the just-compacted/no-new-turn case. Non-compacted
   sessions are unchanged.

Tested:
- bun test src/services/compact/autoCompact.test.ts (12 pass; throttle + ineffective predicates)
- bun test src/server/__tests__/conversations.test.ts -t "context" (6 pass incl. new compact-boundary scoping case)
Not-tested: 4 pre-existing WebSocket runtime-restart integration tests fail in this env (CLI subprocess code 143); unrelated to these changes.
Confidence: high
Scope-risk: moderate
2026-06-09 13:48:40 +08:00
你的姓名
3cbed50ca9 feat(desktop): point repo + update source to this fork, credit fork maintainer
- Update source: desktop/package.json publish.owner NanmiCoder -> 706412584 (electron-updater
  now checks this fork's releases via app-update.yml). Also bump homepage and the legacy
  Tauri updater endpoint to the fork.
- About page: GITHUB_REPO / issues / releases / changelog and the displayed repo name now
  point to 706412584/cc-haha. Sidebar repo link and ActivitySettings profile subtitle too.
- Attribution: original author (NanmiCoder + Bilibili/Douyin/Xiaohongshu social links) is
  preserved unchanged. Added a new "Fork maintainer" credit (706412584) with i18n keys
  settings.about.forkMaintainer / forkMaintainerHint across en/zh/zh-TW/jp/kr.

Why: this is a self-maintained fork shipping its own builds; pointing the updater at upstream
risked overwriting custom builds with upstream releases and hid the fork's own releases.

Tested:
- bunx vitest run src/pages/ActivitySettings.test.tsx src/__tests__/generalSettings.test.tsx (59 pass)
- cd desktop && bun run lint (clean)
Not-tested: full verify (long-running)
Confidence: high
Scope-risk: moderate
2026-06-09 04:36:56 +08:00
你的姓名
87e37d30d1 feat: v0.5.5 — more built-in agents, composer skill/plugin picker, AskUserQuestion notice
Built-in agents:
- Add debugger, security-reviewer, refactor, migration, docs-writer, performance, commit-pr built-ins and register them in getBuiltInAgents().

Desktop:
- AskUserQuestion prompts no longer vanish silently: when an unanswered question card is cleared by turn end (message_complete / error) — e.g. a malformed question call — a visible system notice (chat.questionDropped) is appended instead. Added across en/zh/zh-TW/jp/kr.

Versioning:
- Bump desktop/package.json to 0.5.5 and add release-notes/v0.5.5.md.

Tested:
- bunx vitest run desktop/src/stores/chatStore.test.ts (99 pass, incl. dropped-question notice for message_complete + error, and no notice for non-AskUserQuestion)
- bun test src/tools/AgentTool/builtInAgents.test.ts (pass)
- cd desktop && bun run lint (clean)
- build:windows-x64 + package-smoke (PASS, Claude-Code-Haha-0.5.5-win-x64.exe)
Not-tested: full bun run verify / check:server (long-running)
Confidence: high
Scope-risk: moderate
2026-06-09 03:25:36 +08:00
你的姓名
433239dbaa chore(repo): untrack local tooling dirs (.codegraph, .vscode)
These were accidentally committed in the previous catch-all commit. They're local tool artifacts (.codegraph index, .codegraph/daemon.pid runtime PID, personal IDE settings) that don't belong in the repo. Adding them to .gitignore and removing from tracking via git rm --cached so working copies stay intact.
2026-06-09 02:47:43 +08:00
你的姓名
fee22c80ca 提交 2026-06-09 02:35:06 +08:00
你的姓名
7b81092aad feat(desktop): composer skill/plugin picker, send-now, theme follow-system, read empty pages
ChatInput / SkillPickerMenu:
- + menu now has 4 entries: file, slash command, Skills, Plugins.
- Skills/Plugins open an inline picker above the composer (mirrors the file-search popover: ↑↓ navigate, Enter pick, Esc dismiss). Picking inserts a "@skill:<name>" / "@plugin:<name>" token at the cursor — visible to the user, also legible to the agent as "use this skill/plugin".

Queue "send now":
- sendQueuedMessageNow no longer aborts the running turn or promotes-and-waits — it just sends the message immediately. Removes the prior "Request was aborted" / code 143 retry loop caused by mid-turn stop_generation.

Theme follow-system persistence:
- settingsStore.loadSettings now honours a locally stored "system" theme over whatever concrete value the server last saw, since the server intentionally rejects "system". Fixes follow-system reverting to a concrete theme on every restart.
- uiStore.test.ts: updated the toggleTheme cycle test to include the system step (was a stale failure independent of these changes).

FileReadTool:
- Treat pages: "" as undefined ("read whole file") so the agent's empty-string slips don't trap the tool in a validation-error retry loop.

i18n:
- Added chat.openSkills / openPlugins / skillPicker.{title,empty} / pluginPicker.{title,empty} / sendNow across en / zh / zh-TW / jp / kr.
- Did NOT touch unrelated settings.skills.recommended.* / settings.plugins.* additions present in the working tree from another in-progress task.

Tested:
- bunx vitest run src/stores/chatStore.test.ts -t "message queue" (10 pass)
- bunx vitest run src/stores/settingsStore.test.ts src/stores/uiStore.test.ts (32 pass)
- bun run lint (desktop tsc --noEmit, clean)
- Windows package-smoke after build:windows-x64 (PASS)
Not-tested: full bun run check:desktop / check:server (long-running)
Confidence: high
Scope-risk: narrow
2026-06-09 02:27:17 +08:00
你的姓名
644914a8b4 feat(agents): add test-author + code-reviewer built-ins, project-level game-developer
Built-in additions (always shipped):
- test-author: writes/runs tests for changed code, detects framework, reports changed-line coverage.
- code-reviewer: read-only static review, finds bugs/smells/security issues, ends with REVIEW: APPROVE | CHANGES_NEEDED.

Project-level addition under .claude/agents/:
- game-developer: covers Unity/Unreal/Godot and web JS engines. The system prompt forces engine-version detection (ProjectVersion.txt, *.uproject, project.godot, package.json), prefers querying real symbols already in the project (codegraph) over recalling APIs, and falls back to verifying uncertain APIs against the engine's official docs — to defend against API hallucinations across versions.

.gitignore: switched ".claude/" to ".claude/*" with explicit "!.claude/agents/**" so project agents can ship in the repo while user-private files like .claude/settings.json stay ignored.

Tests:
- builtInAgents.test.ts: asserts test-author + code-reviewer are registered, game-developer is NOT a built-in.
- gameDeveloperAgent.test.ts: parses .claude/agents/game-developer.md through the runtime parser and asserts the strengthened guidance (codegraph, official docs, ProjectVersion.txt) is present.

Tested: bun test src/tools/AgentTool/builtInAgents.test.ts src/tools/AgentTool/gameDeveloperAgent.test.ts (5 pass)
Not-tested: full bun run check:server (long-running)
Confidence: high
Scope-risk: narrow
2026-06-09 02:19:05 +08:00
你的姓名
e66751bfe9 fix(desktop): fully pause message queue on Stop (multi-idle safe)
v0.5.3's one-shot skip only blocked a single auto-drain, but a Stop emits
multiple idle events (message_complete + status idle); the second one still
drained one queued message and flipped the button back to running. Replace
the one-shot skip with a sticky queueDrainPaused flag: Stop pauses all
auto-drains until the user sends their next message (which clears it and
fires immediately, ahead of the queue). The queue then resumes FIFO on the
next idle. Bumps desktop app to 0.5.4 with release notes.

Tested: tsc --noEmit; Vitest chatStore 94 passed (incl. multi-idle stop /
priority message / queue resume); browser end-to-end (Stop keeps button at
Run with queue intact, priority send immediate, queue resumes); Windows x64
NSIS package + package-smoke PASS.
Scope-risk: narrow
Confidence: high
2026-06-08 04:20:33 +08:00
你的姓名
408e04e500 fix(desktop): don't flush message queue on user Stop; remove donation section
Stop no longer flushes the queue: a user-initiated Stop now skips exactly the
auto-drain triggered by its resulting idle (one-shot guard). The queue stays
intact, the next message the user sends fires immediately (priority, not
queued — even with items waiting), and the queue resumes FIFO on the next idle
after that turn. Queue enqueue/drain/edit mechanics are otherwise unchanged.

Also removes the Buy-Me-a-Coffee / donation section from README (zh/en).
Bumps desktop app to 0.5.3 with release notes.

Tested: desktop tsc --noEmit; Vitest chatStore 94 + pages/ActiveSession 43
passed (incl. stop-no-flush / priority-message / queue-resume case); Windows
x64 NSIS package + package-smoke PASS.
Scope-risk: narrow
Confidence: high
2026-06-08 03:58:09 +08:00
你的姓名
e179b36e83 feat(desktop): chat message queue + provider model UX + bundled skills
Desktop chat message queue (v0.5.2):
- Queue follow-up messages (Enter) while the agent is busy; FIFO auto-drain
  on true idle. Collapsible, height-capped to-do panel above the composer so
  it never inflates the composer. Per-item move-to-top/edit/delete/clear.
  Never drains mid-stream, during tool execution, or while a permission
  prompt is pending; Stop cancels the current turn but keeps the queue.
  Per-session isolation. In-memory only (no persistence schema change).

Also included in this commit:
- Settings provider form: model-id comboboxes + fetch-models + per-slot
  context-window auto-fill.
- Bundled skills: defineGoal, pdf, screenshot.
- Bump desktop app to 0.5.2 with release notes.

Tested: desktop tsc --noEmit; Vitest chatStore/pages/ActiveSession 136 passed
(incl. 7 queue tests); browser end-to-end queue smoke; Windows x64 NSIS
package + package-smoke PASS.
Not-tested: full check:desktop suite has pre-existing unrelated failures
(electron packaging, theme cycling, ThinkingBlock label); check:server.
Scope-risk: moderate
Confidence: medium
2026-06-08 03:31:12 +08:00
你的姓名
d88c10c2df fix(server): defer runtime-config restart until session is idle
In-progress generations were interrupted when a set_runtime_config
arrived mid-stream: the handler immediately stopped and restarted the
CLI subprocess. Track per-session busy state from outbound status
events and queue the restart, applying it (with the latest override,
collapsing multiple toggles) once the session returns to idle.

Bumps desktop app to 0.5.1 with matching release notes.

Tested: desktop tsc --noEmit; Windows x64 NSIS package + package-smoke.
Not-tested: check:server, desktop Vitest.
Scope-risk: narrow
Confidence: medium
2026-06-08 02:18:00 +08:00
你的姓名
1923d5553d feat: 服务商模型下拉、跟随系统主题、会话级思考开关、NSIS 自定义安装
桌面端:
- 服务商表单的 4 个 Model ID 改成可下拉+可手动输入;新增"获取模型"按钮
  通过 ${baseUrl}/v1/models 拉取(自动适配 OpenAI Bearer / Anthropic x-api-key)
- 配色主题新增"跟随系统",监听 prefers-color-scheme 自动切换 light/dark
- ModelSelector 新增"思考模式 开/关"会话级开关,默认回退全局
- NSIS 启用许可页+ 自定义安装路径;首次复用 LICENSE 作 license.txt

Server:让会话级 thinking 真正生效
- WS schema 加 thinkingEnabled?: boolean
- handler 把字段写进 runtimeOverrides、纳入 prev/next diff、持久化进 jsonl
- sessionService.appendSessionMetadata / getSessionLaunchInfo / 摘要解析
  补充 thinkingEnabled 字段
- resolveDesktopThinkingMode 支持三态:override 优先(true=enabled、
  false=disabled),undefined 回落到全局 alwaysThinkingEnabled
- RuntimeSettings.thinking 类型放宽为 'enabled' | 'disabled',
  下游 conversationService 已支持,原生拼成 --thinking 参数

i18n:zh / zh-TW / en / jp / kr 同步
2026-06-07 18:39:02 +08:00
程序员阿江(Relakkes)
760295dcd8 test(desktop): update chatBlocks thinking assertions to the done-state label
Commit 449ff0b0 changed completed thinking blocks to render the
'thinking.labelDone' title ("Thought"/"已思考") instead of the static
"Thinking" label, and updated ThinkingBlock.test.tsx but not
chatBlocks.test.tsx. The three inactive/default-state cases there still
queried the toggle button by /Thinking/ and failed. Match the new
done-state label (/Thought/); the active-state case is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 23:28:11 +08:00
程序员阿江-Relakkes
4bea9ee687
Merge pull request #717 from TW199501/feat/i18n-ja-ko-zh-tw
feat(desktop): add Japanese, Korean, and Traditional Chinese UI locales
2026-06-04 23:25:00 +08:00
程序员阿江(Relakkes)
561ae1aefe feat(desktop): fill jp/kr/zh-TW keys added since the PR branched
Main added two TranslationKeys after this branch's base — 'thinking.labelDone'
(the "Thought" label shown after thinking completes) and
'slashCmd.agent.description' — which the jp/kr/zh-TW locale files did not yet
cover, so the Record<TranslationKey, string> contract failed tsc after merging
main. Add both keys to each new locale:

- jp: '思考完了' / '選択した Agent でプロンプトを実行'
- kr: '사고 완료' / '선택한 Agent로 프롬프트 실행'
- zh-TW: '已思考' / '使用指定 Agent 執行提示'

en/zh are unchanged. `tsc --noEmit` is now clean and the i18n + timestamp
suites pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 23:20:55 +08:00
程序员阿江(Relakkes)
54b8c3f841 Merge commit '892f5e19' into i18n-fill-717 2026-06-04 23:13:43 +08:00
程序员阿江(Relakkes)
892f5e193a fix(desktop): restore permission mode after exiting plan mode (#623)
桌面端用 bypassPermissions 进入计划模式、退出后,权限选择器停留在"计划模式",
导致下一个工具调用又弹权限申请。两层根因都修:

- 回传链路断裂:服务端 WS handler 把 CLI 广播的权限模式变化(status:null +
  permissionMode)当成 thinking 丢弃,桌面端协议也没有承载权限模式的入站消息,
  CLI 恢复后的权限永远同步不到 UI。新增 permission_mode_changed 回传通道,桌面端
  据此校正选择器(只更新本地、不回发避免回环;未知模式忽略)。

- 重启抹掉 prePlanMode:bypass→plan 因策略"切换涉及 bypass 就重启"而重启 CLI,
  新进程直接以 plan 启动、prePlanMode 为空,ExitPlanMode 只能恢复成 default 而非
  bypassPermissions。收窄 needsRestart 为只在"进入 bypass"时重启;从 bypass 切出
  保持进程不变、走进程内 transition,CLI 才会栈存 prePlanMode 并在退出 plan 时
  正确恢复 bypass,与 TUI 行为一致。

测试:服务端 ws-memory-events 新增权限回传 + 重启策略用例;桌面端 chatStore 新增
回传校正 + 防回环 + 未知模式忽略用例。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 22:29:01 +08:00
程序员阿江(Relakkes)
4a53a1659f fix(proxy): timeout stalled OpenAI-compatible streams (#548)
Guard OpenAI-compatible proxy streaming bodies with the configured AI request timeout so a provider that emits partial SSE and then idles cannot leave proxy consumers waiting forever.

This is a proxy-level fix found while investigating #548; it does not claim to close the broader desktop interruption issue.

Tested: bun test src/server/__tests__/proxy-network-settings.test.ts

Tested: bun run check:server

Confidence: medium

Scope-risk: narrow
2026-06-04 22:27:04 +08:00
程序员阿江(Relakkes)
82e857163f fix(desktop): gate scheduled-task notification poll on server readiness
The desktop notification poller fired on mount, racing the bootstrap that
resolves the dynamic server URL and confirms /health. Its first requests hit
the uninitialized default base URL and failed with "Failed to fetch", logging
spurious client_api_request_failed warnings to the diagnostics panel.

Add a whenDesktopServerReady() signal resolved once initializeDesktopServerUrl
sets the base URL and the healthcheck passes, and gate the poller on it so it
only starts once the server is reachable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 21:31:57 +08:00
程序员阿江(Relakkes)
449ff0b0ff fix(desktop): show "已思考" label after thinking completes (#732)
The thinking block title always rendered the static "思考中"/"Thinking"
label; isActive only toggled the animated dots. Once thinking finished
the dots disappeared but the in-progress text stayed. Switch the label
to "已思考"/"Thought" when the block is no longer active.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 20:19:11 +08:00
程序员阿江(Relakkes)
5688fe10a7 test(e2e): isolate ANTHROPIC_*MODEL env from the models-API fixtures
The models API derives its model list from ANTHROPIC_MODEL and the
ANTHROPIC_DEFAULT_{HAIKU,SONNET,OPUS}_MODEL env vars. A developer who exports
these for a custom provider (e.g. MiniMax) leaked them into the no-provider
fixture: the four collapsed to one model, the default model became the
custom one, and switched-model names fell back to the raw id — failing five
"available models" / "default model" assertions on that machine while passing
on clean CI.

Clear those four vars in the e2e setup and restore them in teardown, matching
the existing CLAUDE_CONFIG_DIR / CLAUDE_CLI_PATH isolation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 20:12:54 +08:00
程序员阿江(Relakkes)
837337901d fix(diagnostics): stop logging expected states as errors and isolate tests
The diagnostics panel filled with red ERROR/WARN entries during normal use
because the capture layer treated "wrote to console" as "something failed",
with no severity gatekeeping. Much of the noise was also test runs leaking
into the user's real ~/.claude/cc-haha/diagnostics.

- diagnosticsService: default unclassified events to info (not error);
  drop writes under NODE_ENV=test when no CLAUDE_CONFIG_DIR is set; document
  the console-capture contract (expected states use console.debug/info).
- index: don't install console/process capture under bun test.
- api/diagnostics: an ingested event with missing severity defaults to info.
- oauthRefreshLog (new): token refresh failure is gracefully handled and is
  never an error — expected expiry (401/403/revoked) logs at debug, anything
  else at warn. Wired into both Haha OAuth services.
- conversationService: classify cli_runtime_exit severity by exit code, so
  clean/SIGTERM/SIGKILL exits are info and only abnormal codes stay error
  (real "chat died" crashes remain perceivable).
- ws/handler: streaming partial tool-input JSON is normal — debug, not warn.

Tests: oauth-refresh-log + cli-exit-severity unit tests; diagnostics-service
covers the info default and the test-isolation guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 20:12:45 +08:00