cc-haha/scripts/dev-coordinator-e2e.ps1
小橙子 e6332810ed
feat(coordinator): research-fork opt-in — let coordinator mode use a single fork path for cache-sharing fan-out (#22)
Resolves the only path that could plausibly cut coordinator-mode token
cost without changing architecture. Coordinator mode normally rejects
implicit forks (the orchestrator is supposed to delegate to typed
workers); this opt-in carves out a narrow window for cache-sharing
research fan-out only.

Why
===
Coordinator-mode parallel research dispatches several workers of the
same type. Each worker prefills its own system prompt + full tool
schema, easily 10-30k tokens × N workers. Fork's byte-identical API
prefix bypasses that entirely — the parent's prompt cache is fully
reused, marginal cost is just the directive bytes. On a 3-worker
research fan-out we expect ~10x prefill savings for the research
phase.

Out of scope: implementation/verification spawns still go through
typed workers (independent context, independent cache). This PR is
about research only.

What's in this PR
=================
- forkSubagent.ts
  - New helper: isCoordinatorResearchForkEnabled() — gated on
    feature('FORK_SUBAGENT') + isCoordinatorMode() + non-interactive
    + CLAUDE_CODE_COORDINATOR_RESEARCH_FORK=1.
  - New constant: COORDINATOR_RESEARCH_FORK_SUBAGENT_TYPE = 'fork'
    (single source of truth for what callers spell as subagent_type).
  - New ForkMode union: 'normal' | 'coordinator-research'.
  - buildForkedMessages / buildChildMessage now take an optional mode
    parameter (default 'normal'). 'coordinator-research' adds rule 11
    to the boilerplate: "RESEARCH FORK: do not modify files; report
    findings instead". The first 10 rules are byte-identical to the
    normal mode (cache-friendly), and the framing tag at the top is
    unchanged so isInForkChild's recursive-fork guard fires for both
    modes.
- AgentTool.tsx
  - Adds a coord-research-fork branch in the agent-resolution code:
    `subagent_type === 'fork'` + isCoordinatorResearchForkEnabled() →
    routes through the same fork path as an omitted subagent_type.
  - The omitted-subagent_type path in coordinator mode is unchanged
    (still requires explicit type, with the existing helpful error).
  - Computed forkMode threads through to buildForkedMessages.
- invocationLimiter.ts
  - 'fork' default cap = 10 (between verification's 5 and the 8
    fallback). Generous for legitimate parallel research, tight enough
    that pathological loops show up before the budget runs out.
- forkSubagent.test.ts (new, 11 cases)
  - Env-flag gating across all combinations.
  - Boilerplate rule-11 injection in research mode + byte-identical
    first 10 rules across modes (cache contract).
  - Recursive-fork guard fires for both mode variants.
  - Subagent type constant matches the literal "fork".
- invocationLimiter.test.ts
  - One new case: default cap for 'fork' is 10.
- scripts/dev-coordinator-e2e.ps1
  - Adds 8th e2e step: type constant, boilerplate variants, recursion
    guard cross-mode, env-flag gate.

Why opt-in (and double-gated)
=============================
- Lazy-delegation lint and specialist redirect are blacklists: specific
  anti-patterns, near-zero false positives, safe to default on.
- Coordinator research fork is a positive-space affordance: "the
  coordinator MAY use a fork for research." Defaulting it on would
  encourage forks for spawns that legitimately want a clean context
  (e.g. adversarial verification), which conflicts with coordinator's
  independent-worker philosophy.
- Two gates needed: feature('FORK_SUBAGENT') is the bundle-time
  feature flag (we share its plumbing), CLAUDE_CODE_COORDINATOR_RESEARCH_FORK=1
  is the user-level opt-in. Both must be on.

Risks considered
================
- Boilerplate divergence between modes would defeat the parent prompt
  cache. Mitigated: research-mode boilerplate IS the normal-mode
  boilerplate plus rule 11 at the end; first 10 rules are byte-identical
  (covered by a test asserting that property explicitly).
- Recursive-fork guard could miss the new boilerplate. Mitigated: the
  framing tag at the top is unchanged across modes (covered by tests).
- Model abuses 'fork' for non-research spawns. Mitigated: rule 11 in
  the boilerplate forbids file modifications + invocationLimiter cap
  10 contains pathological loops + opt-in flag means nothing changes
  for users who don't turn it on.

Verification
============
- src/tools/AgentTool/forkSubagent.test.ts ............. 11/11 pass
- src/tools/AgentTool/invocationLimiter.test.ts ........ 14/14 pass
- All B1+B2+B3+B4 unit tests .......................... 144/144 pass
- scripts/dev-coordinator-e2e.ps1 ..................... GREEN (8 steps)
- get_diagnostics on changed files .................... clean

Tested: bun test src/tools/AgentTool/forkSubagent.test.ts (11/11); all coordinator unit tests (144/144); scripts/dev-coordinator-e2e.ps1 step 8 GREEN
Not-tested: live coordinator session with the flag on producing actual cache savings (telemetry-only follow-up; the flag is gated and the boilerplate contract is enforced by test)
Confidence: high
Scope-risk: narrow
Constraint: must NOT change buildChildMessage's first 10 rules (cache identity); must NOT change the boilerplate framing tag (recursion guard); must NOT widen the gate to non-coordinator paths (fork stays mutex with normal coordinator delegation by default)
Rejected: defaulting the gate on (positive-space affordance, false-positive cost too high); adding `fork_research` as a new AgentTool input parameter (schema change rejected on B2 by the same logic — too much blast radius); changing fork's tool pool to forbid Edit/Write in research mode (would break useExactTools cache identity; soft-constrained by rule 11 instead)
Directive: when telemetry shows enough adoption to validate the cache savings, consider exposing a /coord-fork slash command alongside the env flag for easier experimentation; do NOT promote to default-on without that data

Co-authored-by: 你的姓名 <you@example.com>
2026-06-12 05:31:50 +08:00

404 lines
16 KiB
PowerShell

# -----------------------------------------------------------------------------
# dev-coordinator-e2e.ps1 — local end-to-end smoke for the coordinator
# behavior changes (B1 + B2).
#
# What it covers (idempotent — safe to re-run):
# 1. Lazy delegation lint (lazyDelegationCheck)
# "based on the findings, fix the bug" → rejected with a re-write hint.
# 2. Specialist router redirect (specialistRouter)
# Coordinator picks `worker` for "code review the auth module" → router
# redirects to `code-reviewer`.
# 3. Invocation limiter near-limit warning (invocationLimiter)
# Last allowed call returns nearLimit=true with a system-reminder
# formatted by formatNearLimitWarning.
# 4. Stall detector transitions (agentStallDetector)
# stalled fires once after threshold; resumed fires once after notify.
# 5. Mode advice (modeAdvice)
# "Migrate the auth subsystem..." → coordinator; "Fix the typo..." →
# normal.
# 6. Task-spec quality (taskSpecQuality, B2)
# Full brief → well-specified; "fix it" → underspecified with a
# missing-dimensions report; reasonable one-liner is NOT flagged.
# 7. Worker-continue advisor (workerContinueAdvisor, B3)
# Same-type worker that already touched the prompt's file(s) is
# surfaced as a continuation candidate; type mismatch and
# file-free prompts are not flagged; error message contract
# preserved (agentId, SendMessage, env flag named).
# 8. Coordinator research fork (forkSubagent, B4)
# Coord-research mode injects the read-only research rule into the
# fork boilerplate while keeping the framing tag intact so the
# recursive-fork guard still fires; subagent_type constant matches
# the literal "fork" callers spell.
#
# This script does NOT spin up the API server or a real LLM call. The five
# B1 behaviors are all enforced inside the AgentTool boundary and are best
# proven via the unit-test suite plus an inline Bun script that exercises
# the real exported modules end-to-end (no mocks). That mirrors the
# "verifier runs the actual command" guidance in AGENTS.md.
#
# Optional flags:
# -SkipUnitTests skip step (a); only run the inline e2e script
# -SkipE2EScript skip step (b); only run the unit-test suite
# -Quiet less ceremony; failures still print
# -----------------------------------------------------------------------------
[CmdletBinding()]
param(
[switch]$SkipUnitTests,
[switch]$SkipE2EScript,
[switch]$Quiet
)
$ErrorActionPreference = 'Stop'
function Write-Step($msg) {
if (-not $Quiet) { Write-Host "==> $msg" -ForegroundColor Cyan }
}
function Write-Ok($msg) {
if (-not $Quiet) { Write-Host " OK $msg" -ForegroundColor Green }
}
function Write-Note($msg) {
if (-not $Quiet) { Write-Host " .. $msg" -ForegroundColor DarkGray }
}
function Fail($msg) {
Write-Host "FAIL: $msg" -ForegroundColor Red
exit 1
}
# Resolve repo root (this script lives in scripts/).
$RepoRoot = Split-Path -Parent $PSScriptRoot
Push-Location $RepoRoot
try {
# Sanity: bun on PATH.
Write-Step "Checking bun is available"
$bunVersion = & bun --version 2>$null
if ($LASTEXITCODE -ne 0 -or -not $bunVersion) {
Fail "bun is not on PATH. Install from https://bun.sh/ and re-run."
}
Write-Ok "bun $bunVersion"
# (a) Unit-test suite — covers each module's contract independently.
if (-not $SkipUnitTests) {
Write-Step "Running B1 unit tests"
$tests = @(
'src/tools/AgentTool/lazyDelegationCheck.test.ts',
'src/tools/AgentTool/invocationLimiter.test.ts',
'src/tools/AgentTool/specialistRouter.test.ts',
'src/tools/AgentTool/agentStallDetector.test.ts',
'src/tools/AgentTool/taskSpecQuality.test.ts',
'src/tools/AgentTool/workerContinueAdvisor.test.ts',
'src/tools/AgentTool/forkSubagent.test.ts',
'src/utils/modeAdvice.test.ts'
)
& bun test @tests
if ($LASTEXITCODE -ne 0) { Fail "B1 unit tests failed (see output above)" }
Write-Ok "B1 unit tests passed"
} else {
Write-Note "skipping unit tests (--SkipUnitTests)"
}
# (b) Inline e2e script — exercises the real exported modules in one run,
# no mocks. Each behavior is asserted; any failure exits non-zero with a
# descriptive message. Written to a temp file because Bun's -e flag
# doesn't support multi-statement TS.
if (-not $SkipE2EScript) {
Write-Step "Running inline B1 end-to-end smoke"
# Place the temp script INSIDE the repo so `import './src/...'` resolves.
$e2eDir = Join-Path $RepoRoot 'scripts'
$e2ePath = Join-Path $e2eDir ".coordinator-e2e-$([Guid]::NewGuid()).ts"
$e2eSource = @'
import {
detectLazyDelegation,
formatLazyDelegationError,
} from '../src/tools/AgentTool/lazyDelegationCheck.js'
import {
suggestSpecialist,
formatSpecialistRedirectMessage,
} from '../src/tools/AgentTool/specialistRouter.js'
import {
_resetLimiterState,
formatNearLimitWarning,
noteInvocation,
} from '../src/tools/AgentTool/invocationLimiter.js'
import { createStallDetector } from '../src/tools/AgentTool/agentStallDetector.js'
import {
analyzeFirstMessageForMode,
formatModeAdviceBanner,
} from '../src/utils/modeAdvice.js'
import {
assessTaskSpec,
formatThinSpecError,
} from '../src/tools/AgentTool/taskSpecQuality.js'
import {
findContinueCandidate,
formatContinueHintError,
type ContinueCandidateTask,
} from '../src/tools/AgentTool/workerContinueAdvisor.js'
import {
buildChildMessage,
COORDINATOR_RESEARCH_FORK_SUBAGENT_TYPE,
isCoordinatorResearchForkEnabled,
isInForkChild,
} from '../src/tools/AgentTool/forkSubagent.js'
function assert(cond: unknown, msg: string): asserts cond {
if (!cond) {
console.error(`ASSERT FAIL: ${msg}`)
process.exit(1)
}
}
// 1. Lazy delegation lint.
{
const lazy = detectLazyDelegation('Based on the findings, fix the auth bug.')
assert(lazy !== null, 'lazy delegation should be detected')
assert(/based on the findings/i.test(lazy!.phrase), 'phrase echoed')
const err = formatLazyDelegationError(lazy!, 'Agent')
assert(err.includes('Based on the findings'), 'error echoes phrase')
assert(err.includes('CLAUDE_CODE_LAZY_DELEGATION_CHECK'), 'error names env flag')
// negative case
const ok = detectLazyDelegation('Fix the null pointer in src/auth/validate.ts:42.')
assert(ok === null, 'specific prompts should not trip the lint')
console.log('1. lazy delegation: OK')
}
// 2. Specialist router redirect.
{
const types = new Set(['worker', 'code-reviewer', 'security-reviewer', 'debugger'])
const suggested = suggestSpecialist('please code review the auth module', types)
assert(suggested === 'code-reviewer', `expected code-reviewer, got ${suggested}`)
const msg = formatSpecialistRedirectMessage('code-reviewer', 'Agent')
assert(msg.includes('code-reviewer'), 'redirect names the specialist')
// negative case
const vague = suggestSpecialist('look at this code', types)
assert(vague === undefined, 'vague prompts should not redirect')
console.log('2. specialist router: OK')
}
// 3. Invocation limiter near-limit warning.
{
_resetLimiterState()
// verification cap = 5 by default. Drive 5 calls; the 5th must report
// nearLimit=true. The 6th must be capped.
const nearFlags: boolean[] = []
for (let i = 1; i <= 5; i++) {
const r = noteInvocation('verification')
nearFlags.push(r.nearLimit)
assert(r.capped === false, `call ${i} should not be capped`)
}
assert(
JSON.stringify(nearFlags) === '[false,false,false,false,true]',
`nearLimit should fire only on the 5th call, got ${JSON.stringify(nearFlags)}`,
)
const r6 = noteInvocation('verification')
assert(r6.capped === true, 'call 6 should be capped')
// formatter shape
const warning = formatNearLimitWarning('verification', {
count: 5,
limit: 5,
capped: false,
nearLimit: true,
})
assert(warning.includes('<system-reminder>'), 'warning is a system-reminder block')
assert(warning.includes('5 of 5'), 'warning shows current/cap')
_resetLimiterState()
console.log('3. invocation limiter near-limit: OK')
}
// 4. Stall detector transitions.
{
const T0 = 1_000_000
const d = createStallDetector({ thresholdMs: 90_000, startedAt: T0 })
let s = d.check(T0 + 1_000)
assert(s.kind === 'idle' && s.transitioned === false, 'idle before threshold')
s = d.check(T0 + 90_000)
assert(s.kind === 'stalled' && s.transitioned === true, 'first stalled check transitions')
s = d.check(T0 + 100_000)
assert(s.kind === 'stalled' && s.transitioned === false, 'subsequent stalled checks do not transition')
d.notify(T0 + 105_000)
s = d.check(T0 + 105_001)
assert(s.kind === 'resumed' && s.transitioned === true, 'resumed transition once after notify')
s = d.check(T0 + 110_000)
assert(s.kind === 'idle', 'returns to idle after resume')
console.log('4. stall detector: OK')
}
// 5. Mode advice.
{
const adv1 = analyzeFirstMessageForMode(
'Migrate the auth subsystem from express-session to JWT across the API server.',
)
assert(adv1?.suggestedMode === 'coordinator', `migration → coordinator, got ${adv1?.suggestedMode}`)
const adv2 = analyzeFirstMessageForMode('Fix the typo in README.md')
assert(adv2?.suggestedMode === 'normal', `typo → normal, got ${adv2?.suggestedMode}`)
// banner: should produce an actionable message when active mode mismatches advice
const banner = formatModeAdviceBanner('normal', adv1)
assert(banner !== null, 'banner should be returned on mismatch')
assert(/coordinator/i.test(banner!), 'banner mentions coordinator')
assert(
banner!.includes('CLAUDE_CODE_COORDINATOR_MODE'),
'banner references the real launch mechanism, not a fake slash command',
)
assert(!/\/coordinator\b/.test(banner!), 'banner must not reference a nonexistent /coordinator command')
// banner: null when active mode matches
assert(formatModeAdviceBanner('coordinator', adv1) === null, 'banner null on match')
console.log('5. mode advice: OK')
}
// 6. Task-spec quality (B2).
{
const good = assessTaskSpec(
'Fix the null pointer in src/auth/validate.ts:42. Add a guard before user.id, run the auth tests, and report the result.',
)
assert(good.quality === 'well-specified', `full brief should be well-specified, got ${good.quality}`)
const thin = assessTaskSpec('fix it')
assert(thin.quality === 'underspecified', `"fix it" should be underspecified, got ${thin.quality}`)
const err = formatThinSpecError(thin, 'Agent')
assert(err.includes('Agent'), 'thin-spec error names the tool')
assert(
err.includes('CLAUDE_CODE_COORDINATOR_TASK_SPEC_STRICT'),
'thin-spec error names the opt-in flag',
)
// A reasonable one-liner must NOT be flagged (false-positive guard).
const ok = assessTaskSpec('Run the test suite and report which tests fail.')
assert(ok.quality !== 'underspecified', `reasonable one-liner should not be underspecified, got ${ok.quality}`)
console.log('6. task-spec quality: OK')
}
// 7. Worker-continue advisor (B3).
{
const T0 = 1_700_000_000_000
const candidates: ContinueCandidateTask[] = [
{
agentId: 'agent-old',
agentType: 'worker',
description: 'auth investigation',
startTime: T0,
touchedFiles: ['src/auth/validate.ts', 'src/auth/session.ts'],
isCompleted: true,
},
{
agentId: 'agent-mismatch-type',
agentType: 'code-reviewer',
description: 'reviewed auth',
startTime: T0 + 1000,
touchedFiles: ['src/auth/validate.ts'],
isCompleted: true,
},
]
// Overlap on the right type → continue
const c = findContinueCandidate({
prompt: 'Fix the null pointer in src/auth/validate.ts:42',
subagentType: 'worker',
candidates,
options: { now: T0 + 5000 },
})
assert(c !== null, 'should recommend continuation when same-type worker overlaps')
assert(c?.agentId === 'agent-old', `expected agent-old, got ${c?.agentId}`)
assert(c?.sharedFiles.includes('src/auth/validate.ts'), 'shared file echoed')
// Type mismatch → no recommendation
const c2 = findContinueCandidate({
prompt: 'Fix src/auth/validate.ts',
subagentType: 'docs-writer',
candidates,
options: { now: T0 + 5000 },
})
assert(c2 === null, 'type mismatch should not recommend continuation')
// Prompt with no file refs → no recommendation
const c3 = findContinueCandidate({
prompt: 'fix the bug',
subagentType: 'worker',
candidates,
options: { now: T0 + 5000 },
})
assert(c3 === null, 'prompt without file refs should not recommend continuation')
// Error message contract
const err = formatContinueHintError(c!, 'Agent', 'SendMessage')
assert(err.includes('agent-old'), 'error contains agent id')
assert(err.includes('SendMessage'), 'error names SendMessage tool')
assert(err.includes('CLAUDE_CODE_COORDINATOR_CONTINUE_HINT'), 'error names env flag')
console.log('7. worker-continue advisor: OK')
}
// 8. Coordinator research fork (B4).
{
// Subagent type the coordinator spells.
assert(
COORDINATOR_RESEARCH_FORK_SUBAGENT_TYPE === 'fork',
`coord research fork type should be 'fork', got ${COORDINATOR_RESEARCH_FORK_SUBAGENT_TYPE}`,
)
// Default mode boilerplate has 10 rules, no research-only rule.
const normal = buildChildMessage('directive')
assert(normal.includes('10. REPORT'), 'default boilerplate has rule 10')
assert(!normal.includes('11.'), 'default boilerplate has no rule 11')
assert(!normal.includes('RESEARCH FORK'), 'default boilerplate has no RESEARCH FORK marker')
// Coordinator-research mode adds rule 11 about read-only investigation.
const research = buildChildMessage('investigate src/auth', 'coordinator-research')
assert(research.includes('11. RESEARCH FORK'), 'research mode adds rule 11')
assert(research.includes('Do NOT modify files'), 'research mode forbids modifications')
// Recursive-fork guard fires for both modes (boilerplate framing tag preserved).
const wrap = (text) => [{ type: 'user', message: { content: [{ type: 'text', text }] } }]
assert(isInForkChild(wrap(normal)), 'guard catches normal fork child')
assert(isInForkChild(wrap(research)), 'guard catches coord-research fork child')
assert(!isInForkChild(wrap('plain user prompt with no fork tag')), 'guard ignores plain user text')
// Env flag gate: requires both coord mode AND the opt-in flag.
const prevCoord = process.env.CLAUDE_CODE_COORDINATOR_MODE
const prevFlag = process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK
try {
delete process.env.CLAUDE_CODE_COORDINATOR_MODE
process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK = '1'
assert(!isCoordinatorResearchForkEnabled(), 'gate off without coord mode')
process.env.CLAUDE_CODE_COORDINATOR_MODE = '1'
delete process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK
assert(!isCoordinatorResearchForkEnabled(), 'gate off without env flag')
process.env.CLAUDE_CODE_COORDINATOR_MODE = '1'
process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK = 'true'
assert(!isCoordinatorResearchForkEnabled(), 'gate requires exact "1"')
} finally {
if (prevCoord === undefined) delete process.env.CLAUDE_CODE_COORDINATOR_MODE
else process.env.CLAUDE_CODE_COORDINATOR_MODE = prevCoord
if (prevFlag === undefined) delete process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK
else process.env.CLAUDE_CODE_COORDINATOR_RESEARCH_FORK = prevFlag
}
console.log('8. coordinator research fork: OK')
}
console.log('\nB1 end-to-end smoke: ALL PASS')
'@
Set-Content -LiteralPath $e2ePath -Value $e2eSource -Encoding UTF8
try {
& bun run $e2ePath
if ($LASTEXITCODE -ne 0) { Fail "B1 e2e smoke failed (see output above)" }
Write-Ok "B1 e2e smoke passed"
} finally {
Remove-Item -LiteralPath $e2ePath -Force -ErrorAction SilentlyContinue
}
} else {
Write-Note "skipping inline e2e script (--SkipE2EScript)"
}
if (-not $Quiet) {
Write-Host ""
Write-Host "B1 coordinator behavior: GREEN" -ForegroundColor Green
}
} finally {
Pop-Location
}