A Claude Code plugin that runs a three-agent development pipeline against a GitHub project board. Items flow from triage → development → review without you touching them. Part of the claude-workbench marketplace.
You work out of a GitHub project board. New issues land with no acceptance criteria. Well-refined ones sit in Ready. Work-in-progress has open draft PRs. PRs waiting on review pile up.
This plugin runs a team of agents on a local 20-minute clock that move items through the pipeline for you — and records what the reviews teach as it goes, so the team stops repeating itself:
- Inspector Lestrade (Sonnet) — triage. Reads items in the
Inboxlane, writes acceptance criteria as a managed follow-up comment on the issue (the description is left untouched). Before scoring, the draft AC is checked by four blind lens sub-agents (malicious-compliance, testability, completeness, edge-case) that each try to find a gap Holmes and Watson would otherwise hit downstream — a much more expensive place to catch it, since Lestrade only ever sees an issue once while Watson and Holmes cycle back on it every bounce. Real gaps get folded in with one bounded tightening pass, never a re-verification loop; a lightweight comment marks when the AC changed this way. Scores WSJF, moves them toBacklogfor your review. The WSJF write also lands two GitHub-native issue attributes The Index derives server-side: the issue Type (PBI) and an issue-level Priority (Urgent/High/Medium/Low) mapped from the WSJF — org-repo-only and best-effort. Also runs blocker + consolidation sweeps: after a repo gets fresh triage work, he re-reads all of its open issues and (1) marks blocked-by dependencies (native GitHub issue dependencies, additive only) so blocked items stay out of Dr. Watson's queue, and (2) consolidates follow-ups — foldsexpand-fromcomments into the issue they target and merges unmistakable near-duplicate follow-ups into the earliest anchor (native duplicate-close, high bar, ambiguous clusters flagged not closed) so the backlog stops sprawling. - Dr. Watson (Opus, $10/run cap) — development. Has two modes: The Index mode (Dispatch-driven, picks the top
Ready/In Progressitem, clones the repo, writes code and tests against AC, opens a PR, moves toIn Review) and Direct mode (invocable as a sub-agent from Claude Code or Cowork for ad-hoc dev work — no The Index calls, just runs the/developskill in a sub-agent context). Both modes follow the/developskill for the actual coding. - Sherlock Holmes (Opus parent, Sonnet lenses, $7/run cap) — code review. Reviews open PRs, approves or requests changes. Escalates to you after 3 change rounds — but your input resets that count: comment, review, or weigh in on the PR and the window restarts from your last word, so an escalated PR you've decided on gets a fresh review instead of bouncing straight back. Reviews fan out across blind, read-only lens sub-agents (AC conformance, correctness, security, test honesty), with every blocker adversarially verified before it lands — a 3-agent red-team/blue-team/auditor pipeline handles security-lens findings every round and every other finding on the PR's first review, a single skeptic handles the rest on a re-review (the fullest check lands where it can prevent a second round, not just where the finding is scariest). After verification, Holmes (and only Holmes — no sub-agent touches the vault) checks each surviving finding against the memory vault for relevant context — a documented decision that reframes it, a past incident that reinforces it — always re-verified against the current tree before it's trusted, never used to waive a real defect or mark an AC item met. Only the parent writes, so there's still exactly one App-signed verdict. Falls back to a single inline pass when the fan-out is unavailable. Findings route by the coherent unit of work — what the issue is really about — then coupling and locality: a finding that belongs to the unit blocks and is fixed in this PR, even in untouched code the diff never caused, because a half-delivered unit is itself the defect. On top of that, anything actionable in the code a PR touched blocks (request changes, however minor), a hard correctness/security/test defect blocks wherever it lives, and — coupling beating locality — untouched code the diff made stale, inconsistent, or wrong blocks too. The self-test: block and fix here if EITHER the diff caused it OR it belongs to the coherent unit — a follow-up only when both are false. The precise routing rule lives in Holmes's canonical review contract (
agents/holmes.md, §4e/§5). The non-blocking follow-up tier is then gated by materiality, default-deny: it holds only findings unrelated to the unit, and most of those — one-off cosmetics (naming, small duplication, style) — are noted in the verdict, not tracked. Only an unrelated latent hazard (security/data-integrity/correctness not live enough to block) or systemic/substantial debt (a schedulable chunk with its own testable "done") earns one tracked issue, taggedTracked under:, capped at one new anchor per PR — the materiality bar that stops the follow-up flood. When Holmes does track, he expands the earliest related open issue in place (a comment Lestrade folds into its acceptance criteria) rather than opening a near-duplicate, and opens a new anchor viacreate_issue(App-signed as Holmes, board-added andPBI-typed) only when nothing related exists. A class of sites violating one invariant (a containment guard, a null-check, a helper every caller owes) is swept whole: if the class belongs to the unit it's folded into the PR; if it's an unrelated anti-pattern agents will replicate it becomes one umbrella issue with a checkbox per site — either way closing the class at once rather than minting a fresh single-site issue every review, the treadmill that otherwise turns one finding into an endless#A → #B → #Cchain. A finding gets the same disposition regardless of verdict, so a clean PR never generates more tracked work than a messy one: on a change request, unit-belonging findings are blockers Watson folds into the same PR, unrelated cosmetics are optional (fix if cheap, else skip), and the unrelated hazard/debt tier is tracked exactly as on approval.
A fourth component — Dispatch — is the local scheduled task that polls the board every 20 minutes and fires the right agent for each pending item. Dispatch is the only thing that's scheduled; the three agents run as dispatched subprocesses.
Lestrade, Watson, and Holmes used to have no memory of their own reviews: nothing recorded what Holmes rejected, why, or how it got fixed, so the same lessons got re-taught every review (test-honesty is ~38% of all rejections, fail-open ~11%, doc-drift ~11%). There's no separate harvesting agent for this — Holmes is the only one who holds both halves of a rejection (what he flagged, and whether the next push actually fixed it), so he records it himself, live, at re-review: one atomic vault note per bounce or AC-dispute event, categorized against a fixed taxonomy, plus an incrementally-refreshed dev-team/top-lessons.md digest (recurring categories, frequency-ranked, each with the concrete prevention rule). Watson reads that digest and searches the vault for anything task-specific before coding; Lestrade reads it before writing acceptance criteria, so the pipeline gets smarter instead of repeating itself on both sides — what gets built and what gets asked for.
/plugin marketplace add mike-bronner/claude-workbench
/plugin install workbench-dev-team@claude-workbench
That installs the agents, the Dispatch prompt, and the bundled skills (see below). Nothing is scheduled yet.
The plugin ships three skills for general use, plus the agents' own reference skills:
developandgit-commit— universal development standards. They register themselves globally viasession-warmup.md, which workbench-core picks up at session start and injects into~/.claude/CLAUDE.md. They apply to every Claude Code / Cowork session, not just dev-team agents. Both are also packageable as.skillfiles for Claude Chat (Mac app) where plugins aren't supported but skills are. Require workbench-core 0.2.0+ for the session-warmup discovery mechanism — install it first if you don't already have it (Claude Code does not enforce plugin install order).orchestrate— runs the team as background sub-agents from any interactive session (see below). Asession-warmup.mdhint makes every session aware the team is available for delegation.
Four more skills exist for the agents rather than for you. comms-style is how Lestrade, Watson, and Holmes write every piece of prose that isn't code — ticket comments, PR bodies, review verdicts — modeled on ASD-STE100 (Simplified Technical English). holmes-review, watson-pipeline, and lestrade-triage hold each agent's long, situational procedure: each agent prompt in agents/ is a thin router that keeps the always-relevant rules inline and points at skills/<name>/references/ for the detail it only needs at one moment — Holmes's review phases and sub-agent prompt skeletons, Watson's eleven-step Index-mode pipeline, Lestrade's acceptance-criteria lenses and Sweep mode. The router pattern mirrors git-commit, and it keeps the per-dispatch prompt small without putting any rule out of reach.
Plugin configuration lives in a slash command (/workbench-dev-team:setup), not a skill — see the Setup section below.
Universal dev workflow + standards: orient before writing, plan before coding, atomic commits, every change gets a test, no committed secrets, lint before pushing. Triggers whenever code is being implemented, fixed, refactored, or tested — manual or agent-driven.
Includes a decision protocol that requires presenting three options to the human (with reasoning and a recommendation) for any meaningful fork — implementation approach, library choice, scope decisions, naming. The human decides, the agent executes. Trivial choices (mechanical translation, following existing repo conventions, one-line obvious fixes) are exempt.
Also defines the commit approval gate (see Commit approval gate below): no git commit without explicit human approval of the diff and message.
Used by Watson internally in both operating modes. Also invocable directly in any plugin-aware Claude session.
Generates commit messages using Conventional Commits + Gitmoji format. Triggers whenever a commit message is being composed — manual, scripted, or agent-driven (including Watson's PRs).
Format example:
feat: ✨ Add email validation endpoint.
Fixes: #789
Full type and emoji references at skills/git-commit/references/.
Turns the current session into the team's orchestrator: dispatches Lestrade, Watson, and Holmes as background sub-agents (Agent tool, run_in_background), passing each agent's model from the shared config; maintains a roster table of who's working on what; relays verdicts and decision forks back to you; follows up on running agents via SendMessage. The main conversation stays lean — sub-agents do the heavy work in their own contexts and return summaries.
Watson supports prose-driven Direct mode for ad-hoc dev work with no board item. Lestrade and Holmes are Index-coupled and need a board item ID.
The skill also routes GitHub actions to the right executor. Two rules: (1) agent work products (formal reviews, AC, status moves) only ever go through The Index, signed as the dispatched agent — never gh; (2) your own actions (comments you dictate, merges you order) go through gh under your identity, on any repo. Whether a repo is Index-governed is answered by check_repo_access — a server-side tool that checks The Index GitHub App's installation list (spec in THE_INDEX_HANDOFF_ROUTING.md; until it ships, the skill degrades to a list_items scan and says so). Merges are never delegated to agents and only happen on your explicit request.
Every git commit requires explicit human approval. Non-negotiable. Enforced twice, because prose alone drifts:
- Prose — the
developskill instructs the agent to present the staged diff and the proposed commit message, then wait for an explicit yes before committing. One approval covers one commit. - Machinery — a plugin
PreToolUsehook (hooks/hooks.json → hooks/scripts/commit-approval-gate.sh) detectsgit commitin any Bash invocation (compound commands,-C/-cflags, env prefixes included) and returnspermissionDecision: "ask", forcing the harness to prompt — even in auto-accept permission modes, even if the model forgot the prose.
Pipeline carve-out. Headless ask prompts auto-deny, so an ungated rule would deadlock every scheduled Watson run at its first commit. The hook therefore allows commits through when the autonomous Index pipeline is demonstrably running: a live-PID /tmp/watson.lock (which Watson's Index mode acquires before any work) or WORKBENCH_DEV_TEAM_PIPELINE=1 in the environment. In the pipeline, board dispatch is the approval and Holmes review + your PR merge is the human gate. A stale lock (dead PID) does not bypass.
Known edge: while a scheduled Watson run is live, its lock also exempts concurrent interactive sessions on the same machine. The prose protocol still applies there; the window is the few minutes of a pipeline tick.
Tests: hooks/scripts/test-commit-approval-gate.sh (19 cases — detection, non-commit silence, carve-outs).
Per-agent model, effort, fallback chain, and budget caps live in a single file, written by /workbench-dev-team:setup with these defaults and never overwritten on re-run (your edits survive plugin updates):
// ~/.claude-workbench/dev-team-config.json
{
"agents": {
"lestrade": { "model": "sonnet", "effort": "high", "fallback": "haiku" },
"holmes": { "model": "opus", "effort": "xhigh", "fanout": true, "lensModel": "sonnet", "maxBudgetUsd": 7.00, "fallback": "sonnet" },
"watson": { "model": "opus", "effort": "xhigh", "maxBudgetUsd": 10.00, "fallback": "sonnet,haiku" }
}
}Both dispatch paths read it:
- Scheduled (Dispatch) passes
--model,--effort(only when set),--fallback-model(only when set), and--max-budget-usdon eachclaude -pinvocation. CLI flags override agent frontmatter (verified empirically), so a config edit takes effect on the next tick — no plugin files to touch. - Interactive (
orchestrate) passes the config'smodelas the Agent tool's per-invocation model override. The Agent tool has no per-invocation effort, budget, or fallback parameter — interactive sub-agents inherit the session's effort level, andmaxBudgetUsd/fallbackapply to the scheduled path only (a model error there surfaces immediately for you to handle).
Holmes also carries two optional review knobs: fanout (bool, default true) toggles its multi-lens review fan-out, and lensModel (default: Holmes's own model) sets the model its lens and skeptic sub-agents run on. Both default cleanly when absent. The optional fallback knob (any agent) is a comma-separated model list handed to --fallback-model, so a dispatch degrades to the next model when the primary is overloaded or unavailable — e.g. a retired model — instead of failing; maxBudgetUsd caps per-run spend (Watson defaults to 10.00; Holmes's is optional). All default cleanly when absent.
The agent definitions carry matching frontmatter defaults (model: sonnet|opus), so direct Agent-tool dispatch without the skill still lands on the right model. effort is deliberately not in frontmatter: frontmatter effort would override the session level — including Dispatch's --effort flag — turning the config knob into a no-op. Missing file, missing key, or malformed JSON all fall back to the defaults above; dispatch never blocks on config problems.
After /plugin install, run this in any Claude Code session:
/workbench-dev-team:setup
It walks you through cadence selection (20 or 30 min), Keychain seeding for any missing credentials, The Index MCP registration, log directory and agent-config creation, and Dispatch scheduled-task registration. Idempotent — re-run any time to refresh the OAuth token, re-register the MCP, or change cadence. Your dev-team-config.json edits are never overwritten.
Prerequisites on your machine: gh (authenticated), jq, security (built into macOS).
You'll be prompted in chat for any of these Keychain entries that aren't already present:
| Entry | Purpose |
|---|---|
the-index-mcp / client-id |
The Index OAuth client ID |
the-index-mcp / client-secret |
The Index OAuth client secret |
github-cli / token |
GitHub token for dispatched agents (auto-extracted from your existing gh auth login Keychain entry when present) |
claude-code / oauth-token |
Claude Code OAuth token for scheduled claude -p invocations. Get one with claude setup-token |
- Verifies prerequisites (
gh,jq,security) and Keychain credentials; prompts for anything missing. - Fetches an OAuth bearer token from The Index (client_credentials grant, 1-year lifetime).
- Registers The Index MCP with Claude Code at user scope, passing the bearer via
--header. This makesmcp__the-index__*tools available to every future Claude Code session, including the dispatched agents. - Creates the log directory at
~/.claude-workbench/dev-team-logs/and writes the default agent config to~/.claude-workbench/dev-team-config.jsonif (and only if) it doesn't already exist. - Registers the scheduled Dispatch task by calling
mcp__scheduled-tasks__create_scheduled_task(orupdate_scheduled_taskif it already exists) directly from the running session. Task ID:workbench-dev-team-dispatch. Cron:*/20 * * * *(or*/30if you chose 30 min).
Re-run the skill any time you need to refresh the OAuth token, re-register the MCP, or change the Dispatch cadence. Also re-run it after a plugin update that changes Dispatch's flow (a change to scheduled-tasks/orchestrator.md itself, not the agent contracts): step 5 reads that file, strips its frontmatter, and passes the body as the scheduled task's prompt, so the deployed task holds a baked-in copy. A plugin update refreshes the file on disk but not the running task — the re-run redeploys the fresh orchestrator body. (Agent definitions and skills are read live per dispatch, so those need no re-run — only the scheduled Dispatch prompt is baked in.)
GitHub webhook ──► The Index (MCP server, OAuth 2.1)
▲
│ MCP tool calls
│
scheduled Dispatch task (every 20 min, local)
│
▼ nohup claude -p --agent ... &
┌───────────┬───────────┬───────────┐
▼ ▼ ▼
Lestrade Holmes Watson
(Sonnet) (Opus,$7) (Opus,$10)
▲ │ │
│ ▼ writes ▼ reads/searches
└──── memory vault (dev-team/top-lessons.md +
review-learnings notes) ────┘
Every 20 minutes, the Dispatch scheduled task wakes up and:
- Calls
mcp__the-index__list_unrefined_items()→ fires Lestrade for each item returned, plus one blocker sweep (Repo sweep: <owner/repo>) per distinct repo among those items. - Calls
mcp__the-index__list_review_items()→ fires Holmes for each item returned. - Calls
mcp__the-index__list_development_items(limit=1)→ fires Watson on the top item (if any). - Exits.
Each dispatch is fire-and-forget via nohup claude -p --agent workbench-dev-team:<name> ... &; disown. Watson alone can run for hours; Dispatch never blocks.
All "what's pending in each lane" logic lives server-side in The Index's MCP tools. Dispatch never interprets item status, field changes, or priority — it just asks The Index "what's pending in each lane?" and fires the matching agent per returned item. Adding a new dispatch rule means editing The Index, not this plugin.
- Inspector Lestrade and Sherlock Holmes are idempotent within a tick. Status lanes (
nullandIn Review) act as the serialization. - Dr. Watson picks from
In ProgressORReady(In Progress first — that's the resume path for crashed runs). A host-local PID mutex at/tmp/watson.lockprevents two Watsons from stepping on the same item. Released explicitly (rm -f /tmp/watson.lock) in cleanup and on every early exit — never via a shelltrap, which would delete the lock the moment the tool-call shell returns.
| Scenario | Tokens |
|---|---|
| Idle Dispatch tick (no work in any lane) | ~1–3K on default model — three MCP calls + exit |
| Lestrade triage | 5–8K Sonnet tokens + four blind Sonnet lens sub-agents verifying the draft AC |
| Lestrade blocker sweep | Sonnet tokens scaling with open-issue count (reads every open title + body in the repo); fires only on ticks that triaged new items |
| Holmes review | Opus parent + four blind Sonnet lens sub-agents + adversarial skeptic per review (security findings, and every finding on the PR's first review: 3-agent red/blue/auditor panel instead — up to 30 verification dispatches on a busy first review vs. 10 on a re-review); capped at $7 per run |
| Watson development | Full Opus session, capped at $10 per run |
Dispatch runs on your Claude Code default model (scheduled tasks don't expose a model selector). Since each tick is under 3K tokens, the default model's cost is negligible even if it's Sonnet.
Two ways to invoke the same agents, same definitions:
- Unattended (default). The scheduled Dispatch task polls The Index every 20 minutes and dispatches via
claude -p --agent. This is what/workbench-dev-team:setupregisters. - Interactive. Any Claude Code session can dispatch an agent directly via the Agent tool, e.g.,
Agent(subagent_type: "workbench-dev-team:lestrade", ...). For multi-agent delegation with config-driven models, background execution, and roster tracking, use theorchestrateskill — it wraps this path with the full protocol. Useful for manual triage, one-off runs, ad-hoc dev work (Watson Direct mode), or debugging without waiting for the next scheduled tick.
- Agent logs.
~/.claude-workbench/dev-team-logs/<agent>-<item>-<timestamp>.log— full agent output per dispatch. - Review-learnings notes.
dev-team/review-learnings/<repo>-pr<n>-<date>.md(one note per bounce/escalation, written by Holmes at re-review) anddev-team/top-lessons.md(the frequency-ranked digest Watson and Lestrade read) in your memory vault. Nothing there yet means no PR has bounced or been AC-disputed since this shipped. - Scheduled task panel. Claude Code's scheduled-tasks panel shows the Dispatch task's run history and next-run time.
- Project board. Items flow Inbox → Backlog → Ready → In Progress → In Review → Approved / Escalated. Status drift (items stuck in a column) is your canary.
| Problem | Fix |
|---|---|
claude mcp list shows the-index as Failed to connect |
Your OAuth token has probably expired or been revoked. Re-run /workbench-dev-team:setup to fetch a fresh token and re-register. |
Dispatch logs the-index unreachable |
curl https://the-index.mikebronner.dev/mcp to check the endpoint. If 500, The Index has a middleware bug (should be 401). |
Watson stuck (process hung, /tmp/watson.lock stale) |
rm /tmp/watson.lock. Next tick will resume. |
| Agent not found | claude agents should list workbench-dev-team:lestrade, :watson, :holmes. If not, reinstall the plugin. |
Item stuck in In Review with no PR |
Holmes couldn't find a PR for the issue. Check gh pr list -R <repo> --search <issue>. |
| Scheduled task isn't firing | Check Claude Code's scheduled-tasks panel. The Mac must be awake (this is a local scheduler). |
- Local execution. Dispatch runs on your Mac. If the host is off, no work moves. Fine for home/dev setups; move Dispatch to an always-on box if you need 24/7 coverage.
- Budget caps.
--max-budget-usd 10.00limits Watson's per-run spend; Holmes carries an optional cap too (default7.00), since its lens fan-out is the only uncapped, multi-agent lane. Complex work may hit the ceiling and leave the item inIn Progress; the next tick resumes. (The $10 figure was originally sized for Fable's 2× Opus pricing; on Opus it now buys roughly twice the tokens per run.) - The Index must be reachable. If the MCP server is down, all three list tools fail and Dispatch logs
the-index unreachableand exits cleanly. The next tick retries. - OAuth token lifetime. The Index issues 1-year tokens via client_credentials. Re-run
/workbench-dev-team:setupannually (or whenever you rotate the OAuth client secret).
If /workbench-dev-team:setup fails at step 5 (scheduled-task registration), choose "Skip" when re-prompted to register the schedule, then register manually from any Claude Code session:
mcp__scheduled-tasks__create_scheduled_task with
taskId: "workbench-dev-team-dispatch"
cronExpression: "*/20 * * * *" # or */30 for 30-min cadence
description: "Dispatch — poll The Index every 20 min and fire workbench-dev-team agents on pending items."
prompt: <body of scheduled-tasks/orchestrator.md, frontmatter stripped>
An earlier iteration targeted Anthropic's cloud-hosted routines (/fire endpoint, configured via /schedule → RemoteTrigger) for event-driven dispatch. Two things killed it:
- The 15-routine-runs-per-day cap on online routines. Dispatch every 20 minutes = 72 fires/day, seven times over.
- Added operational surface. Fire-token storage, a The Index-side webhook dispatcher, per-transition idempotency — a lot of moving parts to event-drive what a 20-minute poll handles just as well.
A local 20-minute poll has higher worst-case latency but zero per-fire cost, simpler failure modes, and no token rotation burden. Given that Watson can run for hours, 20-minute dispatch latency is noise.