Skip to content

Latest commit

 

History

History
113 lines (70 loc) · 61.5 KB

File metadata and controls

113 lines (70 loc) · 61.5 KB

Changelog

All notable user-visible changes to CASCADE are documented here. The format is loosely based on Keep a Changelog; the project does not use strict semver for releases.

Unreleased

Added

  • Dashboard: build a project's worker image from a Dockerfile, with a source selector, live build status, and Rebuild (MNG-1725, spec 023 plan 5 of 5 — final). The Worker Image card in Project → Settings → General now lets a superadmin choose the image source — Global default, Referenced image (the spec-022 control, unchanged), or Dockerfile. For the Dockerfile source, paste only the extra layers (RUN / COPY / ENV …) into a textarea and Set (calls projects.update({workerDockerfile})) — CASCADE supplies the pinned FROM cascade-worker base and builds the image router-side. The two override sources are mutually exclusive: selecting one hides the other's control, matching the backend invariant. The status display separates the active image (workerImageStatus: pending / building / verified / failed) from the most recent build attempt (workerImageBuildStatus: building / failed), so a project running its last-good image while a rebuild fails reads "Verified … · last rebuild failed: <reason>" rather than a misleading "Failed"; a Building… spinner shows for a first build and the card polls (WORKER_IMAGE_POLL_MS) while workerImageStatus === 'building' || workerImageBuildStatus === 'building'. A Rebuild button (Dockerfile source only) calls projects.rebuildWorkerImage to re-run the build against a refreshed base without editing the content. The whole card stays hidden for non-superadmins. Operator docs (README.md, docs/getting-started.md) walk through writing extra layers → save → watch build/verify (or read a failure) → rebuild, and call out the mutual exclusivity and the single-daemon constraint (a Dockerfile-built image is local to the router that built it). Completes the worker-Dockerfile feature end-to-end across schema, spawn resolution, build engine, set surfaces, and dashboard. Closes MNG-1725.

  • Dashboard: set/clear a project's worker image with live verified/pending/failed status (MNG-1699, spec 022 plan 4 of 4 — final). A superadmin can now manage a project's per-project worker image from the dashboard: a new Worker Image card in Project → Settings → General shows the global default as the input placeholder, accepts a reference (Set), reverts to the global default (Clear), and reflects the router-side validation lifecycle inline — a Verifying… spinner that polls while pending (same approach as the run-status pages), a Verified — pinned to @sha256:… badge once the digest is resolved, or a Validation failed: <reason> badge naming the missing requirement. The control is wired to the existing projects.update mutation (set sends workerImage, clear sends null) and is hidden entirely for non-superadmins, mirroring the backend gate. Completes the feature end-to-end across CLI, API, and dashboard; the operator walkthrough in docs/getting-started.md covers deriving a custom image FROM the Cascade worker base, making it available in both the registry-backed and self-hosted/local topologies, setting it from the dashboard, and confirming verified. Closes MNG-1699.

  • Configure a per-project .cascade/setup.sh wall timeout (CLI + dashboard) (MNG-1701). A new nullable setup_timeout_ms column on projects is a per-project override for the wall-clock timeout of the repository setup script, plumbed through the proven maxInFlightItems / snapshotTtlMs / workerImage precedent (migration → Drizzle schema → ProjectConfigSchema → configMapper → projectsRepository → tRPC projects.create/update → CLI → dashboard form). Semantics: null/unset or 0 means no per-project wall timeout (the global worker/watchdog container timeout remains the safety net); a positive value is passed as wallTimeoutMs to the setup.sh runCommand call in setupRepository (src/agents/shared/repository.ts). The idle timeout stays disabled (idleTimeoutMs: 0) regardless — the field only controls the wall timeout. The Zod config schema and both tRPC inputs use .nonnegative() (not .positive()) and the CLI/dashboard use truthiness-safe guards so an explicit 0 (disable) is a first-class transmitted value rather than being dropped. Set it with cascade projects update <id> --setup-timeout-ms 1800000 (or --setup-timeout-ms 0 to disable), inspect it via cascade projects show <id> (Setup Timeout (ms)), or use the "Setup script timeout (ms)" field in the project Settings tab. This also makes the unset/0 default explicitly disable the setup.sh wall timeout (folding in PR #1463's behavior — npm ci on large React Native repos can take 10-15+ min) while a positive value re-introduces a per-project bound. Closes MNG-1701.

  • Configure a per-project worker image (CLI + API): superadmin-set, digest-pinned, validated fail-closed by a router-side smoke-test (MNG-1698, spec 022 plan 3 of 4). A superadmin can now pin a per-project worker image via cascade projects update <id> --worker-image <ref> and the projects.update / projects.create tRPC mutations (non-superadmins get FORBIDDEN). Malformed references are rejected synchronously with BAD_REQUEST (nothing persisted); a valid reference is stored as pending and an eager router-side validation job is enqueued on the dashboard-jobs queue (the Docker socket is router-only). The router handler pulls the image, pins its immutable @sha256: digest from RepoDigests, and runs the extended cascade-compatible-worker-image smoke-test (cascade-tools, node, git, an engine CLI, plus the python shim and Playwright — the same checks tests/docker/worker-runtime-tools/run-test.sh asserts) inside a one-shot docker run --rm. On success the project is marked verified with the digest plan 2 launches; any failure (unpullable image, missing digest, missing tool) marks it failed with a precise reason and never leaves it stuck in pending — fail-closed, so a bad image can never launch. cascade projects update <id> --clear-worker-image reverts to the global default (all four columns nulled, no validation). cascade projects show <id> renders the reference and its lifecycle (pending / verified → <digest> / failed: <reason>), and projects.defaults exposes the global WORKER_IMAGE. Every set/clear emits a structured, grep-stable project_worker_image_changed audit log line (actor + project + from→to). Closes MNG-1698.

  • The per-agent updateChannel setting is now documented (MNG-1687). Each agent type has an optional updateChannel (none / scm-only / pm-only / both, default both) that gates where it posts communication-only status updates — independently for the PM (work-item comments) and SCM (pull-request comments / reviews) surfaces. The catalog, resolver, and posting matrix live in src/config/updateChannel.ts; the per-agent value is stored in the agent_configs.update_channel column (NULL / absent / unrecognized → inherit the default both) and read at runtime via resolveUpdateChannel(project, agentType). The channel silences acks, progress updates, lifecycle status comments, agent summaries / reviews, and the agent's own PM/SCM posting tools (PostComment; PostPRComment, UpdatePRComment, CreatePRReview, ReplyToReviewComment) in both the native-tool and LLMist engine paths — but never PR creation, status moves, label writes, checklist sync, PR linking, friction reports, or the "eyes" reaction. Documented in CLAUDE.md / AGENTS.md (kept byte-identical), docs/architecture/04-agent-system.md (gated-vs-not-gated surface map), and docs/architecture/08-config-credentials.md (per-agent config field). Closes MNG-1687.

  • cascade-tools now suggests the closest command when an agent typos a topic or subcommand (MNG-1442). bin/cascade-tools.js registers an oclif command_not_found hook that turns command typos into the same structured spec-014 envelope every other CLI failure emits: JSON on stdout, a one-line prose summary on stderr, and a runnable did you mean hint when within the shared Levenshtein budget (MNG-1440). Unknown top-level topics (cascade-tools sm get-pr-diff) surface a topic enumeration in expected and a hint that preserves the user's trailing segments (did you mean 'cascade-tools scm get-pr-diff'?). Known-topic / unknown-subcommand typos (cascade-tools pm reaad-work-item) surface the topic's subcommand enumeration and a corrected hint (did you mean 'cascade-tools pm read-work-item'?). Far-away typos drop hint but still surface expected so the agent has a concrete recovery path. Exit code is 2 for unknown-command — preserved from oclif's historical command_not_found default and distinct from every other envelope's exit code 1; existing exit-code consumers see no change. The hook lives at src/cli/_shared/command-not-found-hook.ts, intentionally inside _shared/ so oclif's command-discovery glob in bin/cascade-tools.js excludes it. It is wired through pjson.oclif.hooks so oclif loads it dynamically only when needed — no static import is added, which preserves the existing friendly dist/cli/bootstrap.js missing path in the entrypoint. The pure suggestion logic lives in src/cli/_shared/commandSuggestions.ts (MNG-1441) and is unit-tested directly without booting oclif; candidates come strictly from the loaded oclif config (config.commandIDs plus non-hidden pjson.oclif.topics), so the cascade-tools binary never suggests dashboard topics that its discovery glob excludes. Existing unknown-flag handling from createCLICommand() is untouched. Closes MNG-1442.

Fixed

  • cascade-tools multiline text and large diff I/O are now hardened against shell-quoting footguns and stdout truncation (MNG-1059). The shared CLI factory at src/gadgets/shared/cli/params.ts now rejects invocations that pass --*-file - for two or more file-input flags (e.g. --body-file - --comments-file -) before any readFileSync(0, ...) call — stdin (fd 0) can only be drained once per process, and the previous behavior silently truncated one of the two agent payloads. The rejection emits a structured flag-parse error envelope (error.flag: "body-file,comments-file", hint: "Pass at most one --*-file -; for the others, write the payload to a temp file and pass --<flag>-file <path>.") so agents can self-correct on the next attempt. Direct file paths remain pairwise-compatible — --body-file - --comments-file /tmp/comments.json and --body-file /tmp/body.md --comments-file - both work as before. The native-tool system prompt now renders a "cascade-tools shell-safety rules" section that documents the one-stdin-consumer invariant and provides safe heredoc / temp-file patterns for one and two payloads. The prompt renderer also suppresses inline --body '...' / --text '...' examples whose content contains backticks, code fences, $(...), or newlines when a file-input companion is declared, redirecting the agent at the safer --*-file <path> form instead. File-input flag descriptions for --body-file, --text-file, --description-file, --details-file, and --comments-file explicitly call out markdown / multiline / backticks. Closes MNG-908, MNG-910, MNG-917, MNG-1046.

  • cascade-tools scm get-pr-diff gains an --outputFile <path> escape hatch for large diffs and one-line JSON patches (MNG-1045). When --outputFile is set, the full multiline Markdown diff is written to the requested path on disk and stdout returns a compact JSON summary {outputFile, fileCount, bytes, pathFilter} instead of the raw payload — sidestepping terminal-truncation issues with hundreds-of-kilobytes one-line JSON diffs. Default behavior is preserved: without --outputFile, get-pr-diff returns the formatted Markdown directly. The review-agent skipped-files guidance now points operators at this form (cascade-tools scm get-pr-diff --prNumber <N> --path <path> --outputFile /tmp/pr-diff.md) when a file would otherwise truncate. The --outputFile flag is declared as cliOnly: true on the underlying ToolDefinition so it appears in the CLI + agent-facing manifest but is excluded from the SDK Gadget Zod schema (gadgets return strings in-process and cannot deliver a file path back through that contract).

Changed

  • Worker spawn now resolves a verified per-project worker image and pins it to a digest (MNG-1697, spec 022 plan 2 of 4). resolveSpawnSettings (src/router/worker-spawn-settings.ts) introduces an explicit SpawnSettings.effectiveBaseImage — the project's verified per-project image digest when configured, otherwise the global routerConfig.workerImage. With no per-project image this is a pure no-op refactor: effectiveBaseImage === routerConfig.workerImage and resolved image, snapshot behavior, and pull-fallback are identical to before. When a project sets workerImage with workerImageStatus: 'verified' and a non-empty workerImageDigest, the worker launches from that digest (the Resolved spawn settings log now records projectWorkerImage, globalWorkerImage, and effectiveBaseImage so a post-mortem can confirm which image won). A configured-but-unverified image (pending / failed / missing digest) throws a terminal WorkerImageResolutionError and never silently falls back to the global default; the dispatch-error classifier maps that error to terminal so BullMQ skips the retry budget. container-manager.ts now classifies pull-fallback, snapshot reuse, and the snapshot-404 fallback against effectiveBaseImage instead of the global default — so a per-project base image is pulled-on-missing exactly like the global base, a custom-image run without a snapshot is not misclassified as a snapshot reuse, and a stale snapshot 404 falls back to the verified digest (not the global image). A missing-but-pullable custom digest is pulled once and retried; an unobtainable/invalid custom digest fails loud with a grep-stable terminal error instead of relaunching on the global default. The container launch posture is unchanged (only Memory / MemorySwap / NetworkMode / AutoRemove + labels — no mounts, no privileged). Dormant for operators until the set/validation flow lands (plan 3); exercised here via seeded ProjectConfig.

  • Worker image now ships a Python shim and a shared Playwright Chromium cache as native-session baseline tools (MNG-1055). Dockerfile.worker installs python3 + python-is-python3 so both python and python3 resolve to the same Debian-owned interpreter (closing the friction cluster MNG-887/897/926/934/947/957/973/1010/1024/1033/1039/1044). It also installs a pinned @playwright/test plus Chromium with browser dependencies into /ms-playwright (closing MNG-998/1048), then makes that cache owned by the runtime node user so project .cascade/setup.sh scripts pinned to a different Playwright revision can install the missing Chromium revision into the inherited $PLAYWRIGHT_BROWSERS_PATH. The native-tool env filter (src/backends/shared/envFilter.ts) now allowlists PLAYWRIGHT_BROWSERS_PATH as an exact match so the cache is reachable from every native-tool engine subprocess; the broader PLAYWRIGHT_* prefix is intentionally left out to preserve the defense-in-depth env-allowlist posture. A new Docker smoke script (tests/docker/worker-runtime-tools/run-test.sh) validates python, python3, python -c 'import json', the Playwright Chromium launch path, the env var, and node-user write access to the cache, and is wired into CI (docker-build-check) and both deploy workflows (deploy.yml / deploy-dev.yml) before the worker image is pushed, so a broken baseline cannot reach :latest / :dev tags. The native-tool system prompt (src/backends/shared/nativeToolPrompts.ts) exposes the guaranteed tools to agents under a new "Guaranteed runtime tools" section, and the engine-backends architecture doc, README, and Getting Started prerequisites describe the runtime baseline (including the image-size implication).

Documentation

  • Friction reporting is now documented for operators and provider contributors. Architecture docs cover the optional PM Friction slot (lists.friction for Trello, statuses.friction for JIRA/Linear), ReportFriction, and cascade-tools pm report-friction --details-file -. The integration guide explains that friction reports use existing provider createWorkItem plus optional moveWorkItem, so providers do not need a new adapter method or a DB-backed friction index. Resilience docs describe the JSONL sidecar/outbox retry path, missing-slot behavior, and non-blocking drain failures. See Trello card Rvv7VVd5.

  • Trigger architecture docs now describe the migrated trigger contracts. Added guidance for canonical TRIGGER_EVENTS, shared PM/GitHub result builders, first-match dispatch, structured skip vs bare null, no-agent results, deferred bare-job re-checks, router outcome decision reasons, PM coalescing, capacity scope, dispatch failure compensation, and wedged-lock diagnostics. Migration note for future trigger contributors: new handlers should import event constants, use the shared builders, return structured skips for claimed-but-non-dispatched events, and reserve bare null for "continue to later handlers." See Trello card qUbPtALY.

Fixed

  • Review-agent compact diffs now use local git as the authoritative patch source. GetPRDiffContext is built from the checked-out PR workspace (origin/<base>...HEAD) instead of GitHub's potentially clipped pulls.listFiles().patch body, while GitHub remains the source of changed-file ordering and counts. Files without a locally verified patch are explicitly listed in SKIPPED FILES, skipped-file guidance now points to cascade-tools scm get-pr-diff --prNumber <N> --path <path>, and PR context prepared logs include patch-source counts, token counts, skip reasons, and bounded local-vs-GitHub hunk/size mismatches. See Linear issue MNG-739.

  • Linear and JIRA inline checklist writes are now idempotent across provider/tool retries. The shared markdown checklist engine now upserts by exact ### {Checklist Name} heading and exact item text, merges duplicate inline sections into the first matching section, preserves non-checkbox prose from collapsed duplicate sections, preserves checked state when duplicate rows disagree, and keeps Trello on native checklist APIs. Linear also caches accepted description writes in-process so stale readback no longer overwrites or replays checklist appends into duplicate blocks. See Linear issue MNG-741.

  • cascade-tools scm reply-to-review-comment now accepts --body-file, and create-pr-review --comment handles one inline comment object ergonomically. ReplyToReviewComment now exposes the same generated body file-input contract as the other SCM comment commands. Array-of-object CLI params still prefer JSON arrays, but one top-level JSON object is normalized to a one-item array, while null and primitive JSON values fail early with the structured json-parse envelope instead of reaching the GitHub client. See Linear issues MNG-736, MNG-731, and MNG-729.

  • Linear inline checklist creation now preserves freshly written descriptions across stale readback windows. Linear can briefly return stale issue descriptions after accepting updateIssue(), so CASCADE records successful writes under the existing per-issue description lock and uses that in-process value as the next mutation's base while the provider catches up. Parallel AddChecklist flows no longer fail with Checklist not found in description when one flow creates a checklist heading immediately before another appends items. See Linear issue MNG-685.

  • Linear and JIRA inline checklist updates no longer lose sibling checklist rows during concurrent updates. Both providers rewrite the whole issue description for checklist mutations, so their read/mutate/write path is now serialized per provider/work item with a stale-safe temp-file lock and provider-only retry semantics. The shared inline checklist parser also keeps scanning through prose, indented detail lines, and bullet detail lines until the next heading, so ReadWorkItem reports every visible checkbox row under a checklist heading. See Linear issue MNG-656.

  • ReadWorkItem examples now render PM IDs as runnable bare CLI values. Native-tool prompt guidance and cascade-tools pm read-work-item --help now show --workItemId abc123 instead of JSON-string-literal forms like --workItemId '"abc123"'. The CLI also strips one accidental outer quote layer for ReadWorkItem IDs only, so a copied bad example no longer sends literal quote characters to the PM provider. See Trello card M5f9T1D7.

  • cascade-tools scm create-pr-review now accepts --body-file <path> and --body-file -. This matches the generated CreatePRReview guidance and the existing CreatePR / PostPRComment file-input pattern for long Markdown bodies. See Trello card 7kmo42o6.

  • PM cascade-tools runtime failures now exit non-zero with structured failure envelopes. PM core commands now throw on fatal provider/API failures, letting createCLICommand() emit {"success":false,"error":{"type":"runtime","message":"..."}} instead of wrapping prose like Error posting comment: ... inside success:true data. Intentional non-fatal PM outcomes such as guarded move no-ops and friction retry queueing remain successful command results. See Trello card lU9mHLJT.

  • Native-tool prompt examples for enum, scalar, number, and primitive-array flags now render as runnable CLI syntax instead of JSON string literals. Agent-facing guidance now shows forms such as cascade-tools scm create-pr-review --event APPROVE and repeatable primitive arrays as --labels bug --labels docs, while object and array-of-object flags continue to render shell-quoted JSON payloads. See Trello card l9Sira7y.

  • resolve-conflicts agent no longer silently skips when GitHub's async mergeability computation hasn't resolved by the time the pull_request webhook is processed (spec 020). PRConflictDetectedTrigger previously exhausted a 2×2s synchronous retry budget and silently discarded the event when mergeable === null — because GitHub never sends a follow-up webhook once mergeability resolves, the resolve-conflicts agent never fired. The trigger now returns TriggerResult.deferredRecheck, which causes the router to schedule a bare BullMQ delayed re-check job ~45s later via scheduleCoalescedJob (deduped per PR). The worker re-dispatches via the trigger registry to get fresh mergeability state. Multiple rapid webhooks for the same PR coalesce to a single re-check job. If mergeability is still null after the re-check fires, a Sentry event is captured under tag mergeability_recheck_exhausted and a WARN log is emitted — not a silent discard. Observed live on ucho/PR #329 (2026-05-07). See spec 020.

Added

  • Sentry alerts now materialize as real PM work items in the configured alerts slot (spec 019). The alerting trigger previously minted a synthetic sentry:issue:<id> workItemId, which caused a Trello 400 error on the budget gate for projects with a cost custom field, silently killing every alerting run. The trigger now calls materializeAlertWorkItem('sentry', issueId, project, hints), which creates (or idempotently retrieves) a Trello card / JIRA issue / Linear issue in the PM alerts slot and returns its native ID. Budget tracking, lifecycle transitions, and label writes all work correctly on the resulting card. A partial UNIQUE index on (project_id, external_source, external_id) in work_items ensures a second Sentry alert on the same issue produces the same PM card, not a duplicate. Configure the alerts slot in the PM wizard's Status Mapping step; validation pre-flight emits a pm-category error when the slot is unset and an alerting trigger is enabled. See spec 019.

  • Alerting agent now investigates Sentry alerts and files bug investigation work items (spec 018, plan 1 of 2). The alerting agent had been wired end-to-end except for its system prompt template — definition YAML, capabilities, trigger handlers, context pipeline, and Sentry integration were all in place, but src/agents/prompts/templates/alerting.eta was missing, so the worker crashed at agent boot with ENOENT when the first prod-traffic Sentry alert arrived (cascade project, 2026-05-06). This plan ships the prompt: a three-phase investigator (parse pre-loaded event → confirm root cause via source reads → file or comment) with an explicit INVESTIGATE-AND-FILE-ONLY guardrail. The agent does not edit source, commit, push, or open PRs — that property is enforced at the capability layer (no fs:write, no scm:*), pinned by a static test that asserts the resolved gadget allowlist excludes WriteFile, CreatePR, and CreatePRReview. When the trigger context provides an existing work item, the agent comments on it; otherwise it creates a new bug investigation work item in the configured backlog. Output structure is predictable: Investigate: <ErrorType> in <Function> (<file>:<line>) title and a 4-6 sentence + bullets description. Engine-agnostic prose; reuses partials/environment for the shared preamble. See spec 018. Plan 2 of 2 closes the silent-failure path that masked this gap (worker boot failures will produce visible failed run rows, exit code 2, Sentry capture under worker_boot_failure).

Changed

  • Worker boot-failure visibility: boot-time agent failures now produce visible failed runs (spec 018, plan 2 of 2). The worker now creates the agent_runs row before plan resolution so template-load, model-resolution, context-pipeline, definition-lookup, and run-record failures are not silently converted into invisible successful jobs. Boot failures are marked failed with the structured cause in the run row, captured to Sentry under worker_boot_failure, and re-thrown so the worker exits with code 2; ordinary in-execution crashes keep the existing exit/result semantics. The router crash-reason formatter labels exit code 2 as Worker boot failed, Sentry-driven alerting runs receive stable synthesized workItemIds (sentry:issue:<id> / sentry:metric:<org>:<title>), and a conformance test now fails CI when a YAML-registered agent type has no matching prompt template. See spec 018.
  • Pipeline-capacity gate now enforces maxInFlightItems for PM status-changed triggers (spec 017, plan 2 of 3). The gate at src/triggers/shared/pipeline-capacity-gate.ts is the hard cap on the active pipeline (TODO + IN_PROGRESS + IN_REVIEW work items) introduced after a prior incident where a human moved three cards into TODO simultaneously and three concurrent implementation runs fired against a project pinned to maxInFlightItems: 1. The gate calls getPMProvider() to count in-flight items, but for every PM status-changed trigger the call threw No PMProvider in scope because the three PM router adapters (src/router/adapters/{linear,trello,jira}.ts) wrapped trigger dispatch in their per-PM-type credential AsyncLocalStorage scope but NOT in PM-provider scope (the GitHub adapter at src/router/adapters/github.ts:280 already had both wrappings). The gate fell through to its conservative branch (WARN: pipeline-capacity-gate: PM provider unavailable, allowing run and return false) — silently no-op for the only triggers that actually need it. 32 occurrences/day on cascade-router (verified 2026-04-29). The fix introduces a shared helper withPMScopeForDispatch(project, dispatch) at src/router/adapters/_shared.ts that the three PM router adapters consume, mirroring the GitHub adapter's correct shape. The gate's "PM provider unavailable" branch is converted from WARN + return false (allow) to ERROR-level + Sentry capture under stable tag pipeline_capacity_gate_no_pm_provider + return true (block) — once the routine path establishes scope, hitting that branch is a real AsyncLocalStorage scope leak operators need to investigate. A static-guard test at tests/unit/integrations/pm-router-adapter-pm-scope.test.ts enforces the wrapping invariant per adapter; CLAUDE.md gains a "Capacity-gate invariant" passage in the Architecture section. See spec 017.
  • PM-ack dispatch consolidation: Linear-based PM-focused agents now post their PM-side ack comment (spec 017, plan 1 of 3). PM-focused agents (e.g. backlog-manager) triggered from a GitHub webhook used to silently skip their PM-side ack on Linear projects: the router-adapter's local postPMAck helper had if (pmType === 'trello') / if (pmType === 'jira') branches but no Linear branch, so Linear-based projects fell through to a WARN: Unknown PM type for PM-focused agent ack, skipping and never saw the "🔧 On it" comment that Trello/JIRA projects got (24 silent skips per day on cascade-router, all from ucho, verified 2026-04-29). A near-identical helper at src/triggers/shared/pm-ack.ts already had the Linear branch — pure parallel-path drift. The fix introduces a single consolidated helper dispatchPMAck at src/router/pm-ack-dispatch.ts that indexes the manifest registry directly and invokes manifest.platformClientFactory(projectId).postComment(...) — no per-PM-type literal branching anywhere on the dispatch surface. Both legacy call sites delegate. The PM manifest conformance harness gains a per-provider dispatchPMAck reaches this provider without throwing assertion, and a static-guard test pins "no pmType === '<literal>' branching" against all three call sites; adding a future PM provider to the registry lands the dispatch path for free. Genuinely-unknown PM types (configuration error: project pinned to a deleted provider) now log at ERROR + capture to Sentry under stable tag pm_ack_unknown_pm_type instead of a silent WARN. See spec 017.
  • Progress-comment lifecycle: post-agent cleanup hook now skips when an in-run gadget already deleted the comment (spec 017, plan 3 of 3). The post-agent deleteProgressCommentOnSuccess hook used to read sessionState.initialCommentId, fall back to result.agentInput.ackCommentId when session state was empty, and issue a redundant DELETE — but "session state cleared by a gadget" was indistinguishable from "session state never populated", so the fallback fired and re-deleted comments that were already gone. GitHub returned 404 and WARN: Failed to delete progress comment after agent success was logged 72 times per day on cascade-router (live audit on 2026-04-29). Adds an explicit initialCommentIdConsumed: boolean flag on SessionStateData. Both deleteInitialComment (gadget-driven) and clearInitialComment (sidecar-driven) now set the flag to true after disposing of the comment. The post-agent hook checks the flag first and skips the entire deletion path — including the legacy agentInput.ackCommentId fallback — when consumed. As defense in depth, githubClient.deletePRComment now treats HTTP 404 as success (RFC-7231 idempotency) and logs at DEBUG instead of letting the error bubble as a WARN; other HTTP errors (5xx, 401, network) continue to throw. The legacy fallback to agentInput.ackCommentId continues to work for code paths that never populate session state. See spec 017.
  • PM image delivery: Linear GraphQL fixture + extraction-coverage regression test (spec 016, plan 3 of 3). Captures a reconstructed Linear Issue GraphQL payload at tests/fixtures/linear-issue-with-screenshot.json containing extension-less and extensioned inline-pasted images (description + comment bodies) plus formal Attachment records (Slack/GitHub/Sentry link previews) that must NOT be mistaken for inline images. The unit test at tests/unit/pm/linear/extraction-coverage.test.ts pins the contract and fails loudly with a specific URL-missing message if Linear ever changes its payload shape in a way that loses inline images. Documents the conclusion in src/integrations/README.md: Issue.description markdown is canonical for Linear inline images; Issue.attachments is the wrong surface (formal Attachment records, not pastes). No production code change — this plan ships the regression net for the contract Plans 1+2 established. See spec 016.
  • PM image delivery: runtime cascade-tools pm read-work-item gadget now delivers images on disk (spec 016, plan 2 of 3). The runtime gadget that agents call mid-run for a work item used to return text only — its "Pre-fetched Images" section listed URL refs but no local file paths, so an agent that needed to re-read a work item (e.g. after a teammate added a screenshot) had no way to actually see the new image. After this plan, the gadget downloads any image media present and writes it to .cascade/context/images/work-item-<id>-img-<index>.<ext> (extension derived from the resolved Content-Type MIME), then returns text whose new "Local Image Files" section lists actual file paths the agent's file-read tool can consume. Failed downloads are surfaced in a "Failed Image Downloads" subsection so they're never silently dropped. Same diagnostic log line as the boot path ([image-pipeline] work-item-fetch summary) — operators see consistent shape across boot and runtime fetches. Closes the mid-run pickup gap. See spec 016.
  • PM image delivery: extension-less Linear pasted-image URLs are no longer dropped at the pre-download MIME filter (spec 016, plan 1 of 3). Linear's https://uploads.linear.app/<uuid> URLs (with no file extension in the pathname) used to fall through mimeTypeFromUrl to application/octet-stream and were silently filtered out by filterImageMedia before the download loop ran. The fix introduces an image/* wildcard sentinel for trusted PM-provider upload hosts (allowlisted by hostname); isImageMimeType now accepts the wildcard, and the download response's Content-Type header resolves it to a concrete MIME (image/png, etc.) before any image is written. The shared downloadAndPrepareImages helper consolidates the per-provider download dispatch (jira/linear/trello) so both the boot-path and the runtime gadget (spec 016 plan 2) share one code path. Adds AC#5's grep-stable diagnostic line — [image-pipeline] work-item-fetch summary — emitted once per work-item-fetch with stable fields (provider, workItemId, urlsDetected, urlsAfterFilter, urlsDownloaded, urlsFailed, urlsByMimeType). Closes the silent screenshot-drop bug class verified live on 2026-04-26 (ucho/MNG-357). See spec 016.
  • Router dispatch capacity now waits for a slot; transient Docker errors retry; terminal errors fail fast (spec 015, plan 2 of 2). Replaces guardedSpawn's synchronous "No worker slots available" throw with an in-process slot-waiter (default 5min timeout, configurable via SLOT_WAIT_TIMEOUT_MS). Adds a dispatch-error classifier that splits transient (ECONNREFUSED / ECONNRESET / ENOTFOUND / HTTP 429 / container-name 409 / SLOT_WAIT_TIMEOUT) from terminal (TypeError / ZodError / image-not-found-after-fallback). Both cascade-jobs and cascade-dashboard-jobs queue defaults now specify attempts: 4 with backoff: { type: 'exponential', delay: 5000 } (~75s total before exhaustion). Terminal errors are wrapped in BullMQ's UnrecoverableError so retries skip. Combined with plan 015/1, the original silent black-hole failure mode (verified live on 2026-04-26 via ucho/MNG-350) is fully closed: no more lost jobs on transient capacity misses or Docker hiccups, no more wedged locks. CLAUDE.md updated with the new "Dispatch failure semantics" passage. See spec 015.
  • Router dispatch failures now release in-memory locks via the BullMQ failed event (spec 015, plan 1 of 2). Hooks worker.on('failed') on both cascade-jobs and cascade-dashboard-jobs queues to call a new releaseLocksForFailedJob compensator that releases the work-item lock, agent-type concurrency counter, and recently-dispatched dedup mark for any job whose dispatch fails. Closes the stranded-lock half of the prod incident verified on 2026-04-26 (ucho/MNG-350): a transient capacity miss was leaving the in-memory work-item lock wedged for 30 minutes, silently rejecting subsequent webhooks for the same trio. Also splits the webhook decision-reason vocabulary into three states — Job queued (success), Awaiting worker slot: … (in-flight, healthy), Work item locked (no active dispatch): … (wedged-lock canary, fires a Sentry capture tagged wedged_lock_canary so any regression in compensation is loud). Plan 2 closes the lost-job half (wait-for-slot, retry budget, error classifier). See spec 015.
  • cascade-tools scm create-pr-review: --comment alias + --comments-file escape hatch (spec 014, plan 2 of 2). The command now accepts --comment (singular) as an alias for --comments — the exact muscle-memory mistake from prod run 5d993b04 now resolves correctly. Added --comments-file <path> (and - for stdin) as a JSON-parsed file alternative for long payloads that don't survive shell quoting. Zero edits to shared infrastructure (cliCommandFactory, manifestGenerator, nativeToolPrompts, errorEnvelope) — the two declarative fields on createPRReviewDef.parameters.comments.cliAliases + createPRReviewDef.cli.fileInputAlternatives are everything. Proves spec 014's single-entrypoint invariant: a new or evolved gadget should never need to touch shared machinery. See spec 014.
  • cascade-tools agent ergonomics: truthful system prompt, runnable --help, structured error envelope (spec 014, plan 1 of 2). The system-prompt renderer that describes every cascade-tools command to agents now tells the truth about array-shaped parameters — no more silent s-stripping of names, no more <string> (repeatable) claim for array-of-object flags (they correctly render as --<flag> '<json>' now, with aliases appended via | and a one-line runnable JSON example inlined from the tool definition's examples block). Every CLI failure — flag-parse, JSON-parse, missing-required, enum-mismatch, unknown-flag, auth, runtime — emits a single structured envelope on stdout ({"success":false,"error":{type,flag?,message,got?,expected?,hint?,example?}}) plus a short prose summary on stderr for humans, replacing the ad-hoc mix of this.error() prose and {success:false,error:"<string>"} flat shapes. Mistyped flags get a "did you mean" suggestion via Levenshtein match against declared canonical names + aliases. --help now renders def.examples as copy-pasteable shell invocations under an EXAMPLES section. Root-caused by prod run 5d993b04-6e05-4ae1-b7de-8c274cf3496b where a review agent wasted ~2½ min fighting the prior pre-014 surface and ultimately dropped an inline PR comment. See spec 014 + authoring guide at src/gadgets/README.md.
  • cascade-tools now streams subprocess output live (spec 013). The shared subprocess helper (on top of execa + tree-kill) forwards child stdout/stderr to the parent's stderr line-by-line as it arrives, emits a heartbeat line on stderr every 30 seconds of child silence (configurable), enforces both an idle-silence timeout (default 120s) and a wall-clock timeout (default 600s) with SIGTERM→SIGKILL escalation, and kills the full process tree on timeout. git push and git commit invoked by scm create-pr pass tighter per-caller timeouts and now return captured hook output in the result on success (previously discarded). Result shape is backward-compatible — { stdout, stderr, exitCode } preserved; new optional reason: 'idle-timeout' | 'wall-timeout' surfaces when the helper killed the child. Motivation: LLM-driven CASCADE agents watching an output file could not distinguish a slow pre-push hook (~60s of silence) from a hung process, leading to retry loops that burned 5–10+ minutes of run budget. See spec 013.
  • cascade-tools command bootstrap not found warning silenced (spec 013). The oclif command-loader glob now excludes bootstrap.js, which is a side-effect import from bin/cascade-tools.js, not a command.
  • Linear and JIRA checklists are now inline markdown, not sub-issues / subtasks. Acceptance criteria, implementation steps, and other checklist items added by CASCADE agents (via AddChecklist / AddChecklistItem) now live as - [ ] / - [x] markdown checkboxes inside the parent issue's description, under a ### {Checklist Name} heading. Previously these created full sub-issues (Linear) or subtasks (JIRA) — one per item — which cluttered boards and inflated backlog counts (a single split could create 30+ orphan items). The PMProvider interface is unchanged; only the Linear and JIRA adapter internals changed. Trello continues to use its native checklist API. Forward-only — existing sub-issues / subtasks created before this change are not migrated. See spec 008 and the new "Checklist implementation by provider" section in src/integrations/README.md.

Internal

  • PM webhook-UX manifest migration complete (spec 012). Closes the final gap from spec 011 — every PM wizard step, without exception, now renders via the manifest path. Plans 012/1-3 migrated Trello, JIRA, and Linear webhook steps into per-provider adapters composed of the shared WebhookUrlDisplayStep + provider-specific UX: Trello and JIRA each wrap the shared step with a programmatic "Create Webhook" button + active-webhooks list + per-webhook delete + curl fallback template (via existing webhooks.* tRPC endpoints with the {trelloOnly|jiraOnly} discriminator); Linear wraps it with a "Manual Webhook Setup Required" banner + ProjectSecretField (self-managing LINEAR_WEBHOOK_SECRET persistence) + 5-step manual-setup instructions. Plan 012/4 deleted the legacy WebhookStep, LinearWebhookInfoPanel, useWebhookManagement, useLinearWebhookInfo, the legacy pm-wizard-webhooks-step.test.ts file, and the -webhook id-skip filter introduced by plan 011/4. pm-wizard-common-steps.tsx now only exports SaveStep. pm-wizard.tsx iterates manifestDef.steps without exception; the legacy webhook slot is gone. No operator-visible regression. See spec 012.
  • PM wizard shared-component migration complete (spec 011). Migrates the three production PM provider wizards (Trello, JIRA, Linear) off their per-provider step files and onto the shared StandardStepKind components landed by spec 010. Five plans landed: shared-components widenings (container-pick / project-scope searchable?: boolean → cmdk Combobox; webhook-url-display optional inline signing-secret input; 7th StandardStepKind: custom-field-mapping wired to manifest.createCustomField); Trello migration (OAuth popup stays as kind: 'custom' via TrelloOAuthStep; labelDefaults? + fieldDefaults? forward-edit additive widenings pre-populate Create inputs); JIRA migration (task/subtask mapping stays as kind: 'custom' via IssueTypeMappingStep; free-text label mode exercises the empty-providerLabels path); Linear migration (credentials, team picker, status/label/project-scope all shared; webhook step composes shared WebhookUrlDisplayStep with ProjectSecretField for LINEAR_WEBHOOK_SECRET); cleanup (the three pm-wizard-{trello,jira,linear}-steps.tsx files deleted, ≈1,085 lines of legacy UI retired). Plan 011/4 also fixed a latent regression plans 011/2 + 011/3 introduced: pm-wizard.tsx hardcoded 3 manifest step slots from the spec-006 era; it now iterates over manifestDef.steps dynamically, rendering one WizardStep per entry. Legacy WebhookStep (programmatic webhook registration for Trello/JIRA + Linear signing-secret UX) retained in its own slot — migration into the manifest path is follow-up scope. No operator-visible wizard-UX change beyond consistency: every provider now has searchable pickers, and Trello/JIRA gain inline custom-field create affordances. See spec 011.
  • PM integration hardening follow-ups complete (spec 010). Finishes the PM-layer cleanup started by spec 009. Three plans landed: generic pm.discovery.createLabel(providerId, containerId, name, color?) and pm.discovery.createCustomField(providerId, containerId, name) tRPC mutation endpoints + optional createLabel / createCustomField manifest hooks replace five per-provider wizard call sites; the currentUser DiscoveryCapability is declared on all three real providers (Trello /members/me, JIRA /rest/api/3/myself, Linear viewer) and served through the unified pm.discovery.discover endpoint; six real shared React step components (credentials, container-pick, status-mapping, label-mapping, webhook-url-display, project-scope) now live at web/src/components/projects/pm-providers/steps/*.tsx and the wizard generator dispatches to them via a STANDARD_STEP_COMPONENTS registry. Existing Trello/JIRA/Linear wizards continue to use their spec-006-era per-provider step adapters; the shared path is additive — a new PM provider with purely-standard steps now writes zero per-provider step components. The new-provider-surface snapshot guard is tightened to include the six step files. No operator-visible changes. See spec 010.
  • PM integration hardening (spec 009). Makes the PMProviderManifest a behavioral contract rather than a wiring convention. Five plans landed: branded StateId / LabelId / ContainerId types in src/pm/ids.ts (state-name-vs-ID confusion is now a compile error at direct-adapter call sites); manifest-owned Zod configSchema for each provider (the central src/config/schema.ts imports from src/integrations/pm/<provider>/config-schema.ts — #1138/#1142 drift class becomes a round-trip CI failure); unified pm.discovery.discover(providerId, capability, args) tRPC endpoint driven by manifest.discoveryCapabilities; behavioral conformance harness (tests/unit/integrations/pm-conformance.test.ts) runs round-trip + lifecycle + webhook-verify + trigger-self-hook against every registered provider; single registration entrypoint at src/integrations/entrypoint.ts (router, worker, CLI, dashboard all import one file — guarded by entrypoint-usage.test.ts); shared _shared/auth-headers.ts helpers enforced by provenance test; tests/unit/pm/linear/regression-2026-04.test.ts locks in fixes for six Linear bug classes (#1112/#1117/#1118/#1119/#1131/#1133/#1134/#1137/#1138/#1139/#1142). No operator-visible changes. See spec 009.
  • PM integration plug-and-play (infrastructure). Introduced PMProviderManifest as the canonical per-provider contract — one object declares credentials, webhook route and verifier, router adapter, trigger handlers, platform client, job-id extractor, and optional label-creation hook. Landed pmProviderRegistry, a conformance test harness (tests/unit/integrations/pm-conformance.test.ts), shared helpers (_shared/auth-headers.ts, _shared/webhook-verifier.ts, _shared/label-id-resolver.ts, _shared/project-id-extractor.ts), a new pm.discovery tRPC router, and a frontend provider-wizard registry with a generic step renderer. Dormant in this release — Trello, JIRA, and Linear continue to register through the legacy path; they migrate onto the manifest in follow-up PRs. No operator-visible changes. Closes plan 006/1 of spec 006.
  • PM integration plug-and-play (Trello migrated). Trello's webhook signature verifier, router adapter, triggers, platform client, job-id extractor, wizard steps, and label/custom-field creation hooks are now composed via a single trelloManifest + trelloProviderWizard. Extended the ProviderWizardDefinition contract with an optional useProviderHooks field so provider-specific React hooks run inside a shell component — ManifestProviderWizardSection — rather than at the wizard root; this is how we satisfy the React rules-of-hooks while still keeping Trello's Discovery/LabelCreation/CustomFieldCreation hook composition per-provider. The conformance harness now exercises Trello alongside the test fixture (22 shared tests × provider). Trello's legacy registrations in bootstrap.ts stay for now because nine-plus call sites still use pmRegistry.get('trello') — plan 006/5 migrates those callers and deletes the legacy lines. No operator-visible changes. Closes plan 006/2 of spec 006.
  • PM integration plug-and-play (JIRA migrated). JIRA joins Trello on the manifest pattern with jiraManifest + jiraProviderWizard. verifyWebhookSignature uses the shared makeHmacSha256Verifier factory (Trello's bespoke scheme didn't fit, so this is the first consumer). Wizard steps + discovery / custom-field hooks moved into jiraProviderWizard.useProviderHooks; the JIRA-specific branches and hook instantiations are gone from pm-wizard.tsx. worker-env.ts::extractProjectIdFromJob JIRA branch removed (registry path handles it). Conformance harness now exercises Trello + JIRA + TestProvider (33 shared assertions × provider). Same deferrals as 006/2: bootstrap.ts JIRA registration stays until plan 006/5 migrates the pmRegistry.get('jira') callers. No operator-visible changes. Closes plan 006/3 of spec 006.
  • PM integration plug-and-play (Linear migrated — all PM providers now on manifest). linearManifest + linearProviderWizard complete the migration for all three PM providers. Linear uses the shared makeHmacSha256Verifier({ headerName: 'linear-signature' }) factory. This plan also consolidates three divergent copies of Linear auth/label logic: src/router/platformClients/linear.ts and src/router/bot-identity-resolvers.ts both switch to the shared linearAuthHeader helper, and src/pm/linear/adapter.ts::resolveLabelId delegates to the shared _shared/label-id-resolver. The divergent copies that shipped the Bearer-prefix and silent-label-drop bugs are physically deleted from the codebase. pm-wizard.tsx collapses: with all 3 providers on the manifest, the non-manifest fallback path is gone — every PM provider renders via ManifestProviderWizardSection. src/triggers/builtins.ts is now manifest-only for PM (SCM + alerting still on legacy). Conformance harness runs 44 assertions (11 × TestProvider + Trello + JIRA + Linear). Same deferrals as 006/2 + 006/3: bootstrap.ts Linear registration stays until plan 006/5 migrates the ~dozen pmRegistry.get(...) callers. No operator-visible changes. Closes plan 006/4 of spec 006.
  • PM integration plug-and-play (legacy cleanup — spec 006 complete). src/integrations/bootstrap.ts deleted. SCM (GitHub) + alerting (Sentry) self-register via new src/github/register.ts and src/sentry/register.ts side-effect modules; PM registers via its existing manifest barrel. src/pm/registry.ts becomes a read-only delegate over pmProviderRegistry so the 9 unmigrated pmRegistry.get(...) call sites (webhook handlers, manual runner, credential scope, lifecycle, GitHub adapter) keep working without changes — the adapter transparently reads from the manifest registry, making it the single source of truth for PM provider lookups. register() on the adapter is a deprecation warn. Transitional note removed from the PM integrations README; CLAUDE.md pointer updated to the final state. A follow-up PR will migrate individual call sites to pmProviderRegistry directly and consolidate the per-provider createXxxLabel tRPC endpoints under pm.discovery.* — both are additive cleanups that don't block the spec's ACs. Spec 006 is complete. Closes plan 006/5 of spec 006.

Added

  • Linear PM — optional Project scope. Operators can now narrow a Linear-backed CASCADE project to a specific Linear Project (initiative) in the PM wizard's "Board / Project Selection" step. When set, CASCADE only responds to issues that belong to that Linear Project; webhooks for issues outside the scope are silently dropped by the router (with a structured logger.info entry), outbound listings are scoped to the project, and newly-created issues (including checklist sub-issues) inherit the project. Leave the new selector empty to preserve existing team-wide behavior. Because Linear's data model requires every issue to belong to a team and scopes workflow states per team, status mappings stay team-scoped. For cross-team Linear Projects, CASCADE responds to the intersection of the configured team and project only (sibling-team issues in the same project are ignored). No migration required — existing Linear integrations are unaffected. (Spec 005.)
  • Linear status mapping — full parity with Trello and JIRA. The Linear PM wizard's Field Mapping step now exposes all eight CASCADE stages that drive agent dispatch (backlog, splitting, planning, todo, inProgress, inReview, done, merged) in lifecycle order, instead of only four. An operator can now map a Linear workflow state to any of splitting, planning, todo, or merged and have the corresponding agent (splitting, planning, implementation, backlog-manager) dispatch on issue transitions — previously these four stages were unreachable from Linear because the wizard had no slot to save them. Existing Linear integrations upgrade in place: the four new slots render as "not set" on next wizard visit; pre-existing mappings are untouched. No migration required. The normalized ProjectPMConfig.statuses type widens to declare the full nine-stage vocabulary (including debug, reserved for a future trigger), so providers can no longer silently drift from the trigger layer's dispatch map. (Spec 003, plan 1/1.)
  • Linear wizard — inline webhook signing-secret field and accurate events list. The Webhooks step of the Linear PM wizard now renders a ProjectSecretField bound to LINEAR_WEBHOOK_SECRET directly beneath the webhook URL, so operators can paste Linear's signing secret in place instead of navigating to the Credentials tab. The "Enable events" instructions now list the three event families CASCADE actually consumes — Issues (status transitions), Comments (bot @mentions), and Issue Labels ("Ready to Process") — each with a one-line rationale tracing back to the registered trigger handlers. (Spec 002, plan 2/2.)

Changed

  • Review-agent context shape: compact diffs instead of full files. The review agent's pre-fetched PR context now consists of compact per-file diffs (using GitHub's file.patch) rather than full file contents. Files that can't fit the budget — deleted, binary, oversized patch, or cumulative budget exhausted — are surfaced in a structured SKIPPED FILES injection that names each file with a reason and tells the agent how to fetch it on demand (gh pr diff, Read, Grep). This scales with PR size rather than repo size, mitigates LLM context rot, and ensures the agent is aware of (rather than blind to) the omissions. The context budget is REVIEW_DIFF_CONTEXT_TOKEN_LIMIT (200k tokens), replacing the prior 25k full-file cap. The SKIPPED FILES injection is also delivered to the four other agents that share the PR context pipeline (respond-to-ci, respond-to-pr-comment, respond-to-review, resolve-conflicts); explicit prompt guidance is added in review.yaml only. (Spec 001, plan 2/2.)

Fixed

  • Linear wizard no longer demands you re-paste your API key on every edit. The dashboard's credential-resolution helper was picking the first provider in its registered order that declared a matching role name, so a Linear-only project's ('pm', 'api_key') lookup returned TRELLO_API_KEY — which isn't configured, which surfaced as "Linear credentials not configured" in the Board / Project Selection step. Fixed by adding a required provider parameter to the helper and updating every call site. The same fix also corrects Linear webhook signature verification in the router, which was silently resolving to the JIRA webhook secret (and returning null on Linear-only projects, so verification was silently skipped in production). Discovery calls on a project with no PM integration row yet now return a distinguishable "No PM integration configured" error instead of the misleading "credentials not configured". (Spec 004, plan 1/1.)

  • JIRA resolveLifecycleConfig silent-drop of splitting / planning / todo. The JIRA PM wizard accepts mappings for all eight CASCADE stages, but the normalization step that feeds them to PMLifecycleManager was dropping splitting, planning, and todo on the floor. Any agent lifecycle hook that moved JIRA issues to those statuses silently no-op'd. Now passes all eight keys through. No operator action required — existing JIRA mappings start working once the fix deploys. (Spec 003, plan 1/1.)

  • Linear wizard Save — HTTP 500 on projects.integrations.upsert. A check constraint (chk_integration_category_provider) restricted the pm category to trello or jira; Linear support shipped without a matching constraint update, so every attempt to save a Linear PM integration failed with SQLSTATE 23514. Migration 0049 adds linear to the allowed pm providers. (Spec 002, plan 1/2.)

  • Dashboard error logs now surface DB diagnostic fields. Unhandled errors in the Hono app error handler and tRPC error formatter now include PG error code, detail, constraint, table, and column (unwrapped from .cause when Drizzle wraps a pg driver error). Clients still receive a generic "Internal server error" for unexpected INTERNAL_SERVER_ERROR throws — real diagnostics go to stdout for operators to grep. (Spec 002, plan 1/2.)

  • Review agent on external-fork and large PRs. PR checkouts now use the canonical refs/pull/N/head ref, which works for same-repo branches and external-fork branches alike. Previously, a silent git checkout <branch> failure on fork PRs caused the worker to review the base branch (dev) while believing it was on the PR branch, producing confidently wrong reviews. Any git or HEAD-SHA mismatch now fails the run loudly rather than silently continuing. Additionally, every paginated GitHub REST endpoint used in the review setup pipeline now paginates to completion, so PRs with more than 100 changed files are no longer truncated at the first page. (Spec 001, plan 1/2.)