diff --git a/.agents/skills/daily/SKILL.md b/.agents/skills/daily/SKILL.md index 61d293a..4b43ffd 100644 --- a/.agents/skills/daily/SKILL.md +++ b/.agents/skills/daily/SKILL.md @@ -14,8 +14,11 @@ Runs the full public-scan workflow once per day without the user driving it. Mod 4. **De-branded issue text** — finding-first, code-path note, one footer link max. 5. **Never rescan** — dedup candidates against `docs/scans/` slugs. 6. **Kill switch** — if `.daily-paused` exists at repo root, abort. -7. **Manual-disclosure backlog is a blocking warning.** Count the `pending_private_disclosure*` entries in `reports/.daily_state.json`. Any entry pending more than 7 days gets a loud banner at the top of the run (repo, severity, days waiting). Any High/Critical entry pending more than 14 days blocks a new scan entirely — run status-only and tell the user to clear the queue. On 2026-08-21 six drafted reports were found unsent, the oldest 22 days, including an unauthenticated-admin finding; nothing surfaced them because every phase only looked forward at the next scan. +7. **Manual-disclosure backlog is a blocking warning.** Count the `pending_private_disclosure*` entries in `reports/.daily_state.json`. Any entry pending more than 7 days gets a loud banner at the top of the run (repo, severity, days waiting). Any High/Critical entry pending more than 14 days blocks a new scan entirely — run status-only and tell the user to clear the queue. On 2026-08-21 six drafted reports were found unsent, the oldest 22 days, including an unauthenticated-admin finding; nothing surfaced them because every phase only looked forward at the next scan. **If the report cannot be delivered by ANY channel, apply guardrail 10 (`unreachable`) rather than losing a scan a day.** + 8. **State shape.** Run history lives in `runs_recent`; the `runs` key is a vestigial empty list — do not write to it. Ad-hoc keys join one of four families: `pending_private_disclosure*`, `withheld_finding_*`, `excluded_repos`/`excluded_note`, `*_note`. Drafted disclosure emails live in `reports/disclosures/` (gitignored); when one is sent, swap its pending entry for a `sent` record with date and channel. **A post published with its finding withheld while the private send is still pending is `pending_private_disclosure_*` (the family guardrail 7 counts), even when the channel is user-only and unreachable by the pipeline — never `withheld_finding_*`, which is for reports ALREADY delivered privately.** On 2026-09-16 a 31-day undelivered High (liaohch3/claude-tap) was found mis-filed under `withheld_finding_*` and had evaded the banner for ~20 runs; there is no Gmail-Sent evidence source for a Twitter DM, so ask the user whether they've already reached out rather than assuming. +9. **Channel viability is a SELECTION criterion, adaptive to queue depth.** Before committing to a candidate, resolve how a finding would reach the maintainer: `gh api repos/{o}/{r}/private-vulnerability-reporting --jq .enabled`, plus an email in SECURITY.md (root/`.github/`/`docs/`/docs site), README, FUNDING.yml, profile or org profile (commit `noreply` aliases do not count). Queue depth 0–2 → any channel is fine. **Depth 3+ → require an autonomous channel** (PVR on, or a finding that can honestly go in a public issue); skip email-only candidates and say why. **At any depth, never pick a candidate with NO channel** (PVR off AND no email anywhere AND SECURITY.md forbids public issues) — that combination is the only permanent deadlock the series has produced. Measured 2026-09-18: 14 disclosures went autonomous, 12 needed a human send, and **every** pipeline stoppage came from the human-gated half (6-report backlog: 25 days, 4 scans lost; claude-tap: 33 days, 3 scans lost). Adaptive on purpose — PVR-on repos skew mature and commercially backed, so a permanent preference would bias the series and drop the small projects that most need the review. +10. **`unreachable` is the only way to clear an undeliverable report.** Close a `pending_private_disclosure*` without delivery ONLY with all four proofs recorded in the entry: PVR `{"enabled":false}` checked on **two separate days**; no email in any of the locations above; a public issue forbidden by their own SECURITY.md or self-disclosing; and any remaining contact on a platform **the operator has said** they do not use (ask, never assume). Then re-key to `unreachable_disclosure_` so guardrail 7 stops counting it, and add one line to the public post: *"Reported channel unavailable — the maintainer's stated private channel is disabled and no other private contact could be found. Detail remains withheld."* **Keep the detail withheld** — a High with no fix and no maintainer aware serves attackers first. `unreachable` records a failed delivery; it is not a licence to disclose. If a channel opens later, re-key back and send. ## State `reports/.daily_state.json` (gitignored): `{"last_run":"YYYY-MM-DD","last_slug":"...","runs":[...]}`. If `last_run == today`, rate-limit to status-only. diff --git a/.claude/commands/daily.md b/.claude/commands/daily.md index be51774..51a0075 100644 --- a/.claude/commands/daily.md +++ b/.claude/commands/daily.md @@ -14,9 +14,28 @@ Runs the full AI PatchLab public-scan workflow end-to-end, once per day, without 3. **Strict-norm repo detection.** If the target has a real SECURITY.md (beyond GitHub's default), commercial backing, or a visible security team → post-only, or one-vuln-per-issue. Never a grouped "review" issue. 4. **De-branded issue text.** No "scanned by [tool]" header. A single public-write-up link in a footer line at most. Lead with the finding and where the affected API is actually called in the repo. 5. **Never rescan.** Dedup every candidate against existing slugs in `docs/scans/`. -6. **Manual-disclosure backlog is a blocking warning.** Count the `pending_private_disclosure*` entries in the state file. If any has been pending **more than 7 days**, print a loud banner at the top of the run naming each one (repo, severity, days waiting) before doing anything else. If any is **High or Critical and pending more than 14 days**, do not start a new scan — run status-only and tell the user the queue needs clearing first. Rationale: on 2026-08-21 six reports were found sitting unsent, the oldest 22 days, including an unauthenticated-admin finding. Nothing in the pipeline had surfaced them, because every phase only looked forward at the next scan. A report that is written but never sent is worse than one never written: the maintainer does not know, and the series has already published that something was found. +6. **Manual-disclosure backlog is a blocking warning.** Count the `pending_private_disclosure*` entries in the state file. If any has been pending **more than 7 days**, print a loud banner at the top of the run naming each one (repo, severity, days waiting) before doing anything else. If any is **High or Critical and pending more than 14 days**, do not start a new scan — run status-only and tell the user the queue needs clearing first. Rationale: on 2026-08-21 six reports were found sitting unsent, the oldest 22 days, including an unauthenticated-admin finding. Nothing in the pipeline had surfaced them, because every phase only looked forward at the next scan. A report that is written but never sent is worse than one never written: the maintainer does not know, and the series has already published that something was found. **If the report genuinely cannot be delivered by any channel, do not sit on the block — apply guardrail 8 (`unreachable`) instead of losing a scan every day.** 7. **Always the venv interpreter, never bare `python`.** The project ran three months off the shared user-site and an unrelated `pip install` downgraded pydantic to 1.x, after which `import scanner` raised ImportError. A bare `python` here silently scans with a broken interpreter or not at all. -8. **Kill switch.** If `.daily-paused` exists in the repo root, abort immediately with a one-line note. (Create/remove it to pause/resume without code changes.) +8. **`unreachable` is the only way to clear an undeliverable report.** A `pending_private_disclosure*` + entry may be closed WITHOUT delivery only when every channel is provably absent. All four pieces of + evidence are required, and all four go in the state entry: + - `gh api repos/{o}/{r}/private-vulnerability-reporting` → `{"enabled":false}`, checked on **two + separate days** (a maintainer may switch it on after a nudge); + - no email in `SECURITY.md` (root, `.github/`, `docs/`, published docs site), `README`, + `FUNDING.yml`, the maintainer's profile, or the org profile — commit `noreply` aliases do not count; + - a public issue is forbidden by the project's own SECURITY.md, **or** would itself disclose the finding; + - any remaining contact sits on a platform the operator does not use, **and the operator has said so**. + Never assume this — ask, and record the answer. + + Then re-key the entry to `unreachable_disclosure_` (guardrail 6 stops counting it) and add one + honest line to the public post: *"Reported channel unavailable — the maintainer's stated private + channel is disabled and no other private contact could be found. Detail remains withheld."* + + **Keep the detail withheld.** Do not publish a High with no fix and no maintainer aware of it; that + serves attackers before users. `unreachable` records a failed delivery, it is not a licence to + disclose. If a channel opens later, re-key it back and send. + +9. **Kill switch.** If `.daily-paused` exists in the repo root, abort immediately with a one-line note. (Create/remove it to pause/resume without code changes.) ## State & rate-limit - State file: `reports/.daily_state.json` (under gitignored `reports/`). @@ -57,7 +76,35 @@ If mode is `--status-only`, STOP here after pushing doc updates. 2. Drop any repo whose slug already exists in `docs/scans/` (`-.md`). 3. Responsiveness pre-check on the top few: recent *closed* issues + *merged* PRs from ≥2 distinct contributors in the last ~60 days → signals a maintainer who answers. Skip ghost repos. 4. Strict-norm detection: check for `SECURITY.md`, commercial backing in README, named security reviewers. Record the publication mode this implies. -5. Pick exactly ONE best candidate. Record why (stars, activity, focus, norm mode). +5. **Channel-viability pre-check — adaptive.** Before committing to a candidate, resolve how a + finding would actually reach the maintainer, and weigh that against the current manual queue + depth (count the `pending_private_disclosure*` entries in the state file): + + ``` + gh api repos/{owner}/{repo}/private-vulnerability-reporting --jq .enabled + ``` + + plus an email in `SECURITY.md` (root, `.github/`, `docs/`, published docs site), `README`, + `FUNDING.yml`, the maintainer's profile, or the org profile. Commit `noreply` aliases do not count. + + | Manual queue depth | Rule | + |---|---| + | 0–2 | Any channel is acceptable, including email-only. | + | 3 or more | **Require an autonomous channel** — PVR enabled, or a finding that can honestly go in a public issue. Skip email-only candidates and record why in the run summary. | + | any | **Never pick a candidate with no channel at all** — PVR disabled *and* no email anywhere *and* a SECURITY.md forbidding public issues. That exact combination produced the only permanent deadlock in the series. | + + *Measured 2026-09-18:* 14 disclosures went through an autonomous channel, 12 needed a human send. + **Every pipeline stoppage came from the human-gated half** — the six-report backlog (25 days, 4 + scans lost) and claude-tap (33 days, 3 scans lost). The autonomous half has never stopped a run. + This is a structural mismatch, not bad luck: roughly half of all targets land on a channel the + pipeline cannot use, and they accumulate until one crosses guardrail 6. + + **The preference is adaptive on purpose, never absolute.** PVR-enabled repos skew mature and + commercially backed; a permanent preference would bias the series toward that profile and quietly + drop the small projects that most need the review. When the queue is clear, a PVR-off project is a + fair target — that is precisely what the tiering protects. + +6. Pick exactly ONE best candidate. Record why (stars, activity, focus, norm mode, channel). ## Phase 3 — Scan ```bash @@ -65,6 +112,8 @@ If mode is `--status-only`, STOP here after pushing doc updates. ``` Use `--ignore-file` if the repo has obvious sample/example/demo subtrees (until those are shipped as defaults). +Then read `reports//coverage.json`. It is the authoritative record of what the scan examined, written on every run from the scanners' own meta findings and derived **before** `--ignore-file` suppression, so no pattern can hide it. If `complete` is `false`, the finding count is not a clean bill of health: carry the rows verbatim into the post's **Scan coverage** section and say in one sentence what was not examined. This replaces transcribing the Phase 1 preflight by hand — the preflight still runs, because catching a broken Semgrep *before* burning a scan is worth more than reporting it afterwards. + ## Phase 4 — Curate 1. Group findings by rule family. Auto-flag `tests/`, `sample/`, `examples/`, `demos/`, fixtures, placeholders as candidate-FP. 2. Inspect the top 5 real candidates in the actual repo via `gh api repos///contents/` — read the call site, confirm the threat path. @@ -72,7 +121,7 @@ Use `--ignore-file` if the repo has obvious sample/example/demo subtrees (until 4. **Evaluate the quality gate:** is there ≥1 real, exploitability-shaped, high-confidence item? Record the boolean — it decides Phase 5 filing. ## Phase 5 — Publish (gated) -1. **Always:** write `docs/scans/.md` from `docs/templates/scan-post.md`; prepend a new row to the Scans table in `docs/index.md` **and** a new bullet to `docs/scan-log.md` (the full prose archive), and bump the scan counts in both headers. Three files, every time — on 2026-09-06 the log was found eight entries behind the index because this step only named `index.md`. +1. **Always:** write `docs/scans/.md` from `docs/templates/scan-post.md` — including the **Scan coverage** block, copied from `reports//coverage.json` and never hand-written; prepend a new row to the Scans table in `docs/index.md` **and** a new bullet to `docs/scan-log.md` (the full prose archive), and bump the scan counts in both headers. Three files, every time — on 2026-09-06 the log was found eight entries behind the index because this step only named `index.md`. 2. **If quality gate TRUE and repo not strict-norm:** file a focused courtesy issue on the target (de-branded, with code-path note + concrete fix). If a finding has a clean one-line/one-file fix, also fork → branch → PR referencing the issue. 3. **If repo strict-norm:** post-only, or one issue per critical finding — no grouped issue. 4. **If quality gate FALSE:** post-only (clean-scan write-up). File nothing upstream. diff --git a/AGENTS.md b/AGENTS.md index 7d460f4..d56721d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -57,6 +57,7 @@ AI PatchLab is an AI-assisted security remediation toolkit. The MVP focuses on a - `scanner/ignore.py` - `apply_ignore(findings, patterns)` + `load_ignore_patterns(path)` provide `.gitignore`-style path suppression (used by the `--ignore-file` CLI flag). Empty-file findings are never suppressed. `DEFAULT_SAMPLE_IGNORE_PATTERNS` holds demo/sample/example subtree patterns opted into via `--ignore-samples` - `scanner/models.py` - Normalized `Finding` dataclass + severity/confidence enums + `FINDING_FIELDS` - `scanner/recommendations.py` - Deterministic keyword-based recommendation enrichment +- `scanner/coverage.py` - Per-scanner coverage derived from meta findings (`ToolCoverage`, `build_coverage`, `EXPECTED_TOOLS`, `is_complete`); feeds `reports/coverage.json` and the report's Scan Coverage block - `scanner/confidence.py` - Centralized `Finding.confidence` rules (one function per scanner + `confidence_for_meta_finding` for shared `not-installed` / `scan-error` / etc.) - `scanner/report.py` - JSON + Markdown report writers (severity-grouped, "Top Findings" highlight block, patch suggestion blocks); also exposes `filter_by_min_severity` and `select_top_findings` - `scanner/config.py` - Disabled-by-default AI review configuration loaded from environment / `.env` (`AI_PATCHLAB_*`) @@ -64,12 +65,13 @@ AI PatchLab is an AI-assisted security remediation toolkit. The MVP focuses on a - `scanner/scanners/` - Scanner adapters (`semgrep.py`, `gitleaks.py`, `trivy.py`, `dependency_scan.py`, `ai_review.py`) plus `common.py` placeholder helper and `__init__.py` registry (`SCANNERS`) - `scanner/tools/` - External scanner process runners (`semgrep_runner.py`, `gitleaks_runner.py`, `trivy_runner.py`, `pip_audit_runner.py`, `ai_review_runner.py`) - `reports/` - Generated security reports (`security_report.json`, `security_report.md`) +- `reports/coverage.json` - Per-scanner coverage manifest written on every scan (what each tool actually examined, plus a `complete` flag) - `reports/raw/` - Raw scanner JSON outputs (`semgrep.json`, `gitleaks.json`, `trivy.json`, `pip-audit.json`, `ai-review.json` when enabled) - `src/` - Legacy scaffold entry point (kept for template parity) - `src/main.py` - Legacy entry point (`python -m src.main`) - currently a loguru-wired async stub with TODOs - `.github/workflows/ci.yml` - CI: ruff + black + pytest on Python 3.11 and 3.13 - `reports/disclosures/` - Drafted private disclosure emails awaiting a manual send (gitignored) -- `tests/` - pytest tests (one module per scanner: `test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`) +- `tests/` - pytest tests (one module per scanner: `test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`) - `tests/conftest.py` - Shared fixtures (`mock_db`, `mock_http_client`, `mock_discord`, `test_config`, session `event_loop`) - `examples/` - Reference patterns to read before implementing - `PRPs/` - Active Product Requirements Prompts @@ -174,6 +176,8 @@ export AI_PATCHLAB_AI_REVIEW_COMMAND=/path/to/ai-review-wrapper - Any finding built with `confidence_for_meta_finding(...)` must also set `is_meta=True` so `--min-severity` cannot drop it - Each external tool runner lives in `scanner/tools/_runner.py`, returns a frozen `*Result` dataclass, writes raw JSON to `reports/raw/.json`, and uses `subprocess.run(..., shell=False, check=False)` with captured stdout/stderr - Do not call subprocesses directly from `scanner/scanners/*` - go through the runner module +- Coverage is derived from the raw `collect_findings` output **before** `apply_ignore` (`scanner/run_scan.py`) - `--ignore-file` does not exempt meta findings, so deriving it later would let a path pattern hide the fact that a tool never ran +- Adding a scanner to `SCANNERS` requires adding its `Finding.tool` value to `scanner/coverage.py:EXPECTED_TOOLS`; `tests/test_coverage.py::TestRegistryDrift` fails until you do - `Finding.confidence` values come from `scanner/confidence.py` - never inline `confidence="high"` / `"medium"` / `"low"` in a scanner adapter; add or reuse a rule function instead ### Fingerprint adapter contract @@ -226,9 +230,10 @@ export AI_PATCHLAB_AI_REVIEW_COMMAND=/path/to/ai-review-wrapper - Log architectural decisions in `DECISIONS.md` - Check existing ADRs before making structural changes - Record date, decision, context, and consequences -- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers +- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact ## Known Gotchas +- **Roughly half of scan targets land on a disclosure channel the pipeline cannot use.** Measured 2026-09-18: 14 disclosures went through an autonomous channel (GitHub private vulnerability reporting, or a public issue), 12 needed a human to send an email or a DM. **Every pipeline stoppage in the project's history came from the human-gated half** — a six-report backlog (25 days, 4 scans lost) and one undeliverable High (33 days, 3 scans lost). `/daily` guardrail 9 now makes channel viability a target-SELECTION criterion, tiered on queue depth, and guardrail 8 defines `unreachable` as the only way to close a report that no channel can carry. The tiering is deliberate: PVR-enabled repos skew mature and commercially backed, so a permanent preference would bias the series away from the small projects that most need review - Semgrep is a **Python program on the shared user-site interpreter**, not a standalone binary like gitleaks/trivy. Anything that breaks that interpreter breaks Semgrep too — a pydantic downgrade on 2026-08-20 made `semgrep --version` raise ImportError and every scan would have silently lost 52% of its coverage. The project `.venv` does NOT protect it. Check `semgrep --version` before trusting a scan; repair with `python -m pip install --user --upgrade "pydantic>=2.11" "httpx>=0.27"` - ALWAYS run through `.venv` (`.venv/Scripts/python.exe` on Windows). The project ran three months off the shared user-site; on 2026-08-20 an unrelated `pip install` downgraded pydantic to 1.x and httpx to 0.21 and every import broke, hours after a scan had passed - Meta findings survive `--min-severity` but are NOT yet exempt from `--ignore-file` suppression diff --git a/CLAUDE.md b/CLAUDE.md index c3ea4f1..d801394 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -53,6 +53,7 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS. - `scanner/ignore.py` — `apply_ignore(findings, patterns)` + `load_ignore_patterns(path)` provide `.gitignore`-style path suppression of findings (used by the `--ignore-file` CLI flag). Empty-file findings are never suppressed. `DEFAULT_SAMPLE_IGNORE_PATTERNS` holds the demo/sample/example subtree patterns opted into via the `--ignore-samples` flag - `scanner/models.py` — Normalized `Finding` dataclass + severity/confidence enums + `FINDING_FIELDS`; `Finding.is_meta` flags scanner-infrastructure findings that `--min-severity` must never drop - `scanner/recommendations.py` — Deterministic keyword-based recommendation enrichment +- `scanner/coverage.py` — Per-scanner coverage derived from meta findings (`ToolCoverage`, `build_coverage`, `EXPECTED_TOOLS`, `is_complete`); feeds `reports/coverage.json` and the report's Scan Coverage block - `scanner/confidence.py` — Centralized `Finding.confidence` rules (one function per scanner + `confidence_for_meta_finding` for shared `not-installed` / `scan-error` / etc.) - `scanner/report.py` — JSON + Markdown report writers (severity-grouped, "Top Findings" highlight block, patch suggestion blocks); also exposes `filter_by_min_severity` and `select_top_findings` - `scanner/config.py` — Disabled-by-default AI review configuration loaded from environment / `.env` (`AI_PATCHLAB_*`) @@ -60,12 +61,13 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS. - `scanner/scanners/` — Scanner adapters: `semgrep.py`, `gitleaks.py`, `trivy.py`, `dependency_scan.py`, `ai_review.py`, plus `common.py` placeholder helper and `__init__.py` registry (`SCANNERS`) - `scanner/tools/` — External scanner process runners: `semgrep_runner.py`, `gitleaks_runner.py`, `trivy_runner.py`, `pip_audit_runner.py`, `ai_review_runner.py` - `reports/` — Generated security reports (`security_report.json`, `security_report.md`) +- `reports/coverage.json` — Per-scanner coverage manifest written on every scan (what each tool actually examined, plus a `complete` flag) - `reports/raw/` — Raw scanner JSON outputs (`semgrep.json`, `gitleaks.json`, `trivy.json`, `pip-audit.json`, `ai-review.json` when enabled) - `src/` — Legacy scaffold entry point (kept for template parity) - `src/main.py` — Legacy point d'entrée (`python -m src.main`) — currently a loguru-wired async stub with TODOs - `.github/workflows/ci.yml` — CI: ruff + black + pytest on Python 3.11 and 3.13 (the 3.13 leg catches stdlib removals such as PEP 594 dropping `cgi`) - `reports/disclosures/` — Drafted private disclosure emails awaiting a manual send (gitignored with the rest of `reports/`) -- `tests/` — Tests pytest (`test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`) +- `tests/` — Tests pytest (`test_scanner_foundation.py`, `test_semgrep_scanner.py`, `test_gitleaks_scanner.py`, `test_trivy_scanner.py`, `test_dependency_scan.py`, `test_ai_review.py`, `test_patch_suggestions.py`, `test_recommendations.py`, `test_meta_findings.py`, `test_confidence_field_rules.py`, `test_coverage.py`) - `tests/conftest.py` — Fixtures partagées (`mock_db`, `mock_http_client`, `mock_discord`, `test_config`, session `event_loop`) - `examples/` — Code de référence — LIRE AVANT D'IMPLÉMENTER (api_client, config, discord_alert, mysql, playwright_scraper, scheduler, service) - `PRPs/` — Product Requirements Prompts (actifs) @@ -114,6 +116,8 @@ This project can optionally include a parallel Codex/OpenAI runtime via `AGENTS. - Each external tool runner lives in `scanner/tools/_runner.py`, returns a frozen `*Result` dataclass, writes the raw JSON to `reports/raw/.json`, and uses `subprocess.run(..., shell=False, check=False)` with captured stdout/stderr - New scanners must follow the same registry + runner split — do not call subprocesses directly from `scanner/scanners/*` - Any finding built with `confidence_for_meta_finding(...)` must also set `is_meta=True` so `--min-severity` cannot drop it +- Coverage is derived from the raw `collect_findings` output **before** `apply_ignore` (`scanner/run_scan.py`) — `--ignore-file` does not exempt meta findings, so deriving it later would let a path pattern hide the fact that a tool never ran +- Adding a scanner to `SCANNERS` requires adding its `Finding.tool` value to `scanner/coverage.py:EXPECTED_TOOLS`; `tests/test_coverage.py::TestRegistryDrift` fails until you do - `Finding.confidence` values come from `scanner/confidence.py` — never inline `confidence="high"` / `"medium"` / `"low"` in a scanner adapter; add or reuse a rule function instead ### Fingerprint adapter contract @@ -277,9 +281,10 @@ $env:AI_PATCHLAB_AI_REVIEW_COMMAND = "C:\tools\ai-review-wrapper.cmd" - Before making a structural decision, check DECISIONS.md for precedent - Use the architect agent (`/architect` or Task tool) for complex decisions - Format: ADR (Architecture Decision Record) — date, decision, context, consequences -- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers +- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact ## Known Gotchas +- **Roughly half of scan targets land on a disclosure channel the pipeline cannot use.** Measured 2026-09-18: 14 disclosures went through an autonomous channel (GitHub private vulnerability reporting, or a public issue), 12 needed a human to send an email or a DM. **Every pipeline stoppage in the project's history came from the human-gated half** — a six-report backlog (25 days, 4 scans lost) and one undeliverable High (33 days, 3 scans lost). `/daily` guardrail 9 now makes channel viability a target-SELECTION criterion, tiered on queue depth, and guardrail 8 defines `unreachable` as the only way to close a report that no channel can carry. The tiering is deliberate: PVR-enabled repos skew mature and commercially backed, so a permanent preference would bias the series away from the small projects that most need review - Semgrep is a **Python program on the shared user-site interpreter**, not a standalone binary like gitleaks/trivy. Anything that breaks that interpreter breaks Semgrep too — a pydantic downgrade on 2026-08-20 made `semgrep --version` raise ImportError and every scan would have silently lost 52% of its coverage. The project `.venv` does NOT protect it. Check `semgrep --version` before trusting a scan; repair with `python -m pip install --user --upgrade "pydantic>=2.11" "httpx>=0.27"` - ALWAYS run through `.venv` (`.venv/Scripts/python.exe` on Windows). The project ran for three months off the shared user-site and on 2026-08-20 an unrelated `pip install` downgraded pydantic to 1.x and httpx to 0.21: `scanner`, `fingerprint` and every test stopped importing, while the scan pipeline had passed hours earlier. A global interpreter is a shared mutable dependency - Meta findings (`Finding.is_meta=True`) survive `--min-severity` but are NOT yet exempt from `--ignore-file` suppression - a path pattern can still hide a coverage warning. Any new finding built with `confidence_for_meta_finding(...)` must also set `is_meta=True`; the two always travel together diff --git a/DECISIONS.md b/DECISIONS.md index a3d58c5..f5c453a 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -134,6 +134,35 @@ Pour les décisions qui requièrent un round de discussion avant `accepted`. +### ADR-015: Coverage is a report artifact, not a finding + +**Date:** 2026-09-17 +**Status:** accepted + +**Context:** `cloudflare/security-audit-skill` (MIT, published 2026-06-18) is an agent skill, not a scanner - 220 KB of methodology plus two zero-dependency JSON validators, with no analysis code of its own. It orchestrates isolated sub-agents through six phases and is the seed of Cloudflare's internal vulnerability harness. Its published design principles converge almost exactly on the curation doctrine this project derived independently from 105 public scans: adversarial validation by a fresh agent, "defense-in-depth gaps are not vulnerabilities", "severity requires impact", and an explicit `needs_validation` verdict for facts that are not visible in source. That convergence is evidence the doctrine is right; it is not a reason to adopt the mechanism, which is non-deterministic, un-CI-able, priced per run, and refuses to execute target code without an OS-enforced sandbox. Two of its structural ideas do not exist here and are cheap to port. + +**Decision Drivers:** +- Must not weaken determinism or reproducibility - they are the only properties this project has that an LLM auditor cannot copy (must-have) +- Must address a failure mode already observed in the field, not a hypothetical one (must-have) +- Must fit MVP discipline: no new dependencies, no agent orchestration, files under 300 lines (must-have) +- Should convert per-scan curation judgement into an accumulating asset instead of discarding it (should-have) + +**Considered Options:** +- **A - Adopt the skill as the pipeline.** Pros: far deeper reasoning than any SAST, free, strong brand. Cons: not reproducible, cannot gate CI, needs a sandbox most environments lack, no published precision data at three months old. Rejected as a replacement; retained as a complementary point-in-time tool. +- **B - Ignore it.** Pros: zero cost. Cons: forfeits two structural ideas that answer a documented recurring failure. Rejected. +- **C - Adapt two ideas, drop the rest.** Port (1) coverage as a first-class artifact and (2) deterministic rejection records with a retained reason. Discard the six phases, sub-agent isolation, the artifact-promotion procedure, `coverage_id` encoding, and agent budgets - none apply to a synchronous subprocess pipeline. **Selected.** +- **D - Build a competing agent skill.** Pros: plays to the curated FP corpus. Cons: head-on competition with a Cloudflare-branded 10k-star MIT skill on distribution, and it abandons determinism to fight on their ground. Rejected now; revisit only as a *curation* layer over any scanner's output once the rejection corpus from (2) exists. + +**Decision:** Take option C. Coverage stops being inferred from `info`-severity meta findings scattered through the findings list and becomes `reports/coverage.json` plus a coverage block rendered **before** the findings in `security_report.md`, carrying an explicit `complete: bool`. Rejection records follow as a separate change: anything the scanner suppresses itself (ignore patterns, secrets baselines, a framework rule fired against a project that does not declare the framework) is retained with a machine-readable reason instead of being deleted. + +**Consequences:** +- Positive: the honesty that ADR-013 bought at the `Finding` level is promoted to the report level. `is_meta` already marks every ingredient and every adapter already emits them, so this is a rendering and reconciliation change, not new detection. A reader can no longer mistake "nothing was looked at" for "nothing was found" - the exact confusion behind the three field incidents: Semgrep silently losing 52% of coverage after an unrelated pydantic downgrade, tracecat's two core files drawing zero rules with `paths.skipped` empty, and `scan_dependency` reading root-only so a monorepo with no top-level manifest rendered as clean. +- Positive: it is a real product differentiator. Scanners report findings; almost none report what they failed to examine. +- Negative: reports get longer and some will now open by announcing they are incomplete. That is the intended cost. +- Negative: rejection records grow the report surface and need their own suppression discipline, or they become the noise they were meant to remove. +- Risks: scope creep toward porting more of the skill. The mitigation is this ADR - anything beyond the two named ports needs its own decision. Per-tool unit counts (files scanned, manifests audited) are deliberately excluded from v1 because they would couple the coverage module to runner internals; they are a follow-up once the artifact exists. +- Relates to ADR-013 (meta findings exempt from severity filtering) and ADR-014 (field-derived confidence tiers); both are prerequisites that made this cheap. + ### ADR-014: Field-derived confidence tiers, measured not guessed **Date:** 2026-08-21 **Status:** accepted diff --git a/PRPs/done/2026-09-17-coverage-manifest.md b/PRPs/done/2026-09-17-coverage-manifest.md new file mode 100644 index 0000000..fec367a --- /dev/null +++ b/PRPs/done/2026-09-17-coverage-manifest.md @@ -0,0 +1,215 @@ +# PRP: Coverage manifest — report what the scan did not look at + +## Overview + +Promote scan coverage from an `info`-severity finding buried in the list to a first-class +report artifact: `reports/coverage.json` plus a `## Scan Coverage` section in +`security_report.md`, with an explicit `complete` flag and a banner when the scan is +incomplete. + +Every ingredient already exists — each adapter emits `is_meta=True` findings for +not-installed, disabled, scan-error, parse-error, partial-coverage and no-manifest states +(ADR-013). What is missing is the reconciliation: today a reader must notice one `info` row +among dozens to learn that Semgrep never ran. This PRP makes that impossible to miss. + +Implements the first half of ADR-015 (option C). The second half — deterministic rejection +records — is a separate PRP and is explicitly **out of scope** here. + +## Dependencies + +- None new. Standard library only (`dataclasses`, `json`, `pathlib`). +- Requires ADR-013 (`Finding.is_meta`) — already shipped. + +## Context & References + +### MUST READ — Load these into your context + +- `DECISIONS.md` → **ADR-015** (this decision), ADR-013 (meta findings), ADR-014 (confidence tiers) +- `scanner/models.py` → `Finding`, `FINDING_FIELDS`, the `is_meta` docstring +- `scanner/scanners/__init__.py` → `SCANNERS` registry (5 entries) +- `scanner/run_scan.py:26-55` → `collect_findings`, then `apply_ignore` (line 52), then `filter_by_min_severity` (line 53), then `write_reports` (line 54) +- `scanner/report.py` → `build_report`, `write_markdown_report`, `write_reports`, `select_top_findings` +- `scanner/confidence.py` → `confidence_for_meta_finding` + +### Critical Gotchas + +1. **Build coverage from the raw `collect_findings` output — before `apply_ignore`.** + `--ignore-file` does not yet exempt meta findings (known gotcha, logged in ROADMAP). + A path pattern matching the repo root can currently suppress a `semgrep-not-installed` + finding. Deriving coverage upstream of line 52 closes that hole for coverage without + touching `apply_ignore` itself. +2. **A tool that emits nothing at all must still appear as a row.** That is the entire + point — "no findings" and "never ran" must not render identically. Absence of a meta + finding is the positive signal (`ran`), so the tool list cannot be derived from the + findings alone; it comes from a declared constant guarded by a registry test. +3. **Tool identifiers are not scanner function names and not finding-id prefixes.** + The five `Finding.tool` values are `semgrep`, `gitleaks`, `trivy`, `dependency-scan`, + `ai-security-review`. Note that `dependency-scan` emits ids prefixed `pip-audit-*`. + Key on `Finding.tool`, never on the id prefix. +4. **`ai-review-disabled` is the default state, not a failure.** AI review is + disabled-by-default by ADR-010. Its status is `not_run`, but it must not on its own + make every ordinary report read as broken — see the `complete` definition below. +5. Files stay under 300 lines. Put derivation logic in `scanner/coverage.py`, not in the + report writer; check `scanner/report.py` length before adding to it. + +## Architecture + +### New Files + +- `scanner/coverage.py` — `ToolCoverage` frozen dataclass, `COVERAGE_STATUSES`, + `EXPECTED_TOOLS`, `build_coverage()`, `is_complete()`. Pure functions over a finding + list; no I/O, no subprocess. +- `tests/test_coverage.py` — derivation, precedence, registry drift, ignore-immunity, + and report-rendering tests. + +### Modified Files + +- `scanner/report.py` — `build_report` accepts and embeds coverage; `write_markdown_report` + renders the banner and table; `write_reports` writes `coverage.json` and returns its path. +- `scanner/run_scan.py` — build coverage right after `collect_findings`, thread it to + `write_reports`, print the coverage path. +- `CLAUDE.md` / `AGENTS.md` / `ROADMAP.md` — housekeeping (Task 6). + +### Data contract + +```python +COVERAGE_STATUSES = ("ran", "partial", "error", "not_run") # best -> worst + + +@dataclass(frozen=True) +class ToolCoverage: + tool: str + status: str # one of COVERAGE_STATUSES + detail: str # human sentence derived from the meta findings + meta_finding_ids: tuple[str, ...] +``` + +Status derivation per tool, **worst status wins** (iterate `COVERAGE_STATUSES` in reverse): + +| Meta finding id suffix | Status | +|---|---| +| `-not-installed`, `-disabled` | `not_run` | +| `-scan-error`, `-json-parse-error`, `-command-error` | `error` | +| `-partial-coverage`, `-no-supported-manifest` | `partial` | +| *(no meta finding for this tool)* | `ran` | + +`reports/coverage.json`: + +```json +{ + "repository": "...", + "generated_at": "UTC ISO-8601", + "complete": false, + "tools": [ + { + "tool": "semgrep", + "status": "partial", + "detail": "...", + "meta_finding_ids": ["semgrep-partial-coverage"] + } + ] +} +``` + +`complete` is `True` only when every tool is `ran`, **except** that a tool whose sole meta +finding is `ai-review-disabled` does not by itself set `complete: false` — an opt-in +feature left off is a configuration state, not a coverage gap. Record that carve-out in a +comment next to the rule; it is the one judgement call in this module. + +## Implementation Plan + +### Task 1: `scanner/coverage.py` + +Write `ToolCoverage`, `COVERAGE_STATUSES`, `EXPECTED_TOOLS`, and +`build_coverage(findings: list[Finding]) -> tuple[ToolCoverage, ...]`. One row per entry in +`EXPECTED_TOOLS`, in registry order. Google-style docstrings, full type hints. +`build_coverage` must be total: an unknown `Finding.tool` is ignored rather than raising. + +### Task 2: Registry drift test + +In `tests/test_coverage.py`, assert `EXPECTED_TOOLS` equals the set of `Finding.tool` +values reachable from `SCANNERS`. Run each scanner against an empty temp repo with no +external tools on PATH and collect the tool names from the meta findings they emit. This +test is the guardrail that makes a new scanner adapter fail loudly until it is added to +coverage. + +### Task 3: Report integration + +- `build_report(repo_path, findings, coverage=None)` → add a `"coverage"` key + (`{"complete": bool, "tools": [...]}`), omitted entirely when `coverage is None` so + existing callers and tests are unaffected. +- `write_markdown_report`: when coverage is present and `complete` is `False`, emit a + one-line blockquote banner immediately after the header — + `> **Incomplete scan.** 2 of 5 tools did not run fully. See Scan Coverage below.` + Then Top Findings (unchanged), then a `## Scan Coverage` table before the grouped + findings. +- `write_reports(..., coverage=None)` → write `reports/coverage.json` when coverage is + present, and add `"coverage"` to the returned path dict. + +### Task 4: Wire `run_scan.py` + +Build coverage from the output of `collect_findings` **before** `apply_ignore`. Pass it to +`write_reports`. Print `Coverage report: ` alongside the existing two lines. + +### Task 5: Tests + +`tests/test_coverage.py` covers: + +1. Each status derivation from a synthetic finding list. +2. Worst-status-wins when a tool emits two meta findings of different classes. +3. A tool with zero findings renders `ran`, not missing. +4. `complete` is `False` when Semgrep is not installed; `True` when the only meta finding + is `ai-review-disabled`. +5. **Ignore-immunity:** an ignore pattern matching the repo root does not remove the + `semgrep-not-installed` row from coverage. +6. Markdown contains the banner when incomplete and does not when complete. +7. `coverage.json` is written, parses, and its keys match the contract. + +### Task 6: Housekeeping + +Update `ROADMAP.md` (mark the item done with date), the Key Directories and Scanner +adapter contract sections of `CLAUDE.md`, and the matching sections of `AGENTS.md`. Flip +ADR-015 `Status:` from `proposed` to `accepted`. Move this PRP to `PRPs/done/`. + +## Final Validation Loop + +```bash +# 1. Lint +.venv/Scripts/ruff.exe check scanner src/ tests/ fingerprint/ + +# 2. Format check +.venv/Scripts/python.exe -m black --check scanner src/ tests/ fingerprint/ + +# 3. Tests +.venv/Scripts/python.exe -m pytest tests/ -v + +# 4. Smoke test — self-scan, then confirm the artifact exists and is honest +.venv/Scripts/python.exe scanner/run_scan.py --repo "." --reports-dir reports +cat reports/coverage.json +``` + +## Success Criteria + +- [ ] `reports/coverage.json` is written on every scan, with one row per registered scanner +- [ ] A scan where Semgrep is missing renders a banner and `"complete": false` +- [ ] A scan where every tool ran renders no banner and `"complete": true` +- [ ] An `--ignore-file` pattern cannot remove a row from coverage +- [ ] Adding a scanner to `SCANNERS` without updating `EXPECTED_TOOLS` fails a test +- [ ] `scanner/report.py` and `scanner/coverage.py` are both under 300 lines +- [ ] Full lint + format + test suite green on Python 3.11 and 3.13 +- [ ] No new dependency added + +## PRP Quality Checklist + +- [x] Decision recorded in an ADR before implementation (ADR-015) +- [x] Known gotchas named with the field incident behind each +- [x] Scope explicitly bounded (rejection records deferred to a second PRP) +- [x] Every new behaviour has a matching test in the plan +- [x] Housekeeping task included + +## Confidence Score: 8/10 + +The derivation is mechanical and the raw material already exists, which is why this is +high. The point deducted is the `complete` carve-out for `ai-review-disabled` — it is a +judgement call, and if the report ends up reading as "incomplete" on every ordinary scan +the flag will be ignored within a week. Validate that on the self-scan before accepting. diff --git a/ROADMAP.md b/ROADMAP.md index 0babc28..5407219 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -51,6 +51,13 @@ - [x] JSON + Markdown match report with mandatory disclaimer (2026-05-14) - [x] ADR-012 logged (2026-05-14) +## Phase 4.7 - Disclosure throughput (the pipeline's measured bottleneck) +- [x] Channel viability is a target-selection criterion, tiered on manual-queue depth (2026/09/18) +- [x] `unreachable` disposition with four written evidence criteria, closing the deadlock class (2026/09/18) +- [x] Outcome column on the public index documented, and every value in use covered by the legend (2026/09/18) +- [ ] Track the autonomous/human-gated split per month, to see whether the tiering actually moves it +- [ ] Re-key sweep: assert at startup that no `withheld_finding_*` entry lacks a delivery record (the claude-tap mis-key evaded guardrail 6 for ~20 runs) + ## Phase 4.6 - Field-driven curation (from the public scan series) - [x] Meta findings exempt from `--min-severity` (2026/08/21) - [x] Semgrep partial-coverage finding built from the `errors` array, not `paths.skipped` (2026/08/21) @@ -61,6 +68,8 @@ - [ ] Report lockfile and open-floor dependency sets as separate rows naming the install path, rather than one merged verdict - [ ] Timeouts on the semgrep, trivy and gitleaks runners (pip-audit done; the others share the same hang risk) - [ ] Exempt meta findings from `--ignore-file` suppression as well as `--min-severity` +- [x] Coverage manifest: `reports/coverage.json` + a `## Scan Coverage` block rendered before the findings, so "never ran" cannot render as "found nothing" (2026/09/18) - ADR-015; published posts carry it via `docs/templates/scan-post.md` +- [ ] Retain deterministic rejection records with a machine-readable reason instead of deleting suppressed findings - builds the curation corpus (ADR-015, second half; needs its own PRP) ## Phase 4.5 - Polish & Stabilize - [ ] Add integration tests with sample vulnerable repositories diff --git a/docs/index.md b/docs/index.md index 0e65b85..26efc3b 100644 --- a/docs/index.md +++ b/docs/index.md @@ -89,7 +89,8 @@ login and static assets. Fifty-two flagged, none reported. - Critical issues are reported to maintainers under responsible disclosure before being published here in full detail. - Where a project's security policy forbids public vulnerability reports, the - finding is withheld from this page too — those rows read *private*. + finding is withheld from this page too — those rows read *private*, or + *undelivered* where no private channel turned out to exist. - Posts focus on patterns and lessons — not exploit walkthroughs. ## All scans @@ -98,6 +99,15 @@ login and static assets. Fifty-two flagged, none reported. **Real** is what survived curation. The gap between those two columns is the entire job. +**Outcome** reads: **fixed** — the maintainer shipped a fix; *accepted* — a private +advisory was accepted into triage but no fix has shipped yet; *open* — reported +publicly and still open; *declined* — the maintainer considered it and said no, and +the write-up quotes their reasoning; *private* — reported through a private channel +and withheld here until it resolves; *undelivered* — no private channel existed at +all, so the report could not be delivered and the detail stays withheld +indefinitely; *partial* — only part of a multi-item report landed; — — nothing was +filed, which is the usual outcome of a clean scan. + | Date | Repository | Findings | Real | Outcome | | --- | --- | ---: | --- | --- | | 2026-09-15 | [experientiallabs/experiential](scans/experientiallabs-experiential.html) | 77 | 1 real — withheld | private | @@ -126,7 +136,7 @@ entire job. | 2026-08-19 | [Mai-with-u/MaiBot](scans/mai-with-u-maibot.html) | 373 | 1 real — withheld | private | | 2026-08-18 | [roflcoopter/viseron](scans/roflcoopter-viseron.html) | 399 | 1 real — withheld | private | | 2026-08-17 | [zilliztech/memsearch](scans/zilliztech-memsearch.html) | 90 | 1 real | **fixed** | -| 2026-08-16 | [liaohch3/claude-tap](scans/liaohch3-claude-tap.html) | 104 | 1 real — withheld | private | +| 2026-08-16 | [liaohch3/claude-tap](scans/liaohch3-claude-tap.html) | 104 | 1 real — withheld | undelivered | | 2026-08-15 | [TracecatHQ/tracecat](scans/tracecathq-tracecat.html) | 212 | 0 real | — | | 2026-08-14 | [datalayer/jupyter-mcp-server](scans/datalayer-jupyter-mcp-server.html) | 37 | 0 real | private | | 2026-08-13 | [lightseekorg/tokenspeed](scans/lightseekorg-tokenspeed.html) | 181 | 1 real — withheld | private | @@ -139,7 +149,7 @@ entire job. | 2026-08-06 | [nottelabs/notte](scans/nottelabs-notte.html) | 226 | 1 real — withheld | private | | 2026-08-05 | [Vexa-ai/vexa](scans/vexa-ai-vexa.html) | 297 | 1 real — withheld | private | | 2026-08-04 | [ArcReel/ArcReel](scans/arcreel-arcreel.html) | 82 | 1 real — withheld | private | -| 2026-08-03 | [the-momentum/open-wearables](scans/the-momentum-open-wearables.html) | 145 | 2 real | ✅ | +| 2026-08-03 | [the-momentum/open-wearables](scans/the-momentum-open-wearables.html) | 145 | 2 real | **fixed** | | 2026-08-02 | [Observal/Observal](scans/observal-observal.html) | 1,117 | 1 real — withheld | private | | 2026-08-01 | [repowise-dev/repowise](scans/repowise-dev-repowise.html) | 86 | 2 real — withheld | private | | 2026-07-31 | [rocketride-org/rocketride-server](scans/rocketride-org-rocketride-server.html) | 268 | 1 real — withheld | private | diff --git a/docs/scans/liaohch3-claude-tap.md b/docs/scans/liaohch3-claude-tap.md index a868b05..767a57c 100644 --- a/docs/scans/liaohch3-claude-tap.md +++ b/docs/scans/liaohch3-claude-tap.md @@ -10,9 +10,11 @@ date: 2026-08-16 **Repository:** [liaohch3/claude-tap](https://github.com/liaohch3/claude-tap) **Commit scanned:** `901b856f` (main at scan time) **Scan date:** 2026-08-16 -**Disclosure status:** withheld — strict-norm repo (`SECURITY.md` forbids public -vulnerability issues), one real finding described here at class level only. The -private report is a manual step this pipeline cannot take — see the channel note. +**Disclosure status:** withheld — **reported channel unavailable.** The +maintainer's stated private channel (GitHub private vulnerability reporting) is +disabled, and no other private contact could be found. The one real finding stays +described here at class level only, and the full report is held for the moment a +channel opens — see the channel note. ## Summary @@ -270,6 +272,22 @@ with a count and a representative location. - **2026-08-16** — Public post (this page), finding withheld at class level. Private disclosure to the maintainer is a manual step (PVR disabled, no published email); it has not yet been made. +- **2026-09-16** — Finding re-verified still present on `main` (`a05e5582`); no + routing change had shipped. Channel re-checked: still no private route. +- **2026-09-18** — **Closed as undeliverable.** Every sanctioned channel was + checked and none exists: private vulnerability reporting returns + `{"enabled":false}` (checked on three separate days), no email address appears + in `SECURITY.md`, `README.md`, `FUNDING.yml`, the maintainer's profile or the + organisation profile (commit authorship is the GitHub `noreply` alias), and the + project's own `SECURITY.md` forbids a public issue — which would in any case + disclose the finding. The only remaining contact is a social account this + reporter does not hold. + + **The detail stays withheld.** A High-severity issue with no fix shipped and no + maintainer aware of it serves an attacker before it serves a user, so nothing + further is published here. The full report — mechanism, attack chain, and a + proposed fix — is written and held. **If private reporting is enabled on the + repository, or an address is published anywhere, it will be sent the same day.** ## Reproduce diff --git a/docs/templates/scan-post.md b/docs/templates/scan-post.md index ad055a1..02e1c33 100644 --- a/docs/templates/scan-post.md +++ b/docs/templates/scan-post.md @@ -23,6 +23,23 @@ date: YYYY-MM-DD **Total findings:** N (M of interest after curation) +## Scan coverage + +/coverage.json` - never hand-write them. +A tool that did not run reports nothing, which reads identically to a clean +result unless it is stated here. If `complete` is false, say in one sentence +what was not examined, and say it before the findings.> + +| Tool | Status | Detail | +| --- | --- | --- | +| `semgrep` | `ran` | | +| `gitleaks` | `ran` | | +| `trivy` | `ran` | | +| `dependency-scan` | `ran` | | +| `ai-security-review` | `not_run` | disabled by default (ADR-010) | + +**Coverage complete:** yes / no + ## Top findings ### 1. diff --git a/scanner/coverage.py b/scanner/coverage.py new file mode 100644 index 0000000..ce2b974 --- /dev/null +++ b/scanner/coverage.py @@ -0,0 +1,166 @@ +"""Scan coverage derived from scanner-infrastructure findings. + +A finding count only means something next to a statement of what was examined. +Every scanner adapter already emits `is_meta=True` findings when it cannot run, +when it crashes, or when it covered only part of what it was pointed at +(ADR-013). Those signals are rendered at `info` severity among the findings, +where a reader has to notice one row among dozens to learn that Semgrep never +ran. + +This module reconciles them into one row per registered scanner so that +"nothing was looked at" can never render identically to "nothing was found" +(ADR-015). +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any + +from scanner.models import Finding + +COVERAGE_STATUSES = ("ran", "partial", "error", "not_run") +"""Coverage states ordered best to worst - later entries win a tie.""" + +EXPECTED_TOOLS = ( + "semgrep", + "gitleaks", + "trivy", + "dependency-scan", + "ai-security-review", +) +"""`Finding.tool` value of every scanner in `scanner.scanners.SCANNERS`. + +Declared rather than derived from the findings: a scanner that ran cleanly +emits nothing at all, and its absence from the report is exactly the signal +this module exists to prevent. `tests/test_coverage.py` fails if this drifts +from the registry. +""" + +EXPECTED_OFF_FINDING_IDS = frozenset({"ai-review-disabled"}) +"""Meta findings that report an opt-in feature being off, not a coverage gap. + +AI review is disabled by default and stays that way unless explicitly +configured (ADR-010). Counting it as a gap would make every ordinary scan +announce itself as incomplete, and a banner that fires every time is a banner +nobody reads. +""" + +_NOT_RUN_SUFFIXES = ("-not-installed", "-disabled") +_ERROR_SUFFIXES = ("-scan-error", "-json-parse-error", "-command-error") +_PARTIAL_SUFFIXES = ("-partial-coverage", "-no-supported-manifest") + +_STATUS_RANK = {status: index for index, status in enumerate(COVERAGE_STATUSES)} + +_RAN_DETAIL = "Ran without reporting a coverage problem." + + +@dataclass(frozen=True) +class ToolCoverage: + """What one scanner actually managed to examine during a scan.""" + + tool: str + status: str + detail: str + meta_finding_ids: tuple[str, ...] = () + + def __post_init__(self) -> None: + """Validate the normalized status early.""" + if self.status not in COVERAGE_STATUSES: + raise ValueError(f"Unsupported coverage status: {self.status}") + + def to_dict(self) -> dict[str, Any]: + """Return a JSON-serializable coverage row.""" + return { + "tool": self.tool, + "status": self.status, + "detail": self.detail, + "meta_finding_ids": list(self.meta_finding_ids), + } + + +def status_for_meta_finding(finding_id: str) -> str: + """Map a meta finding id to the coverage state it implies. + + Args: + finding_id: Normalized `Finding.id`, e.g. `"semgrep-not-installed"`. + + Returns: + One of `COVERAGE_STATUSES`. An unrecognized meta finding falls back to + `"partial"`: by definition it describes the state of the scan rather + than a defect in the scanned code, and defaulting it to `"ran"` would + reintroduce the silence this module removes. + """ + if finding_id.endswith(_NOT_RUN_SUFFIXES): + return "not_run" + if finding_id.endswith(_ERROR_SUFFIXES): + return "error" + if finding_id.endswith(_PARTIAL_SUFFIXES): + return "partial" + return "partial" + + +def build_coverage(findings: list[Finding]) -> tuple[ToolCoverage, ...]: + """Derive one coverage row per registered scanner. + + Call this on the raw output of `collect_findings`, before `apply_ignore`: + `--ignore-file` does not yet exempt meta findings, so a path pattern + matching the repository root can otherwise suppress the very finding that + says a tool never ran. + + Args: + findings: Normalized findings from every scanner, unfiltered. + + Returns: + One row per entry in `EXPECTED_TOOLS`, in registry order. The worst + state a tool reported wins; findings from unknown tools are ignored + rather than raising. + """ + by_tool: dict[str, list[Finding]] = {tool: [] for tool in EXPECTED_TOOLS} + for finding in findings: + if finding.is_meta and finding.tool in by_tool: + by_tool[finding.tool].append(finding) + + rows: list[ToolCoverage] = [] + for tool in EXPECTED_TOOLS: + metas = by_tool[tool] + if not metas: + rows.append(ToolCoverage(tool=tool, status="ran", detail=_RAN_DETAIL)) + continue + worst = max(metas, key=lambda finding: _STATUS_RANK[status_for_meta_finding(finding.id)]) + rows.append( + ToolCoverage( + tool=tool, + status=status_for_meta_finding(worst.id), + detail=worst.title, + meta_finding_ids=tuple(finding.id for finding in metas), + ) + ) + return tuple(rows) + + +def _counts_as_covered(row: ToolCoverage) -> bool: + """Return true when a row does not represent a gap in what was examined.""" + if row.status == "ran": + return True + return set(row.meta_finding_ids).issubset(EXPECTED_OFF_FINDING_IDS) + + +def incomplete_tools(coverage: tuple[ToolCoverage, ...]) -> tuple[ToolCoverage, ...]: + """Return the rows that represent a real gap in what was examined.""" + return tuple(row for row in coverage if not _counts_as_covered(row)) + + +def is_complete(coverage: tuple[ToolCoverage, ...]) -> bool: + """Return true when every scanner examined what it was pointed at.""" + return not incomplete_tools(coverage) + + +def coverage_payload(coverage: tuple[ToolCoverage, ...]) -> dict[str, Any]: + """Return the JSON-serializable coverage block embedded in reports.""" + return { + "complete": is_complete(coverage), + "incomplete_tool_count": len(incomplete_tools(coverage)), + "tool_count": len(coverage), + "tools": [row.to_dict() for row in coverage], + } diff --git a/scanner/report.py b/scanner/report.py index f2a311c..7cd7fa1 100644 --- a/scanner/report.py +++ b/scanner/report.py @@ -7,6 +7,7 @@ from pathlib import Path from typing import Any +from scanner.coverage import ToolCoverage, coverage_payload from scanner.models import CONFIDENCES, FINDING_FIELDS, SEVERITIES, Finding DEFAULT_TOP_FINDINGS_LIMIT = 5 @@ -67,19 +68,34 @@ def group_findings_by_severity(findings: list[Finding]) -> dict[str, list[dict[s return grouped -def build_report(repo_path: Path, findings: list[Finding]) -> dict[str, Any]: - """Build the complete JSON report payload.""" +def build_report( + repo_path: Path, + findings: list[Finding], + coverage: tuple[ToolCoverage, ...] | None = None, +) -> dict[str, Any]: + """Build the complete JSON report payload. + + Args: + repo_path: Scanned repository root. + findings: Findings to report, already filtered. + coverage: Per-scanner coverage rows from `scanner.coverage`. Omitted + entirely from the payload when `None`, so callers that do not + supply it keep their previous output. + """ grouped = group_findings_by_severity(findings) summary = {severity: len(grouped[severity]) for severity in SEVERITIES} top = [finding.to_dict() for finding in select_top_findings(findings)] - return { + report: dict[str, Any] = { "repository": str(repo_path.resolve()), "generated_at": datetime.now(UTC).isoformat(), "summary": summary, "top_findings": top, "findings_by_severity": grouped, } + if coverage is not None: + report["coverage"] = coverage_payload(coverage) + return report def write_json_report(report: dict[str, Any], report_path: Path) -> None: @@ -95,12 +111,29 @@ def write_markdown_report(report: dict[str, Any], report_path: Path) -> None: f"Repository: `{report['repository']}`", f"Generated at: `{report['generated_at']}`", "", - "## Summary", - "", - "| Severity | Findings |", - "| --- | ---: |", ] + coverage = report.get("coverage") + if coverage and not coverage["complete"]: + lines.extend( + [ + f"> **Incomplete scan.** {coverage['incomplete_tool_count']} of " + f"{coverage['tool_count']} tools did not examine what they were pointed at. " + "A low finding count below does not mean this repository is clean - " + "see Scan Coverage.", + "", + ] + ) + + lines.extend( + [ + "## Summary", + "", + "| Severity | Findings |", + "| --- | ---: |", + ] + ) + for severity in SEVERITIES: lines.append(f"| {severity.title()} | {report['summary'][severity]} |") @@ -121,6 +154,23 @@ def write_markdown_report(report: dict[str, Any], report_path: Path) -> None: ] ) + if coverage: + lines.extend( + [ + "## Scan Coverage", + "", + "One row per configured scanner. A tool that did not run reports " + "nothing, which is indistinguishable from a clean result unless it " + "is stated here.", + "", + "| Tool | Status | Detail |", + "| --- | --- | --- |", + ] + ) + for row in coverage["tools"]: + lines.append(f"| `{row['tool']}` | `{row['status']}` | {row['detail']} |") + lines.append("") + lines.extend(["## Findings", ""]) for severity in SEVERITIES: @@ -192,14 +242,43 @@ def _indent_code_block(value: str) -> str: return "\n".join(f" {line}" for line in value.splitlines()) -def write_reports(repo_path: Path, findings: list[Finding], reports_dir: Path) -> dict[str, Path]: - """Create the reports directory and write JSON plus Markdown reports.""" +def write_reports( + repo_path: Path, + findings: list[Finding], + reports_dir: Path, + coverage: tuple[ToolCoverage, ...] | None = None, +) -> dict[str, Path]: + """Create the reports directory and write the JSON, Markdown and coverage reports. + + Args: + repo_path: Scanned repository root. + findings: Findings to report, already filtered. + reports_dir: Directory the reports are written to. + coverage: Per-scanner coverage rows. When supplied, `coverage.json` is + written beside the reports and returned under the `"coverage"` key. + + Returns: + Mapping of report kind to written path. + """ reports_dir.mkdir(parents=True, exist_ok=True) - report = build_report(repo_path=repo_path, findings=findings) + report = build_report(repo_path=repo_path, findings=findings, coverage=coverage) json_path = reports_dir / "security_report.json" markdown_path = reports_dir / "security_report.md" write_json_report(report, json_path) write_markdown_report(report, markdown_path) + paths = {"json": json_path, "markdown": markdown_path} + + if coverage is not None: + coverage_path = reports_dir / "coverage.json" + write_json_report( + { + "repository": report["repository"], + "generated_at": report["generated_at"], + **report["coverage"], + }, + coverage_path, + ) + paths["coverage"] = coverage_path - return {"json": json_path, "markdown": markdown_path} + return paths diff --git a/scanner/run_scan.py b/scanner/run_scan.py index 9000bdd..2296821 100644 --- a/scanner/run_scan.py +++ b/scanner/run_scan.py @@ -9,6 +9,7 @@ if __package__ in {None, ""}: sys.path.insert(0, str(Path(__file__).resolve().parents[1])) +from scanner.coverage import build_coverage from scanner.git_source import GitCloneError, cloned_repo from scanner.ignore import ( DEFAULT_SAMPLE_IGNORE_PATTERNS, @@ -48,10 +49,20 @@ def run_scan( ignore_patterns = [*DEFAULT_SAMPLE_IGNORE_PATTERNS, *ignore_patterns] findings = collect_findings(resolved_repo, reports_dir) + # Coverage reflects what the scanners reported, before any suppression: + # `--ignore-file` does not yet exempt meta findings, so a pattern matching + # the repository root could otherwise hide the finding that says a tool + # never ran. + coverage = build_coverage(findings) findings = rebase_finding_paths(findings, resolved_repo) findings = apply_ignore(findings, ignore_patterns) findings = filter_by_min_severity(findings, min_severity) - return write_reports(repo_path=resolved_repo, findings=findings, reports_dir=reports_dir) + return write_reports( + repo_path=resolved_repo, + findings=findings, + reports_dir=reports_dir, + coverage=coverage, + ) def run_scan_from_url( @@ -143,6 +154,8 @@ def main(argv: list[str] | None = None) -> int: print(f"JSON report: {report_paths['json']}") print(f"Markdown report: {report_paths['markdown']}") + if "coverage" in report_paths: + print(f"Coverage report: {report_paths['coverage']}") return 0 diff --git a/tests/test_coverage.py b/tests/test_coverage.py new file mode 100644 index 0000000..5097053 --- /dev/null +++ b/tests/test_coverage.py @@ -0,0 +1,292 @@ +"""Tests for scan coverage reporting. + +A report that says "3 findings" while two of five scanners never ran is true +and misleading at the same time. These tests pin the property that makes the +difference visible: every configured scanner gets a row, and a tool that did +not examine what it was pointed at cannot be silently absent. +""" + +from __future__ import annotations + +import json +from pathlib import Path + +from scanner.coverage import ( + COVERAGE_STATUSES, + EXPECTED_TOOLS, + ToolCoverage, + build_coverage, + coverage_payload, + incomplete_tools, + is_complete, + status_for_meta_finding, +) +from scanner.ignore import apply_ignore +from scanner.models import Finding +from scanner.report import build_report, write_reports +from scanner.scanners import SCANNERS +from scanner.tools.gitleaks_runner import GitleaksResult +from scanner.tools.semgrep_runner import SemgrepResult +from scanner.tools.trivy_runner import TrivyResult + +AI_REVIEW_ENV_VARS = ( + "AI_PATCHLAB_AI_REVIEW_ENABLED", + "AI_PATCHLAB_AI_REVIEW_PROVIDER", + "AI_PATCHLAB_AI_REVIEW_COMMAND", + "AI_PATCHLAB_AI_REVIEW_TIMEOUT_SECONDS", +) + + +def _meta(finding_id: str, tool: str, title: str = "Example state", **overrides: object) -> Finding: + """Build a meta finding with defaults for the field under test.""" + base: dict[str, object] = { + "id": finding_id, + "tool": tool, + "severity": "info", + "title": title, + "description": "Example description.", + "file": "app.py", + "line": None, + "recommendation": "Install the tool.", + "confidence": "high", + "is_meta": True, + } + base.update(overrides) + return Finding(**base) # type: ignore[arg-type] + + +def _stub_external_tools(monkeypatch) -> None: + """Report every external scanner as not installed, without touching PATH.""" + for var in AI_REVIEW_ENV_VARS: + monkeypatch.delenv(var, raising=False) + monkeypatch.setattr( + "scanner.scanners.gitleaks.run_gitleaks", + lambda repo_path, raw_report_path: GitleaksResult( + installed=False, raw_report_path=raw_report_path + ), + ) + monkeypatch.setattr( + "scanner.scanners.semgrep.run_semgrep", + lambda repo_path, raw_report_path: SemgrepResult( + installed=False, raw_report_path=raw_report_path + ), + ) + monkeypatch.setattr( + "scanner.scanners.trivy.run_trivy", + lambda repo_path, raw_report_path: TrivyResult( + installed=False, raw_report_path=raw_report_path + ), + ) + + +class TestStatusDerivation: + """Each meta finding id maps to exactly one coverage state.""" + + def test_not_installed_means_not_run(self) -> None: + assert status_for_meta_finding("semgrep-not-installed") == "not_run" + + def test_disabled_means_not_run(self) -> None: + assert status_for_meta_finding("ai-review-disabled") == "not_run" + + def test_scan_error_means_error(self) -> None: + assert status_for_meta_finding("semgrep-scan-error") == "error" + + def test_json_parse_error_means_error(self) -> None: + assert status_for_meta_finding("trivy-json-parse-error") == "error" + + def test_command_error_means_error(self) -> None: + assert status_for_meta_finding("ai-review-command-error") == "error" + + def test_partial_coverage_means_partial(self) -> None: + assert status_for_meta_finding("semgrep-partial-coverage") == "partial" + + def test_no_supported_manifest_means_partial(self) -> None: + assert status_for_meta_finding("dependency-scan-no-supported-manifest") == "partial" + + def test_unknown_meta_finding_defaults_to_partial(self) -> None: + """An unmodelled scan-state signal must not read as full coverage.""" + assert status_for_meta_finding("semgrep-something-new") == "partial" + + def test_every_status_is_declared(self) -> None: + derived = { + status_for_meta_finding(finding_id) + for finding_id in ( + "x-not-installed", + "x-scan-error", + "x-partial-coverage", + "x-unknown", + ) + } + assert derived.issubset(set(COVERAGE_STATUSES)) + + +class TestBuildCoverage: + """Every configured scanner gets exactly one row, in registry order.""" + + def test_every_expected_tool_gets_a_row(self) -> None: + coverage = build_coverage([]) + assert tuple(row.tool for row in coverage) == EXPECTED_TOOLS + + def test_a_tool_that_reported_nothing_ran(self) -> None: + """This is the whole point: silence means success, and must be stated.""" + coverage = build_coverage([]) + assert all(row.status == "ran" for row in coverage) + + def test_meta_finding_sets_the_status_and_detail(self) -> None: + findings = [_meta("semgrep-not-installed", "semgrep", title="Semgrep is not installed")] + row = build_coverage(findings)[0] + assert row.tool == "semgrep" + assert row.status == "not_run" + assert row.detail == "Semgrep is not installed" + assert row.meta_finding_ids == ("semgrep-not-installed",) + + def test_worst_status_wins(self) -> None: + findings = [ + _meta("semgrep-partial-coverage", "semgrep"), + _meta("semgrep-not-installed", "semgrep"), + ] + row = build_coverage(findings)[0] + assert row.status == "not_run" + assert len(row.meta_finding_ids) == 2 + + def test_ordinary_findings_do_not_affect_coverage(self) -> None: + findings = [ + _meta("sql-injection", "semgrep", severity="high", is_meta=False, confidence="medium") + ] + assert build_coverage(findings)[0].status == "ran" + + def test_unknown_tool_is_ignored(self) -> None: + findings = [_meta("mystery-not-installed", "mystery-tool")] + coverage = build_coverage(findings) + assert tuple(row.tool for row in coverage) == EXPECTED_TOOLS + + def test_rejects_an_unsupported_status(self) -> None: + try: + ToolCoverage(tool="semgrep", status="maybe", detail="") + except ValueError as exc: + assert "maybe" in str(exc) + else: # pragma: no cover - guard + raise AssertionError("ToolCoverage accepted an undeclared status") + + +class TestCompleteness: + """`complete` is the flag a reader trusts, so its edges matter.""" + + def test_all_ran_is_complete(self) -> None: + assert is_complete(build_coverage([])) is True + + def test_a_missing_tool_is_incomplete(self) -> None: + coverage = build_coverage([_meta("semgrep-not-installed", "semgrep")]) + assert is_complete(coverage) is False + assert [row.tool for row in incomplete_tools(coverage)] == ["semgrep"] + + def test_disabled_ai_review_alone_stays_complete(self) -> None: + """AI review is off by default (ADR-010) - a banner that always fires is ignored.""" + coverage = build_coverage([_meta("ai-review-disabled", "ai-security-review")]) + assert is_complete(coverage) is True + + def test_disabled_ai_review_does_not_mask_a_real_gap(self) -> None: + coverage = build_coverage( + [ + _meta("ai-review-disabled", "ai-security-review"), + _meta("semgrep-not-installed", "semgrep"), + ] + ) + assert is_complete(coverage) is False + + def test_ai_review_command_error_is_a_real_gap(self) -> None: + coverage = build_coverage([_meta("ai-review-command-error", "ai-security-review")]) + assert is_complete(coverage) is False + + def test_payload_counts_match_the_rows(self) -> None: + payload = coverage_payload(build_coverage([_meta("trivy-scan-error", "trivy")])) + assert payload["complete"] is False + assert payload["incomplete_tool_count"] == 1 + assert payload["tool_count"] == len(EXPECTED_TOOLS) + assert len(payload["tools"]) == len(EXPECTED_TOOLS) + + +class TestIgnoreImmunity: + """`--ignore-file` must not be able to hide the fact that a tool never ran.""" + + def test_coverage_survives_a_pattern_that_suppresses_the_meta_finding(self) -> None: + findings = [_meta("semgrep-not-installed", "semgrep", file="app.py")] + coverage = build_coverage(findings) + + surviving = apply_ignore(findings, ["*.py"]) + + assert surviving == [], "precondition: the pattern must suppress the finding" + assert coverage[0].status == "not_run" + + +class TestRegistryDrift: + """Adding a scanner without adding it to coverage must fail here.""" + + def test_expected_tools_matches_the_scanner_registry(self, tmp_path: Path, monkeypatch) -> None: + _stub_external_tools(monkeypatch) + repo_path = tmp_path / "empty-repo" + repo_path.mkdir() + reports_dir = tmp_path / "reports" + reports_dir.mkdir() + + tools = { + finding.tool + for scanner in SCANNERS + for finding in scanner(repo_path, reports_dir) + if finding.is_meta + } + + assert tools == set(EXPECTED_TOOLS) + + +class TestReportRendering: + """The banner is the part a reader cannot miss, so it is pinned.""" + + def test_no_coverage_key_when_not_supplied(self, tmp_path: Path) -> None: + report = build_report(repo_path=tmp_path, findings=[]) + assert "coverage" not in report + + def test_banner_is_rendered_when_incomplete(self, tmp_path: Path) -> None: + coverage = build_coverage([_meta("semgrep-not-installed", "semgrep")]) + paths = write_reports( + repo_path=tmp_path, + findings=[], + reports_dir=tmp_path / "reports", + coverage=coverage, + ) + markdown = paths["markdown"].read_text(encoding="utf-8") + assert "**Incomplete scan.**" in markdown + assert "1 of 5 tools" in markdown + assert "## Scan Coverage" in markdown + assert "`not_run`" in markdown + + def test_no_banner_when_complete(self, tmp_path: Path) -> None: + paths = write_reports( + repo_path=tmp_path, + findings=[], + reports_dir=tmp_path / "reports", + coverage=build_coverage([]), + ) + markdown = paths["markdown"].read_text(encoding="utf-8") + assert "Incomplete scan" not in markdown + assert "## Scan Coverage" in markdown + + def test_coverage_json_is_written(self, tmp_path: Path) -> None: + paths = write_reports( + repo_path=tmp_path, + findings=[], + reports_dir=tmp_path / "reports", + coverage=build_coverage([_meta("trivy-not-installed", "trivy")]), + ) + assert paths["coverage"].name == "coverage.json" + payload = json.loads(paths["coverage"].read_text(encoding="utf-8")) + assert payload["complete"] is False + assert payload["repository"] + assert payload["generated_at"] + assert {row["tool"] for row in payload["tools"]} == set(EXPECTED_TOOLS) + + def test_no_coverage_file_when_not_supplied(self, tmp_path: Path) -> None: + reports_dir = tmp_path / "reports" + paths = write_reports(repo_path=tmp_path, findings=[], reports_dir=reports_dir) + assert "coverage" not in paths + assert not (reports_dir / "coverage.json").exists()