This file captures the Codex/OpenAI workflow for the project. It should stay aligned with
CLAUDE.mdwhen both runtimes are present.
AI PatchLab is an AI-assisted security remediation toolkit. The MVP focuses on a local Python scanner that accepts a repository path, normalizes security findings from Semgrep, Gitleaks, Trivy, pip-audit, and an opt-in local AI review, and writes JSON plus Markdown reports for remediation planning.
- Treat
INITIAL.md,PRPs/,ROADMAP.md,DECISIONS.md,README.md, andexamples/as shared sources of truth across runtimes - If
CLAUDE.mdexists, keep it aligned with structural, dependency, and workflow changes made through Codex - Prefer adding Codex behavior in
.agents/skills/instead of changing Claude-specific files unless parity requires it
- Always read
ROADMAP.mdat the start of a new conversation to understand the current state - Check
DECISIONS.mdfor past architectural decisions and context - Check
PRPs/for active or completed implementation plans - Read
examples/before implementing new features to match existing patterns - Check
INITIAL.mdfor the current feature requirements - Use the
statusskill at the start of a session for a quick project snapshot
- Python 3.11+
- Local CLI scanner foundation (Semgrep, Gitleaks, Trivy, pip-audit, opt-in local AI review)
- JSON and Markdown report generation
- Data stack selected for repository analysis workflows
- MySQL 8.0 (aiomysql for async, available but not required in v0.1)
- Playwright (optional extra
[scraping], if scraping is needed later) - Discord webhooks (alerts, optional later)
- Loguru (logging)
- pytest + pytest-asyncio + pytest-mock (testing)
- ruff + black (linting/formatting)
- httpx (async HTTP client)
- pydantic + pydantic-settings + python-dotenv (validation + config)
- Standard-library
subprocessfor external scanner runners (no remote endpoints)
fingerprint/- Web template fingerprinting (v0.1, experimental): seed loader, extractors, indexer CLI, web probe, matchers, scoring, JSON+MD match report. Local-first, single-target probe, probabilistic signal only.fingerprint/models.py- Frozen dataclasses (RepoFingerprint,AssetFingerprint,HtmlSignature,MatchResult,MatchSignal) and the canonicalband_for_score()helperfingerprint/config.py-FingerprintConfig(pydantic-settings,AI_PATCHLAB_FINGERPRINT_env prefix)fingerprint/git_seeds.py- Loader + validator forfingerprint/seeds/repos.json+slug_from_repo_urlfingerprint/seeds/repos.json- Curated seed list (committed; expand via PR only)fingerprint/extractors/- Pure functions over a cloned repo:favicon.py,static_assets.py,html_signatures.pyfingerprint/repo_index.py- Clones viascanner.git_source.cloned_repo, runs extractors, writesfingerprint/db/<slug>.jsonfingerprint/run_index.py- CLI:.venv/Scripts/python.exe fingerprint/run_index.py --rebuild/--repo-url <url>fingerprint/web_probe.py- Synchttpx.Clientprobe with robots.txt respect, scheme allowlist, bytes/asset capsfingerprint/matchers/-asset_hash.py,html_regex.py; registered infingerprint/matchers/__init__.py:MATCHERSfingerprint/scoring.py- Bounded weighted score (WEIGHT_VALUES); sharedband_for_scorehelper from modelsfingerprint/run_match.py- CLI:.venv/Scripts/python.exe fingerprint/run_match.py --target <url>->reports/fingerprint/match_<host>_<UTC>.json+.mdfingerprint/report.py- JSON + Markdown writer; disclaimer block is mandatoryfingerprint/db/- Per-repo fingerprint JSONs (gitignored except.gitkeep)reports/fingerprint/- Generated match reportsscanner/- Scanner CLI, finding model, recommendation enrichment, report generation, scanner registryscanner/run_scan.py- CLI entry point (.venv/Scripts/python.exe scanner/run_scan.py --repo <path>or--from-git-url <url>)scanner/git_source.py- Shallow-clone a public git URL into a temp directory via thecloned_repocontext manager; cleanup-on-exit,shell=False, no remote API callsscanner/paths.py-rebase_finding_paths(findings, repo_root)rewrites each finding'sfile(andidwhen it embeds the same path) to a repo-relative POSIX path so reports survive temp-dir cleanupscanner/ignore.py-apply_ignore(findings, patterns)+load_ignore_patterns(path)provide.gitignore-style path suppression (used by the--ignore-fileCLI flag). Empty-file findings are never suppressed.DEFAULT_SAMPLE_IGNORE_PATTERNSholds demo/sample/example subtree patterns opted into via--ignore-samplesscanner/models.py- NormalizedFindingdataclass + severity/confidence enums +FINDING_FIELDSscanner/recommendations.py- Deterministic keyword-based recommendation enrichmentscanner/coverage.py- Per-scanner coverage derived from meta findings (ToolCoverage,build_coverage,EXPECTED_TOOLS,is_complete); feedsreports/coverage.jsonand the report's Scan Coverage blockscanner/verdict_corpus.py- Corpus validation and aggregation (validate_payloadreports every problem in one pass,load_corpusreadscorpus/verdicts/)scanner/run_verdict_check.py- CLI:--check <verdicts.json>(exit 2 on any problem, required by/dailyPhase 4) and--summary(count the whole corpus)corpus/verdicts/- Committed dismissal corpus, one file per scan; the durable half, sincereports/is gitignoredscanner/verdicts.py- Dismissal records (VerdictRecord,summarize_removed,count_by_reason,load_records) with closedSCANNER_REASON_CODES/CURATION_REASON_CODESvocabularies; feedsreports/verdicts.jsonand the report's Dismissed sectionscanner/report_markdown.py- Markdown rendering, split out ofreport.pyto stay under the 300-line ceiling (write_markdown_reportis still re-exported fromscanner.report)scanner/confidence.py- CentralizedFinding.confidencerules (one function per scanner +confidence_for_meta_findingfor sharednot-installed/scan-error/ etc.)scanner/report.py- JSON + Markdown report writers (severity-grouped, "Top Findings" highlight block, patch suggestion blocks); also exposesfilter_by_min_severityandselect_top_findingsscanner/config.py- Disabled-by-default AI review configuration loaded from environment /.env(AI_PATCHLAB_*)scanner/remediation/- Deterministic patch suggestion engine (patch_suggestions.py) for known vulnerability patternsscanner/scanners/- Scanner adapters (semgrep.py,gitleaks.py,trivy.py,dependency_scan.py,ai_review.py) pluscommon.pyplaceholder helper and__init__.pyregistry (SCANNERS)scanner/tools/- External scanner process runners (semgrep_runner.py,gitleaks_runner.py,trivy_runner.py,pip_audit_runner.py,ai_review_runner.py)reports/- Generated security reports (security_report.json,security_report.md)reports/verdicts.json- Dismissal corpus: one row per rule family removed, with a reason code and a countreports/coverage.json- Per-scanner coverage manifest written on every scan (what each tool actually examined, plus acompleteflag)reports/raw/- Raw scanner JSON outputs (semgrep.json,gitleaks.json,trivy.json,pip-audit.json,ai-review.jsonwhen enabled)src/- Legacy scaffold entry point (kept for template parity)src/main.py- Legacy entry point (python -m src.main) - currently a loguru-wired async stub with TODOs.github/workflows/ci.yml- CI: ruff + black + pytest on Python 3.11 and 3.13reports/disclosures/- Drafted private disclosure emails awaiting a manual send (gitignored)tests/- pytest tests (one module per scanner:test_scanner_foundation.py,test_semgrep_scanner.py,test_gitleaks_scanner.py,test_trivy_scanner.py,test_dependency_scan.py,test_ai_review.py,test_patch_suggestions.py,test_recommendations.py,test_meta_findings.py,test_confidence_field_rules.py,test_coverage.py,test_verdicts.py,test_verdict_corpus.py)tests/conftest.py- Shared fixtures (mock_db,mock_http_client,mock_discord,test_config, sessionevent_loop)examples/- Reference patterns to read before implementingPRPs/- Active Product Requirements PromptsPRPs/done/- Archived PRPs (2026-05-12-trivy-integration.md,20260513-phase-3-ai-review-behavior.md)PRPs/templates/- Reusable PRP templates (prp_base.md)docs/- GitHub Pages site (Jekyll, themecayman):_config.yml,index.md(landing + scan log),scans/(per-scan write-ups),templates/scan-post.md(template, excluded from publish)logs/- Log files (gitignored except.gitkeep).agents/skills/- Codex skills for repeatable workflows.claude/- Claude runtime files (commands, agents, pipelines), if present
ez-project-workflow- Core EzProject operating rules for any coding taskkickoff- Interview flow that updatesINITIAL.mdand project docsgenerate-prp- Build a self-contained PRP from a feature spec (with optional Understanding Lock for ambiguous specs)execute-prp- Implement a PRP end to end with validation and housekeepingfix-issue- Diagnose, fix, test, review, and commit a bug end-to-endpipeline- Execute a YAML-defined workflow step by step (feature, bugfix, security, release, custom)tdd- Strict Red-Green-Refactor implementation (Iron Law: no production code without an observed-failing test)review-code- Review the current diff or target filesrefactor- Analyze code smells and execute selected refactorings with per-change validationretrospective- Self-healing retrospective in a fresh context afterexecute-prp(lite mode) or sprint-wide (no arg)next- Read project state and recommend the 1-3 best next actionsupgrade-status- Compare project to latest template; list features not yet adoptedstatus- Produce a compact project status snapshotaudit-project- Audit runtime docs and project scaffoldinghousekeeping- Sync docs and remove temporary artifactscleanup- Identify dead code via import-graph trace; report and ask before deletingsecurity-scan- OWASP top 10 + secrets + vulnerable deps auditdependency-check- Vulnerabilities, outdated packages, compatibility audit; optional safe auto-updateperformance- Static hot-path scan + optional runtime profilingdocument- Auto-generate docs from code (API ref, data dict, architecture, modules)monitor-setup- Health check endpoint + alerting + uptime tracking adapted to stackrollback- Safe git rollback (always revert + stash, never reset --hard)create-skill- Scaffold a new Codex skill for the projectdaily- Autonomous daily scan-and-disclose pipeline (status sweep -> candidate discovery -> one scan -> curation -> gated publication). Full-auto per the 2026-05-28 decision; guardrails: 1 scan/day, quality gate on issue/PR filing, strict-norm detection, kill switch.daily-paused
These Claude slash commands are intentionally not mirrored on the Codex side. Listed in tests/template/test_codex_parity.py as KNOWN_EXCEPTIONS.
/do(smart-router) - natural-language intent routing requires a Claude-side primitive that Codex doesn't expose; the pattern table at.claude/commands/smart-router.mdis Claude-only./idea-to-pr- end-to-end idea -> PR flow requires PR/branching orchestration that's tightly coupled to Claude's tool layer./upgrade-to-project- one-shot MVP -> Project upgrade; the equivalent on the Codex side is to invokeez-upgrade-project.ps1directly.
# Dev
.venv/Scripts/python.exe scanner/run_scan.py --repo "C:\path\to\repo"
.venv/Scripts/python.exe scanner/run_scan.py --repo "." --reports-dir reports
python -m src.main
.venv/Scripts/python.exe -m pytest tests/ -v
.venv/Scripts/python.exe -m pytest tests/ -v -k "test_name"
# Lint & Format
.venv/Scripts/ruff.exe check scanner src/ tests/ fingerprint/
.venv/Scripts/ruff.exe check scanner src/ tests/ fingerprint/ --fix
.venv/Scripts/python.exe -m black scanner src/ tests/ fingerprint/
# Web template fingerprinting (experimental)
.venv/Scripts/python.exe fingerprint/run_index.py --rebuild # Rebuild local DB from seed list
.venv/Scripts/python.exe fingerprint/run_match.py --target https://example.com # Probe a single live URL
# Setup
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pip install -e ".[scraping]" # Optional Playwright extra
cp .env.example .env
# External scanners (install separately, must be on PATH)
semgrep --version
gitleaks version
trivy --version
.venv/Scripts/python.exe -m pip install pip-audit && pip-audit --version
# Optional AI review (disabled by default - see scanner/config.py)
export AI_PATCHLAB_AI_REVIEW_ENABLED=true
export AI_PATCHLAB_AI_REVIEW_PROVIDER=local_command
export AI_PATCHLAB_AI_REVIEW_COMMAND=/path/to/ai-review-wrappermemory- Session persistence when configuredmysql- Direct database access through MCP whenMYSQL_DSNis configured
- Type hints REQUIRED on all functions
- Google-style docstrings on public functions
- Max ~300 lines per file (note:
scanner/remediation/patch_suggestions.pyis ~280 lines) - English identifiers; French comments are fine
- Use
async/awaitwhere appropriate; the scanner core is synchronous on purpose because it shells out viasubprocess.run - Use loguru, never
print()in production code (exception:scanner/run_scan.pyprints report paths to stdout as the CLI surface) - Config via
.env/AI_PATCHLAB_*env vars; never hardcode secrets - Use
pathlib.Pathfor filesystem paths - Use
pydanticfor env-loaded config; use@dataclass(frozen=True)for internal normalized records (Finding,*Result) - External scanner subprocesses MUST use
shell=Falseand an explicit argv list
- Each scanner in
scanner/scanners/exposes onescan_<name>(repo_path: Path, reports_dir: Path) -> list[Finding]function and is registered inscanner/scanners/__init__.py:SCANNERS - Scanners must never raise on a missing or failing external tool - emit a normalized
infofinding instead so the report still completes - Each adapter pipes its findings through
apply_patch_suggestions(enrich_findings(...))before returning - Any finding built with
confidence_for_meta_finding(...)must also setis_meta=Trueso--min-severitycannot drop it - Each external tool runner lives in
scanner/tools/<tool>_runner.py, returns a frozen*Resultdataclass, writes raw JSON toreports/raw/<tool>.json, and usessubprocess.run(..., shell=False, check=False)with captured stdout/stderr - Do not call subprocesses directly from
scanner/scanners/*- go through the runner module - Every suppression step in
run_scanrecords what it removed viasummarize_removed(before, after, reason_code)- a narrower report must never shrink its own numbers silently (ADR-016) reason_codevalues are a closed vocabulary inscanner/verdicts.py; add a code to the module rather than inventing one at a call site, or the corpus stops being countable- Curation rows are hand-written JSON, so the dataclass never runs on them -
/dailyPhase 4 MUST get exit 0 fromscanner/run_verdict_check.py --checkbefore publishing (ADR-017).ruleis required on every row;verdicttakes one of five values and the sentence goes indetail confirmed-realis first-party only; an upstream dependency merely behind a fixed version isdependency-currency- Coverage is derived from the raw
collect_findingsoutput beforeapply_ignore(scanner/run_scan.py) ---ignore-filedoes not exempt meta findings, so deriving it later would let a path pattern hide the fact that a tool never ran - Adding a scanner to
SCANNERSrequires adding itsFinding.toolvalue toscanner/coverage.py:EXPECTED_TOOLS;tests/test_coverage.py::TestRegistryDriftfails until you do Finding.confidencevalues come fromscanner/confidence.py- never inlineconfidence="high"/"medium"/"low"in a scanner adapter; add or reuse a rule function instead
- Extractors in
fingerprint/extractors/are pure functions over a cloned repo path (and the activeFingerprintConfig). No subprocess, no network, no raise on missing files. - Matchers in
fingerprint/matchers/take(RepoFingerprint, TargetSnapshot)and returnlist[MatchSignal]. Registered infingerprint/matchers/__init__.py:MATCHERS. A mismatch is the empty list. - The indexer (
fingerprint/repo_index.py:index_seed) is the only place that performs a git clone; it goes throughscanner.git_source.cloned_repo. - The web probe (
fingerprint/web_probe.py:fetch_target) is the only place that makes outbound HTTP requests. Scheme allowlist:http/https. Honoursrobots.txtand hard byte/asset caps. - The match CLI (
fingerprint/run_match.py) always exits 0 (partial-result discipline). Empty DB, unreachable target, bad scheme, robots-disallowed all still produce a valid report. - The Markdown report ALWAYS includes the
DISCLAIMERblock and never uses attribution words ("confirmed", "proven", "stolen", "copied") - tested intests/test_fingerprint_report.py. - Score banding goes through
band_for_score()infingerprint/models.py- single source of truth used by validator and writer. - The seed list
fingerprint/seeds/repos.jsonis curated. Adding a new entry is a human PR - never auto-discover via the GitHub API.
- MySQL tables in
snake_case - Always include
id,created_at, andupdated_at - Index frequently queried columns
- Append migrations to
docs/schema.sqlwith a dated comment (file not yet created -docs/is empty) - Use parameterized queries only
- Wrap external calls in
try/except - Log errors with context using loguru
- Use retries with backoff for network calls when needed
- Never silently suppress errors
- Use custom exception classes for domain-specific failures
- Scanner runners catch
OSErrorandsubprocess.TimeoutExpired, write a safe empty raw JSON, and return a structured*Resultinstead of propagating
- Use pytest for all tests (
pytest-asynciomode isautoperpyproject.toml) - Mock external calls (API, DB, subprocess) in unit tests
- Cover critical calculation and workflow paths
- Use
pytest-asynciofor async tests - Available markers:
slow,integration - Each scanner adapter has a dedicated
tests/test_<scanner>.pymodule
- Create a branch per feature
- Use descriptive English commit messages
- Format commits as
type: description - Never commit directly to
main - Run validation before committing
- Always check
ROADMAP.mdfirst - Progression:
[ ]todo ->[-]in progress ->[x]done - Add a
YYYY/MM/DDtimestamp when status changes - Update
ROADMAP.mdafter completing each meaningful task
- Log architectural decisions in
DECISIONS.md - Check existing ADRs before making structural changes
- Record date, decision, context, and consequences
- Current ADRs of record: ADR-001 scaffold, ADR-002 data stack, ADR-003 placeholder adapters, ADR-004 Gitleaks, ADR-005 Semgrep, ADR-006 recommendation enrichment, ADR-007 patch suggestions, ADR-008 Trivy, ADR-009 pip-audit, ADR-010 disabled-by-default AI review boundary, ADR-011 centralized scanner confidence rules, ADR-012 probabilistic web template fingerprinting boundary, ADR-013 meta findings exempt from severity filtering, ADR-014 field-derived confidence tiers, ADR-015 coverage is a report artifact, ADR-016 dismissals recorded as counted rule families, ADR-017 a closed vocabulary needs a closer
- Do not tag scan posts by keyword inference. A classifier over post bodies was built and rejected 2026-09-18: validated against 9 posts of known ground truth it gave klavis 6 finding classes where the real finding was dependency CVEs, and got 3 of 9 project families wrong (OpenBiliClaw as "developer tooling", tracecat as "MCP server"). Post bodies discuss false positives and credited defences at length, so matching them tags a clean scan with the class it dismissed. On a site whose whole argument is that pattern-matching produces plausible-but-wrong results, publishing plausible-but-wrong tags is self-refuting. Grouping pages are generated from the curated index table instead — hand-maintained, verified data
- Roughly half of scan targets land on a disclosure channel the pipeline cannot use. Measured 2026-09-18: 14 disclosures went through an autonomous channel (GitHub private vulnerability reporting, or a public issue), 12 needed a human to send an email or a DM. Every pipeline stoppage in the project's history came from the human-gated half — a six-report backlog (25 days, 4 scans lost) and one undeliverable High (33 days, 3 scans lost).
/dailyguardrail 9 now makes channel viability a target-SELECTION criterion, tiered on queue depth, and guardrail 8 definesunreachableas the only way to close a report that no channel can carry. The tiering is deliberate: PVR-enabled repos skew mature and commercially backed, so a permanent preference would bias the series away from the small projects that most need review - Semgrep is a Python program on the shared user-site interpreter, not a standalone binary like gitleaks/trivy. Anything that breaks that interpreter breaks Semgrep too — a pydantic downgrade on 2026-08-20 made
semgrep --versionraise ImportError and every scan would have silently lost 52% of its coverage. The project.venvdoes NOT protect it. Checksemgrep --versionbefore trusting a scan; repair withpython -m pip install --user --upgrade "pydantic>=2.11" "httpx>=0.27" - ALWAYS run through
.venv(.venv/Scripts/python.exeon Windows). The project ran three months off the shared user-site; on 2026-08-20 an unrelatedpip installdowngraded pydantic to 1.x and httpx to 0.21 and every import broke, hours after a scan had passed - Meta findings survive
--min-severitybut are NOT yet exempt from--ignore-filesuppression - Semgrep coverage comes from the
errorsarray, neverpaths.skipped-skippedhas been empty on every series run where rules timed out.scan_semgrepemitssemgrep-partial-coveragenaming eachrule -> filepair that did not run ai-patchlaband2026-05-12are template placeholders replaced during scaffoldingaiomysqlpools must be closed explicitly withawait db.disconnect()infinally- Playwright
networkidlecan time out on SPAs; usedomcontentloadedwhen needed - Discord webhooks: 30 messages/minute rate limit per webhook
- pydantic-settings:
.envvariables are case-insensitive by default - MCP MySQL DSNs require URL-encoding for special characters in passwords
- cPanel MySQL DB and user names are prefixed (e.g.
cpaneluser_dbname) - don't forget the prefix - AI review must remain disabled by default and local-first. Never add a default remote provider, default endpoint, default model, or default token variable. Any future remote/paid provider requires explicit configuration and a new ADR.
- Scanner subprocess invocations MUST use
shell=Falseand an explicit argv list -shell=Trueis the exact anti-pattern the patch engine warns about (seescanner/remediation/patch_suggestions.py:SUBPROCESS_SHELL_SUGGESTION) - Semgrep on Windows: when not on
PATH, the runner falls back toPath.home() / AppData/Roaming/Python/Python313/Scripts/semgrep.exe(seescanner/tools/semgrep_runner.py:PIP_USER_SEMGREP_PATH). The Python minor version is hardcoded - if the user installs under a different Python version, the runner will silently skip Semgrep until the constant is updated - Semgrep UTF-8 output (Windows, 2026-06-11 fix): Semgrep writes its
--outputJSON via Python's default codec (cp1252 on Windows). A repo with non-Latin-1 source (Chinese/Japanese/Korean/emoji) crashes Semgrep mid-write withUnicodeEncodeError, leaving a 0-byte report + exit 2._build_semgrep_envforcesPYTHONUTF8=1+PYTHONIOENCODING=utf-8;scan_semgreptreats an empty report as asemgrep-scan-error. The scan-error isinfoseverity, so--min-severity mediumstill filters it (RESOLVED 2026-08-21: meta findings carryis_meta=Trueand are exempt from--min-severity) - All external scanners are optional - if a tool is missing, the adapter emits a normalized
infofinding (semgrep-not-installed,gitleaks-not-installed,trivy-not-installed,pip-audit-not-installed,ai-review-disabled) instead of failing - pip-audit input resolution order: root
requirements*.txtfirst, thenpylock.*.toml(with--locked), thenpyproject.toml - AI review timeouts default to 120s (
AI_PATCHLAB_AI_REVIEW_TIMEOUT_SECONDS); on timeout the runner writes[]toreports/raw/ai-review.jsonand emits a normalized error finding so the report still completes - Fingerprint web probe respects
robots.txt.urllib.robotparsersplits the live user-agent on/before matching, so a robots.txt block forai-patchlab-fingerprint/0.1is matched asai-patchlab-fingerprint(drop the version when authoring rules) - Fingerprint matching is a SIGNAL, not an attribution. The Markdown report keeps the
DISCLAIMERblock and never claims a match is "confirmed", "proven", "stolen", or "copied" - blocked by a regression test - The fingerprint CLI accepts exactly one
--targetper invocation. No--targets-fileflag exists by design - multi-target scanning requires a new ADR - Fingerprinting must not add DOM-parser dependencies (
beautifulsoup4,lxml) or browser stacks (playwright,selenium). v0.1 stays onre+hashlib+httpxonly
- Keep things simple - MVP first
- Read
examples/before inventing new patterns - Validate each step before moving on
- Do not over-engineer
- Log meaningful architecture decisions in
DECISIONS.md
When an implementation is painful, fix the underlying AI layer - not just the code. The AI layer is CLAUDE.md, AGENTS.md, examples/, .claude/commands/, .claude/agents/, .claude/pipelines/, .agents/skills/, PRP templates.
- Never review the implementer's work in the same context that produced it (writer is biased about its own output)
- After
execute-prp, in a NEW session, run theretrospectiveskill with the PRP name (lite mode) to surface AI-layer drift and propose concrete file edits - Apply the proposed edits only after the user approves
- Periodically run
retrospectivewith no argument for a broader sprint review
- After any feature or fix, update
ROADMAP.md - After structural changes, update
AGENTS.md - If
CLAUDE.mdexists, update the matching sections there too - If new setup steps or commands were added, update
README.md - After completing a PRP, move it to
PRPs/done/with a date prefix - Delete scratch files and obvious temporary artifacts when done
- Never leave documentation out of sync with the code