feat(archetypes): Analyst — answers questions from your local data with SQL and one chart (held) - #4026
Conversation
…th SQL and one chart (held) Adds the `analyst` archetype: answers the operator's questions from their own data files (CSV, TSV, Parquet, JSON, Excel, SQLite) via the data plugin's read-only DuckDB, with one live Vega-Lite chart and a one-line takeaway. It says what it assumed, cites the source file, never invents a number (derived figures included), and keeps answers short. On first run with no data folders, it tells the operator where to set them (Settings ▸ Plugins ▸ Data Analyst ▸ Data folders) and never tries to set them itself. - config/soul-presets/analyst.md: the persona. - archetype-catalog.json: the row is parked in `held` (Josh tests before the picker). - plugin-directory.yaml: `archetype_repos` registers protoLabsAI/analyst-archetype. - tests: test_analyst_archetype_is_held pins the held state and the persona's rules. - docs/guides/fleet.md: names Analyst among the held rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 19 minutes. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (6)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
QA panel review — PASS
code-review-structural · head a7371ec7ed0b · formal
Purely additive PR adding a held "Analyst" archetype across catalog, soul preset, plugin directory, docs, and a pinning test. The panel raised a single nit (protopatch) on a missing encoding='utf-8' in the new test's read_text() calls, which the verifier confirmed against the project's own PLW1514 ruff rule. No blockers, no majors, no disagreements. Verification covered all 1 finding; no gaps.
Findings
| Severity | Location | Finding | Verified | |
|---|---|---|---|---|
| ⚪ | nit | tests/test_bundled_config_assets.py:423 |
read_text() without encoding='utf-8' contradicts the codebase's own PLW1514 guardrail enforced via ruff. | confirmed |
findings JSON (machine-readable)
[
{
"file": "tests/test_bundled_config_assets.py",
"line": 423,
"severity": "nit",
"category": "conventions",
"claim": "read_text() without encoding='utf-8' contradicts the codebase's own PLW1514 guardrail enforced via ruff.",
"evidence": "Three read_text() calls in this test file omit encoding='utf-8'. The project enforces PLW1514 (unspecified-encoding) via ruff and has a dedicated test (test_catalog_utf8_encoding.py) that sweeps operator_api/ for the same pattern, citing four shipped bugs caused by locale-dependent decoding on Windows.",
"source": "protopatch",
"verdict": "confirmed",
"note": "Verified: lines 395 and 412 of the new test call .read_text() without encoding; pyproject.toml ruff config explicitly selects PLW1514 with preview=true and a detailed justification citing four Windows locale bugs."
}
]There was a problem hiding this comment.
Promoting the PASS verdict for head a7371ec7ed0b: all checks terminal-green, zero unresolved review threads. (approve-on-green)
Open findings carried by this approval — non-blocking, but they did not go away:
- nit
tests/test_bundled_config_assets.py:423— read_text() without encoding='utf-8' contradicts the codebase's own PLW1514 guardrail enforced via ruff.
Approving a WARN does not resolve its findings (issue #22).
Adds the Analyst archetype as a held catalog row (Josh tests before the picker). It answers questions from the operator's own data files with read-only SQL (the data plugin's DuckDB) and shows each answer as one live chart and a one-line takeaway.
Bundle: https://github.com/protoLabsAI/analyst-archetype (v0.1.0 in review, protoLabsAI/analyst-archetype#1)
Changes (same shape as brand-launch, #3981)
config/soul-presets/analyst.md: the persona. Its default path is question → schema → query → one chart plus a one-line takeaway. It states its assumptions (date range, metric definition), cites the source file, never invents a number (derived percentages and ratios are computed in SQL too), and keeps answers short. First run: if no data folders are allowlisted, it tells the operator to use Settings ▸ Plugins ▸ Data Analyst ▸ Data folders and never tries to set that itself (data_dirsis operator-only).config/archetype-catalog.json: aheldrow with_held: "Josh tests before the picker (2026-10-03).",id: analyst, iconChartColumn(lucide 0.468 has it;BarChart3is not inicons), the bundle URL,soul_preset: analyst, andrequires_tools: [data_query, data_chart].config/plugin-directory.yaml:archetype_reposregistersprotoLabsAI/analyst-archetype(the catalog→registry guard covers held rows).tests/test_bundled_config_assets.py:test_analyst_archetype_is_heldpins the row as held, never served, preset resolving, bundle URL and contract, plus the persona's hard rules.docs/guides/fleet.md: names Analyst among the held rows.The setup dialog asks for a data folder through the bundle's
config_inputs(data.data_dirs,type: path). No core change was needed.Dependencies
Gates
ruff check .✓ ·lint-imports✓ (4 kept)pytest tests/ -q -n 8: 11843 passed, 17 skippedscripts/live_smoke.py✓Live check
I ran a throwaway instance on :7908 from a worktree with #4025 and this branch merged in, with the version set to 0.192.0 locally, on
anthropic-oauth:claude-sonnet-5-5:data_sourcescall and returned the Settings path. It made noset_configcall.data_dirsset to a synthetic coffee-shop CSV, it answered with a bar chart in the Artifact panel and "Saturday sells the most: $2,351 … 38.6% above the overall daily average", plus the assumptions and the source file. The numbers match an independent recompute.An earlier run had the model compute "28% above" by hand when the real figure was 38%. That's why the persona now requires derived figures to come from SQL.
🤖 Generated with Claude Code