Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
117 changes: 117 additions & 0 deletions .github/workflows/verify-bundle.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# Bundle verify-and-bump loop (ADR 0049). Copy into your bundle repo's
# .github/workflows/ — under examples/ in protoAgent this file is inert.
#
# Two jobs:
# verify — installs THIS manifest's pin set into a scratch agent on a fresh
# protoAgent checkout and probes every declared console view. Runs on
# every PR (so a pin-bump must prove itself), on demand, and weekly
# (so a rotted pin set turns the schedule red instead of rotting silently).
# bump — (every trigger except pull_request: the schedule, manual dispatch,
# and the `member-released` repository_dispatch member repos send on
# release, #2960) checks each tag-pinned member for a newer
# release tag and opens/updates a PR with the bump; that PR then runs
# `verify`. But a PR opened with the repository GITHUB_TOKEN never
# auto-starts a `pull_request` run — GitHub holds it `action_required`
# until a maintainer approves it (see README.md "Pin-bump PR lifecycle",
# #2645). So this job (a) reuses ONE branch/PR per bundle repo instead of
# piling up duplicates every week, and (b) fails loudly — PR comment +
# label + a red job — if the run it just pushed doesn't show up started
# within a bounded wait, so an unapproved candidate reads as a
# maintenance failure instead of quietly rotting.

name: verify-bundle

on:
pull_request:
workflow_dispatch:
# Sent by a member repo's release.yml the moment it publishes a release (the
# member-side step lives in protoAgent's examples/bundles/member-release-notify.yml,
# #2960), so pins bump within minutes of a member release instead of at the
# next cron. The dispatch is a hint; the schedule below stays the backstop.
repository_dispatch:
types: [member-released]
schedule:
- cron: "17 6 * * 1" # weekly, Monday 06:17 UTC

permissions:
contents: write
pull-requests: write
issues: write # labeling the pin-bump PR (`gh pr edit --add-label`) rides the issues API
actions: read # `gh run list`, to detect the approval-required stall below

jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

# The core to verify against — keep in sync with the manifest's
# `verified_against:` (a release tag), or use `main` to verify ahead.
# The repo variable PROTOAGENT_REF overrides `main` while this bundle depends on
# core that hasn't merged yet (e.g. `refs/pull/<n>/head`) — delete the variable
# once it lands, so verify tracks main again.
- uses: actions/checkout@v4
with:
repository: protoLabsAI/protoAgent
ref: ${{ vars.PROTOAGENT_REF || 'main' }}
path: protoagent

- uses: astral-sh/setup-uv@v5
- name: Sync protoAgent deps
working-directory: protoagent
run: uv sync

- name: Verify pin set (install + load + probe declared views)
working-directory: protoagent
run: uv run --no-sync python "$GITHUB_WORKSPACE/scripts/verify_bundle.py" "$GITHUB_WORKSPACE"

bump:
if: github.event_name != 'pull_request'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- name: Check tag-pinned members for newer releases
id: check
run: |
python3 scripts/check_bundle_updates.py protoagent.bundle.yaml | tee "${BUMP_PINS_SCRATCH_DIR:-/tmp}/bumps.txt"
echo "changed=$(git diff --quiet && echo false || echo true)" >> "$GITHUB_OUTPUT"

# ONE candidate per bundle repo (#2645): a stable branch name means a later
# scheduled run force-pushes and updates the existing open PR instead of
# opening a duplicate. One manifest per bundle repo, so "per repo" here is
# just "per repo" — see README.md "Pin-bump PR lifecycle" for the full
# contract. Treat `bump-pins` as bot-owned: it is rewritten wholesale every
# run, so hand edits to it don't survive the next bump. Logic lives in
# scripts/bump_pins_pr.sh (#2669) — a real, unit-tested script rather than
# inline YAML bash, so a change to this contract has automated coverage
# instead of relying on a manual dry-run against a live repo.
- name: Open or update the pin-bump PR
if: steps.check.outputs.changed == 'true'
id: pr
env:
GH_TOKEN: ${{ github.token }}
run: bash "$GITHUB_WORKSPACE/scripts/bump_pins_pr.sh" open-or-update

# Make the approval-required stall visible instead of silent (#2645): poll
# for the `verify` run this push should have queued. If it comes back
# `action_required` (GitHub's recursion guard for GITHUB_TOKEN-authored
# PRs) — or never shows up at all within the bounded wait, which is worse —
# comment on the PR, label it, and fail this job so an unapproved candidate
# turns the weekly schedule red instead of rotting quietly.
#
# `bump-pins` is a permanently-reused branch (one candidate per bundle repo,
# see the step above), so on a repeat run the branch name alone can
# match a STALE prior run (already-approved from last week, or a
# leftover from a since-closed PR) before GitHub has even registered a
# run for the commit just pushed — run creation is async. Filter on the
# exact commit SHA (`gh run list -c/--commit`) so the poll can only ever
# see the run for THIS push, not whatever last happened to be on top of
# the branch's run list.
- name: Flag the approval-required verify run
if: steps.check.outputs.changed == 'true'
env:
GH_TOKEN: ${{ github.token }}
run: >
bash "$GITHUB_WORKSPACE/scripts/bump_pins_pr.sh" flag-approval
"${{ steps.pr.outputs.number }}" "${{ steps.pr.outputs.branch }}" "${{ steps.pr.outputs.sha }}"
146 changes: 144 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,146 @@
# analyst-archetype

The **Analyst** archetype bundle for [protoAgent](https://github.com/protoLabsAI/protoAgent).
v0.1.0 is in review.
The **Analyst** archetype bundle for [protoAgent](https://github.com/protoLabsAI/protoAgent):
an analyst that answers questions from **your own data files**. Point it at a folder of CSV,
TSV, Parquet, JSON, Excel or SQLite files and ask a question. It reads the schema, writes
read-only SQL, and answers with **one live chart and a one-line takeaway**.

**It is read-only, and it reads only the folders you allow.** It never writes to your data
files, makes no network calls, and can't widen its own folder list.

## What an answer looks like

> **Saturday sells the most: $2,351 in average daily revenue, 38.6% above the overall daily average.**
> *(a live bar chart of average revenue by weekday, in the Artifact panel)*
> Assumed: Jul 1 – Sep 30 2026 (92 days). "Sells the most" means average revenue per day;
> each weekday appears 13 times (Wednesday 14), so averages are the fair comparison.
> Source: `analyst-data/coffee_daily_sales.csv`

(A real answer from a synthetic coffee-shop dataset, on `claude-sonnet-5-5`.)

The default path for every question: **question → schema → query → one chart + a one-line
takeaway.**

## The rules it keeps

These are written into the persona (protoAgent's `config/soul-presets/analyst.md`):

1. **It never invents a number.** Every figure comes from a query it ran, derived ones
(percentages, ratios) included. If the data can't answer the question, it says so and
says what's missing.
2. **It cites the source file** (and the table or sheet) for every answer.
3. **It says what it assumed**: the date range, how it defined the metric, which rows it
dropped and why.
4. **It keeps answers short**: one chart, the takeaway, the assumptions line, the source.
5. **It only reads the folders you allowed**, and it never tries to change that setting.

## What's inside

| plugin | source | role |
|---|---|---|
| `data` | [data-plugin](https://github.com/protoLabsAI/data-plugin) (Data Analyst) | Seven `data_*` tools on an embedded, read-only DuckDB (connect, sources, schema, profile, query, chart, export) and the skills `exploring-a-dataset` and `building-a-chart` |
| `artifact` | builtin | Renders `data_chart`'s live Vega-Lite charts, themed to the console |
| `notes` | builtin (on by default) | Metric definitions and findings between sessions |

The persona is protoAgent's `analyst` soul preset. It ships with core, and this bundle names
it, so there is one copy to keep correct.

**Not included:** dashboards, scheduled reports, databases over the network, coding
delegates, a browser. It answers questions; it doesn't build things.

## Create an agent

**Core floor: protoAgent ≥ 0.192.0** on the hub. The data plugin declares
`min_protoagent_version: 0.192.0` (the `vega-lite` artifact kind its charts render through),
and the loader refuses it on older cores. Bundles can't enforce a floor themselves, so it's
stated here.

### From the picker (once the archetype is listed)

The `analyst` row is in protoAgent's archetype catalog but **held**, so the picker doesn't
serve it yet. Once it's listed: **Fleet ▸ New agent ▸ Analyst**, pick your **Data folder**
on the set-up step, then **Create**.

Until then, there's a console route on any ≥ 0.192.0 hub: **Settings ▸ Plugins ▸ Install
from URL** → `https://github.com/protoLabsAI/analyst-archetype`. An installed bundle with an
`archetype:` block registers itself as a picker card. If the hub's core predates the
`analyst` preset, paste the persona into the set-up step's **Advanced ▸ Persona** field.

### From the API

This is the exact body the picker sends. The persona goes in inline, so it doesn't matter
whether the hub's core ships the preset yet:

```bash
curl -fsSL https://raw.githubusercontent.com/protoLabsAI/protoAgent/main/config/soul-presets/analyst.md -o analyst.md

jq -n --rawfile soul analyst.md '{
name: "analyst",
bundle: "https://github.com/protoLabsAI/analyst-archetype",
soul: $soul,
requires_tools: ["data_query", "data_chart"],
config_inputs: { "data.data_dirs": "/absolute/path/to/your/data" }
}' | curl -s -X POST http://127.0.0.1:7870/api/fleet \
-H 'content-type: application/json' \
${PROTOAGENT_TOKEN:+-H "authorization: Bearer $PROTOAGENT_TOKEN"} \
-d @-
```

`7870` is the default instance; use your hub's port. The member inherits the hub's model
connection (`inherit_config: true` is the default).

### Direct install onto an existing agent

```
python -m server plugin install https://github.com/protoLabsAI/analyst-archetype
```

Then enable the suggested list (`data, artifact, notes`) and set the data folder (below).

## First run

The set-up step asks for one thing: the **Data folder** that holds your files. It's
optional. If you skip it, the agent's first answer tells you where to set it:

> **Settings ▸ Plugins ▸ Data Analyst ▸ Data folders**

The value is one or more absolute folders (comma- or newline-separated). It's
**operator-only**: the agent's `set_config` refuses it, and the persona never asks to change
it. Credential files, your home directory as a whole and the agent's own home are refused
even when listed.

Excel workbooks need `openpyxl` on the host; the agent says how to install it when it meets
one. Then ask a question: *"Which weekday sells the most?"*

## Pin lifecycle (ADR 0049)

External members are pinned to release tags, and each pin is a **floor**: installs take the
newest *compatible* release (caret semantics, so for 0.x the minor is the boundary).
`scripts/verify_bundle.py`, run by `.github/workflows/verify-bundle.yml` on every PR and
weekly, installs this manifest into a scratch agent on a fresh protoAgent checkout, loads
every member, probes each declared console view, and checks the capability contract
(`requires_tools`). `scripts/check_bundle_updates.py` opens a bump PR only for an
out-of-range release.

### Pin-bump PR lifecycle

The `bump` job (weekly, on dispatch, and on a member's `member-released` dispatch) reuses
**one** `bump-pins` branch and PR, rewritten wholesale each run, so don't hand-edit it. A PR
opened with the repository `GITHUB_TOKEN` never auto-starts its `pull_request` run: GitHub
holds it as `action_required` until a maintainer approves it. The job detects that, labels
and comments on the PR, and fails, so an unapproved candidate turns the schedule red instead
of going stale unnoticed.

While this bundle depends on core that hasn't been released, the repo variable
`PROTOAGENT_REF` (e.g. `refs/pull/<n>/head`) points `verify` at that ref. Delete it once the
change lands.

Run the verify locally from a protoAgent checkout:

```
uv run --no-sync python /path/to/analyst-archetype/scripts/verify_bundle.py /path/to/analyst-archetype
```

## License

MIT
80 changes: 80 additions & 0 deletions protoagent.bundle.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# A protoAgent plugin BUNDLE (ADR 0040) — the Analyst archetype: answers the operator's
# questions from their own local data files (CSV, TSV, Parquet, JSON, Excel, SQLite) with
# read-only SQL on an embedded DuckDB, and shows each answer as one live chart plus a
# one-line takeaway. This is the dedicated bundle repo protoAgent's archetype-catalog.json
# `analyst` row installs (HELD from the picker until the operator lists it); the persona
# ships as protoAgent's `analyst` soul preset.
#
# Install (onto any host): python -m server plugin install <this-repo-url>
# — or create a fleet member from it (README "Create an agent").
#
# READ-ONLY, AND ONLY WHAT YOU ALLOW. The data plugin opens files only inside the folders
# in `data.data_dirs`, never writes to them, and makes no network calls. `data_dirs` is
# operator-only (`spawns: true`): the agent's set_config refuses it, and the persona is told
# never to try.
#
# CORE FLOOR — protoAgent >= 0.192.0 (bundles carry no min-version, so it is stated here
# and in the README, not enforced by the bundle):
# * data-plugin declares `min_protoagent_version: 0.192.0` (the `vega-lite` artifact kind
# and the `artifact.show` plugin service data_chart renders through, protoAgent #4025),
# and the loader REFUSES it on an older core, so the agent would have no data tools;
# * the `analyst` soul preset ships with the protoAgent PR that added this archetype's
# catalog row. On a core without it, a picker/`soul_preset` create falls back to the
# base persona. Pass the persona inline as `soul` instead (README).
#
# FIRST-RUN SETUP — one optional answer (core >= 0.144.0 renders `config_inputs:`): the
# folder that holds the operator's data. It becomes `data.data_dirs`. Left blank, the agent
# tells the operator where to set it (Settings ▸ Plugins ▸ Data Analyst ▸ Data folders) on
# the first question.
#
# MODEL — deliberately unset: the agent uses the host's model connection.
#
# PIN LIFECYCLE (ADR 0049): external members are pinned to release TAGS; on core >= 0.146
# a pin is a FLOOR (installs take the newest COMPATIBLE tag, caret semantics, so for 0.x
# the minor is the boundary). Bump only to adopt an out-of-range release. The verify
# workflow proves each pin set. Keep member entries on ONE line:
# scripts/check_bundle_updates.py rewrites `ref:` in place.

id: analyst-archetype
name: Analyst
description: >-
An analyst that answers questions from your own data files. Point it at a folder of CSV,
TSV, Parquet, JSON, Excel or SQLite files and ask a question: it reads the schema, writes
read-only SQL against an embedded DuckDB, and answers with one live chart and a one-line
takeaway. It says what it assumed (date range, metric definition), cites the file the
numbers came from, never invents a number, and keeps answers short. It reads only the
folders you allow and never writes to them.

# The core version this pin set was last verified against (ADR 0049): metadata, not
# enforced. See CORE FLOOR above.
verified_against: 0.192.0

plugins:
- { id: data, url: https://github.com/protoLabsAI/data-plugin, ref: v0.1.0 } # Data Analyst: data_connect/sources/schema/profile/query/chart/export + the exploring-a-dataset and building-a-chart skills
- { id: artifact, builtin: true } # renders data_chart's live Vega-Lite charts
- { id: notes, builtin: true } # on by default: definitions and findings between sessions

enabled: [data, artifact, notes] # suggested turn-on list

# Create-time setup prompts (#2934, core >= 0.144.0). `help` renders on core >= 0.183.0.
config_inputs:
- key: data.data_dirs
label: "Data folder"
help: "The folder that holds your data files. The agent can read files inside it and never writes to them. Add more later in Settings ▸ Plugins ▸ Data Analyst ▸ Data folders."
type: path

# No `config:` defaults on purpose. The data plugin's own defaults (row caps, a 20 s query
# timeout, a 1 GB memory limit) are right for a fresh agent, and `data_dirs` is the
# operator's answer above, never a bundle default.

# Agent archetype (ADR 0042 / 0100): presents this bundle as the "Analyst" agent type in
# the new-agent picker once installed. `soul_preset` gives a DIRECT `plugin install <url>`
# the persona on a host whose core ships the preset.
archetype:
label: Analyst
icon: ChartColumn
soul_preset: analyst
blurb: Answers questions from your own data files — CSV, Parquet, Excel, SQLite — with SQL and a live chart. Read-only, and only the folders you allow.
# Capability contract (#2277, ADR 0100): the persona commits to querying and charting.
# Each tool is registered by the data member; verify_bundle.py fails if one isn't.
requires_tools: [data_query, data_chart]
Loading
Loading