Skip to content

test(evals): add a claude plugin eval suite for routing, pushback and gates - #55

Merged
himself65 merged 2 commits into
mainfrom
claude/vigilant-tharp-a1767f
Sep 25, 2026
Merged

himself65 merged 2 commits into
mainfrom
claude/vigilant-tharp-a1767f

Conversation

@himself65

@himself65 himself65 commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Prompt edits to the trade skill, like the v2.16.0 Opus 5.5 audit (#53), had no way to be measured. This adds a small behavioral suite for claude plugin eval (CLI 2.1.269+) under plugins/trade/evals/. It has 9 cases covering routing, pushback handling and the structure gates, graded on what the run did rather than on what the reply says it did.

Case What it checks
route-daily-block-check "NBIS 今天有没有大单" reads daily.md. The reply answers the literal question first, in Chinese, and invents no numbers
route-print-lookup "刚才那笔巨量成交是什么" is a lookup: at most 2 calls to a mocked Unusual Whales server and no GEX / max pain / IV-term / flow-alert pulls. The print is identified and the ladder is offered, not run
route-report-capital-flow "COHR LITE MU 资金流向" reads report.md and states the options-flow proxy basis (口径)
route-import-article-link A research link to save creates a file in the knowledge dir that knowledge_path names, with no Write or Edit under references/. The Write call's content is English, attributed, and carries the not-verified caveat and a bear case
route-link-with-trade-question Control: a link plus a trade question stays in analysis, and nothing is written
pushback-wrong-claim-hold A wrong pushback ("high IV rank, so buy calls") is held, with figures from the earlier turn
pushback-right-claim-update A correct pushback (IV floor plus a catalyst, so long vega) produces an update that says what changed
gate-snow-jade-lizard The bull-conviction count and the P/L matrix (with +35% and +50%) come before any structure, and there is no Jade Lizard. The user's notes are loaded and corpora/ is never read
gate-sell-put-high-ivr A put sale is not approved on IV rank alone; the pitfall 36 VRP gate runs first

No case touches real market data. Runs are sandboxed without MCP servers or API keys, and the lookup case gets its answers from canned data.

Running it

claude plugin eval plugins/trade --scaffold --allow-tools Write Edit --ablation none --no-publish

The skill README's new Evals section covers the flags, cost control, and how to compare two skill versions with worktrees. Authoring notes are in plugins/trade/evals/README.md. Now that #53 has merged, the audit itself can be measured: run v2.15.0 as the baseline and v2.16.0 as the candidate, with the same model and --ablation none.

Verification

  • claude plugin eval plugins/trade --scaffold --allow-tools Write Edit --max-cost-usd 0 loads all 9 cases (schema, graders, mocks, tool grants) with no warnings. It stops before the first run, so it costs nothing. I ran it on the pre-audit skill and again on v2.16.0.

  • All deterministic grader patterns were unit-tested against synthetic inputs, using the harness's own glob and regex semantics.

  • claude plugin validate --strict plugins/trade passes.

  • Haiku pilots, 1 run per case on the pre-audit skill (v2.15.0), about $1.10. They test mechanics, not scores. They confirmed that the mocks, scaffolds, Write grant, file_exists, transcript resume and judges all work.

  • Opus 5.5 smoke run, 1 run per case on v2.16.0, $3.25. It passes on everything the audit targeted:

    • pushback is held when wrong and updated, with the change spelled out, when right
    • both gates run before any recommendation
    • the import lands in the knowledge dir as an English digest with the caveat and a bear case
    • the one-print lookup never runs the daily ladder

    Two real gaps remain and are left for a follow-up decision:

    • The lookup made 3 pulls (candles, dark pool and option trades) where the spec allows 1–2. daily.md says "darkpool or option-trades".
    • The no-data daily plan lists the block filters but skips the activity gate.
  • Four judge rubrics were too strict and are fixed in the second commit. Language checks moved to a deterministic regex. Each revised rubric was validated by replaying the harness judge on the saved Opus replies, which now pass, and on deliberately bad replies (fabricated flow numbers, a "verified" import with no path, a Jade Lizard answer with no count or matrix), which still fail 3/3.

  • Each case has one run so far, so this is a sanity check rather than a measurement.

Reviewer notes

  • No version bump. The suite is outside the release zip (the release workflow only zips skills/*), and the README change is docs only.
  • Harness quirks found while piloting. These are handled and recorded in evals/README.md:
    • Headless runs register the skill as /trade:trade, so the pushback prompts use that form to load the current SKILL.md.
    • The trace grader target holds only what the child emitted: no prompt, no resumed history, no expanded command. "Skill loaded" is therefore checked through tool calls.
    • An llm judge on the trace sees only its first and last 12 lines, so the digest content is graded with regexes over the Write call.
    • Resuming history.jsonl makes each run write <session-id>.jsonl next to it. .gitignore now covers those files and evals/results/.
  • --scaffold runs two scripts as you. They only copy fixtures from their case directory into the run's workspace.
  • Fixture directories can't be named knowledge/, because the root .gitignore ignores that name at any depth. The SNOW fixture is stored as kb/ and copied in.

… gates

Prompt edits to the trade skill, like the v2.16.0 Opus 5.5 audit, had no
way to be measured. This adds a 9-case suite under plugins/trade/evals/
for `claude plugin eval`:

- Routing. 今天有没有大单 goes to daily. A one-print question is a lookup
  in at most two pulls, with the daily ladder offered, not run (mocked
  unusual-whales server). 资金流向 across names goes to report. A research
  link to save becomes an English writedown in the knowledge dir that
  knowledge_path names, never references/. A control case checks that a
  link sent with a trade question stays in analysis.
- Pushback. A wrong claim is held with evidence; a right one produces an
  update that says what changed. Both resume a transcript.
- Gates. The bull-conviction count and P/L matrix run before a Jade
  Lizard recommendation, and the preflight skips corpora/. A put sale is
  not approved on IV rank alone (pitfall 36).

Graders check behavior (reference files read, tool and mock calls, files
created, the content of the Write call), with short PASS/FAIL rubrics
only for answer shape. No case touches real market data.

Docs: run and A/B instructions in the skill README, authoring notes in
evals/README.md, a pointer in CLAUDE.md. .gitignore covers run results
and the session files a resumed case writes next to history.jsonl.

No version bump: the suite is outside the release zip, and the README
change is docs only.
…ke run

A 1-run Opus 5.5 smoke pass on v2.16.0 failed four replies that met the
intent of their rubric:

- import chat-reply: the clause "FAIL on a path outside writedowns/" also
  caught the corpora/ source copy that import.md now requires.
- gate pl-matrix: qualitative cells ("credit kept", "large gain") are the
  style of pitfall 24's own asymmetry table, and a one-line verdict ahead
  of the count and matrix that justify it is fine (conviction-count
  clarified the same way).
- daily answers-honestly and report states-proxy-basis: fewer conditions
  per rubric. Calendar facts, and a 口径 stated for the report the agent
  would run once connected, now count.

Language checks move out of the judge into a deterministic
replies-in-chinese regex (8+ consecutive CJK characters) on the four
Chinese-prompt routing cases. Each revised rubric was validated by
replaying the harness judge on the saved Opus replies, which now pass,
and on bad-reply controls, which still fail 3/3.
@himself65
himself65 merged commit 2148b08 into main Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant