test(evals): add a claude plugin eval suite for routing, pushback and gates - #55
Merged
Merged
Conversation
… gates Prompt edits to the trade skill, like the v2.16.0 Opus 5.5 audit, had no way to be measured. This adds a 9-case suite under plugins/trade/evals/ for `claude plugin eval`: - Routing. 今天有没有大单 goes to daily. A one-print question is a lookup in at most two pulls, with the daily ladder offered, not run (mocked unusual-whales server). 资金流向 across names goes to report. A research link to save becomes an English writedown in the knowledge dir that knowledge_path names, never references/. A control case checks that a link sent with a trade question stays in analysis. - Pushback. A wrong claim is held with evidence; a right one produces an update that says what changed. Both resume a transcript. - Gates. The bull-conviction count and P/L matrix run before a Jade Lizard recommendation, and the preflight skips corpora/. A put sale is not approved on IV rank alone (pitfall 36). Graders check behavior (reference files read, tool and mock calls, files created, the content of the Write call), with short PASS/FAIL rubrics only for answer shape. No case touches real market data. Docs: run and A/B instructions in the skill README, authoring notes in evals/README.md, a pointer in CLAUDE.md. .gitignore covers run results and the session files a resumed case writes next to history.jsonl. No version bump: the suite is outside the release zip, and the README change is docs only.
…ke run
A 1-run Opus 5.5 smoke pass on v2.16.0 failed four replies that met the
intent of their rubric:
- import chat-reply: the clause "FAIL on a path outside writedowns/" also
caught the corpora/ source copy that import.md now requires.
- gate pl-matrix: qualitative cells ("credit kept", "large gain") are the
style of pitfall 24's own asymmetry table, and a one-line verdict ahead
of the count and matrix that justify it is fine (conviction-count
clarified the same way).
- daily answers-honestly and report states-proxy-basis: fewer conditions
per rubric. Calendar facts, and a 口径 stated for the report the agent
would run once connected, now count.
Language checks move out of the judge into a deterministic
replies-in-chinese regex (8+ consecutive CJK characters) on the four
Chinese-prompt routing cases. Each revised rubric was validated by
replaying the harness judge on the saved Opus replies, which now pass,
and on bad-reply controls, which still fail 3/3.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Prompt edits to the trade skill, like the v2.16.0 Opus 5.5 audit (#53), had no way to be measured. This adds a small behavioral suite for
claude plugin eval(CLI 2.1.269+) underplugins/trade/evals/. It has 9 cases covering routing, pushback handling and the structure gates, graded on what the run did rather than on what the reply says it did.route-daily-block-checkdaily.md. The reply answers the literal question first, in Chinese, and invents no numbersroute-print-lookuproute-report-capital-flowreport.mdand states the options-flow proxy basis (口径)route-import-article-linkknowledge_pathnames, with no Write or Edit underreferences/. The Write call's content is English, attributed, and carries the not-verified caveat and a bear caseroute-link-with-trade-questionanalysis, and nothing is writtenpushback-wrong-claim-holdpushback-right-claim-updategate-snow-jade-lizardcorpora/is never readgate-sell-put-high-ivrNo case touches real market data. Runs are sandboxed without MCP servers or API keys, and the lookup case gets its answers from canned data.
Running it
claude plugin eval plugins/trade --scaffold --allow-tools Write Edit --ablation none --no-publishThe skill README's new Evals section covers the flags, cost control, and how to compare two skill versions with worktrees. Authoring notes are in
plugins/trade/evals/README.md. Now that #53 has merged, the audit itself can be measured: runv2.15.0as the baseline andv2.16.0as the candidate, with the same model and--ablation none.Verification
claude plugin eval plugins/trade --scaffold --allow-tools Write Edit --max-cost-usd 0loads all 9 cases (schema, graders, mocks, tool grants) with no warnings. It stops before the first run, so it costs nothing. I ran it on the pre-audit skill and again on v2.16.0.All deterministic grader patterns were unit-tested against synthetic inputs, using the harness's own glob and regex semantics.
claude plugin validate --strict plugins/tradepasses.Haiku pilots, 1 run per case on the pre-audit skill (v2.15.0), about $1.10. They test mechanics, not scores. They confirmed that the mocks, scaffolds, Write grant,
file_exists, transcript resume and judges all work.Opus 5.5 smoke run, 1 run per case on v2.16.0, $3.25. It passes on everything the audit targeted:
Two real gaps remain and are left for a follow-up decision:
daily.mdsays "darkpool or option-trades".Four judge rubrics were too strict and are fixed in the second commit. Language checks moved to a deterministic regex. Each revised rubric was validated by replaying the harness judge on the saved Opus replies, which now pass, and on deliberately bad replies (fabricated flow numbers, a "verified" import with no path, a Jade Lizard answer with no count or matrix), which still fail 3/3.
Each case has one run so far, so this is a sanity check rather than a measurement.
Reviewer notes
skills/*), and the README change is docs only.evals/README.md:/trade:trade, so the pushback prompts use that form to load the currentSKILL.md.tracegrader target holds only what the child emitted: no prompt, no resumed history, no expanded command. "Skill loaded" is therefore checked through tool calls.llmjudge on the trace sees only its first and last 12 lines, so the digest content is graded with regexes over the Write call.history.jsonlmakes each run write<session-id>.jsonlnext to it..gitignorenow covers those files andevals/results/.--scaffoldruns two scripts as you. They only copy fixtures from their case directory into the run's workspace.knowledge/, because the root.gitignoreignores that name at any depth. The SNOW fixture is stored askb/and copied in.