ci: Day-1 eval automation — frontier-release watcher + runbook issues + report generator - #284
Closed
reacher-z wants to merge 1 commit into
Closed
ci: Day-1 eval automation — frontier-release watcher + runbook issues + report generator#284reacher-z wants to merge 1 commit into
reacher-z wants to merge 1 commit into
Conversation
Daily cron polls the public OpenRouter catalog for frontier-vendor models released in the last 2 days, dedups against previous day1-eval issues, and opens a runbook issue: models.yaml entry, v1-lite smoke, full V2 batch, rescore, leaderboard row, and a results-post draft via day1_report.py. Includes a vendor-outreach template per model. The benchmark run itself happens on maintainer infra, never in CI.
Collaborator
|
I don't think we want to and are should catch up evaluation for all newly released models. It's too costly both economically and timely. Also the GitHub issues will be flooded by such new model evaluation ones. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Publicity workstream #1/#2 (per maintainer's 2026-08-16 directive): make Day-1 evaluations of new frontier models a standing, automated habit.
What this adds
.github/workflows/day1-eval.yml— daily cron; when a frontier vendor (Anthropic/OpenAI/Google/xAI/DeepSeek/Moonshot/Zhipu/MiniMax/Qwen/Meta/Mistral) lists a new model on OpenRouter, opens aday1-evalissue with a complete runbook (smoke → full V2 → rescore → leaderboard row → results-post draft). Dedups against prior issues;:freevariants skipped; benchmark runs stay on maintainer infra.scripts/day1_watch.py— the watcher (stdlib only; tested live: correctly lists recent releases).scripts/day1_report.py— turnsbatch-summary.json(+ optional rescore summary) into a ready-to-edit results post: headline numbers, 6-tweet X thread, Chinese blurb.Notes for review
workflow_dispatchcovers manual runs.