Repository navigation
Determine the benchmark type from §6.1's load pattern - #94
Merged
Merged
Conversation
Four §9.1 rows apply differently to agentic benchmarks — §5.3's point minimum and accuracy count, §5.7's Offline point, and §9.1's "Offline point present" — but §8.3 carries no field naming the benchmark type. `offline-point-present` said so in its own docstring and warned instead of erroring. §6.1's load pattern answers it. The reference implementation names its fixed-concurrency agentic scheduler `agentic_inference`, and that name is the only agentic signal in any file a submission carries. - `PointConfig.is_agentic` reads the pattern; `ModelContext.is_agentic` requires the curve's points to agree, since §8.5 makes one result one benchmark. A curve that disagrees is reported by a new `benchmark-type-consistency` rule and read as single-turn — the stricter branch, so one mislabelled point cannot switch off §5.7. - `load-pattern` accepts both §6.1 patterns; `poisson` and friends stay out. - `offline-point-present` inverts for agentic: §5.7 says an agentic submission "neither requires nor may include" an Offline point, so presence is the error. For single-turn, absence is now the ERROR §9.1 specifies rather than a hedge. - `point-count` holds agentic at 7 even where a point declares `dedicated`; that declaration is itself the defect. - `accuracy-coverage` already yields N=4 for agentic, because the Offline row iterates `offline_points`. Its pass message now names which N applied. No fixture declared `offline` at all, so the WARN had been covering a corpus that is non-compliant under v1.0. `regenerate_fixtures.py` now elects the C_max point as the Offline result (§5.7.2 Option 2), which adds no run and leaves every point count unchanged. The builder contract test needed its curve moved rather than patched: it topped out at 1000 against `max_supported_concurrency: 1024`, so no point existed to elect. The builder is untouched — it must not invent §8.3 disclosure (#72), so the declaration lives in the test's input. Overlaps #93, which also admits `agentic_inference` as a load pattern but does not wire the distinction to any rule. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
arav-agarwal2
marked this pull request as ready for review
September 24, 2026 16:28
§3.2 makes the reference implementation the authority for per-model specifications — "canonical weights, dataset, chat template, server parameters, accuracy target" — so the models and thresholds both come from its Agentic Inference example. Models, as #93 adds them: kimi-k3, qwen3.6-35b-a3b, deepseek-v4.1-flash. The accuracy gates do not fit §15's. That gate folds every dataset of a point into one sample-weighted score per metric and gates each point against it; the agentic benchmarks gate three quantities aggregated three different ways, so folding them together gives a number with no meaning under either rule. New `agentic_targets.py` holds the table and `_check_agentic_accuracy` gates them, with §15's gate standing down for a recognised agentic model: - `agentic-accuracy-inline` — per point. "Every Kimi K3 and Qwen3.6-35B-A3B submitted Pareto point must satisfy all of the model-specific accuracy thresholds." - `agentic-accuracy-swebench` — mean-of-N across points, which is §4.3's multi-turn branch: "The arithmetic mean of the N required accuracy results MUST meet the quality threshold; individual results need not." A short set is still gated, on the mean of what is present, and said to be short. - `agentic-osl-range` — per point, against a range, and read from `result_summary.json` rather than the accuracy results. The field is resolved explicitly and never falls back to the windowed `output_sequence_lengths`, which has the same shape and a different value; a silent fallback would gate the wrong number. DSV4 is recognised but every one of its thresholds is TBD upstream, so it reports as ungateable rather than passing silently — a different report from an unrecognised model, which is also distinguished. §3.2 says this list does not belong in a release: it is published "at least 6 weeks before the submission round opens", so a new round should not need a new checker. `data/seed_sets.yaml` and `data/approved_drafters.yaml` are the pattern to follow when that is worth doing; noted at the list rather than done here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
anandhu-eng
approved these changes
Sep 30, 2026
| if not self.valid_points: | ||
| return self | ||
| patterns = sorted({c.runtime_settings.load_pattern for _, c in self.valid_points}) | ||
| if len(patterns) > 1: |
Collaborator
There was a problem hiding this comment.
Pointing out that in rules, it seems offline load pattern is accepted: https://github.com/mlcommons/endpoints_policies/blob/v1.0_rules_dev/endpoints_rules.md#61-load-pattern
| "deepseek-r1", | ||
| "kimi-k3", | ||
| "qwen3.6-35b-a3b", | ||
| "deepseek-v4.1-flash", |
Collaborator
There was a problem hiding this comment.
https://github.com/mlcommons/endpoints/tree/main/examples/10_Agentic_Inference#dsv4-1
We might have to allign on what the name is
Collaborator
Author
There was a problem hiding this comment.
I've changed it from flash, but honestly yea. It looks like this may be something to make an issue on and check during rules today.
- Benchmark type: a dedicated Offline run no longer votes. §6.1 gives it its own load pattern (`max_throughput` in the reference implementation), so a correct single-turn curve was failing `benchmark-type-consistency`. And an agentic curve that wrongly carried one was read as single-turn, which hid the §5.7 violation from `offline-point-present`. An elected C_max point is a fixed-concurrency run and still votes. - SWE-bench mean-of-N takes one value per mandatory band. §5.3 puts the four required results at the mandatory points, and §4.3 asks for one run per region. Results outside the four bands are left out. The rules name no tiebreak for a band with several results, so those are averaged into one value, with a warning, rather than letting a submitter pick which one counts. - Agentic scores are always read as fractions. Both reference scorers return [0, 1], and guessing the unit from the value would read a 0.9% score reported as 0.9 as 90%. A value outside [0, 1] is an error. - `deepseek-v4.1-flash` is now `deepseek-v4-pro`, the reference README's DeepSeek-V4-Pro (DSV4). Qwen is written canonically, as `qwen3_6-35b-a3b`, to match #96. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 30, 2026
arav-agarwal2
added a commit
that referenced
this pull request
Sep 30, 2026
Restore the canonical Llama name lost in #94's merge
arekay-nv
pushed a commit
to arekay-nv/endpoints-submission-cli
that referenced
this pull request
Oct 1, 2026
mlcommons#94's merge resolved the `_ALLOWED_MODEL_NAMES` conflict by taking mlcommons#94's side of the block, which predated mlcommons#96. That brought back `llama3.1-8b` and the comment saying the name comes from system_desc.json, so main rejects the canonical `llama3_1-8b` that its own fixtures, and every submission the builder writes, declare. `valid_standardized` fails `model-name-valid`, along with ten tests. This restores mlcommons#96's name and comment, and keeps mlcommons#94's agentic entries. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Four §9.1 rows apply differently to agentic benchmarks, but §8.3 carries no field naming the benchmark type.
offline-point-presentadmitted as much in its own docstring and warned instead of erroring:§6.1's load pattern answers it. The reference implementation names its fixed-concurrency agentic scheduler
agentic_inference, and that name is the only agentic signal in any file a submission carries:The determination
PointConfig.is_agenticis the per-point signal.ModelContext.is_agenticis the curve-level answer and requires unanimity — §8.5 defines a result as one system, one benchmark model, one dataset, so a curve is one benchmark. A curve whose points disagree is reported by a newbenchmark-type-consistencyERROR and read as single-turn: the stricter branch, so one mislabelled point cannot switch off §5.7.What swings on it
load-patternconcurrencyagentic_inference— both accepted;poissonand friends stay outpoint-countdedicateddeclaration cannot lift it to 8offline-point-presentaccuracy-coverageoffline-point-presentinverts because §5.7 is explicit:and §9.1 asks for "exactly one … for non-agentic benchmarks; none is present for agentic benchmarks".
accuracy-coverageneeded no logic change — N falls out already, since the Offline row iteratesoffline_points, which an agentic curve has none of. Its pass message now names which N applied so a reviewer can see the reading.offline-orderingneeded nothing either: it already no-ops without a dedicated point.Fixture fallout
Flipping the WARN to an ERROR broke 9 tests, because no fixture declared
offlineat all — the hedge had been covering a corpus that is non-compliant under v1.0. Fixed at the source rather than per-fixture:regenerate_fixtures.pynow elects the C_max point as the Offline result (§5.7.2 Option 2), which adds no run and leaves every point count unchanged. 12 fixtures gain one line each;--checkis clean, so it stays idempotent.The builder contract test needed its curve moved rather than patched: it topped out at 1000 against
max_supported_concurrency: 1024, so no point existed to elect. The builder itself is untouched — it must not invent §8.3 disclosure (#72), so the declaration lives in the test's synthetic input.Verification
1,016 tests pass (was 976) · mypy --strict clean · ruff clean · sphinx
-Wclean. 40 new tests covering the determination, the unanimity rule, bothoffline-point-presentbranches, both point minimums, and each agentic gate's aggregation.Second commit: the agentic models and their accuracy gates
§3.2 makes the reference implementation the authority for per-model specifications, so the models and the thresholds both come from its Agentic Inference example. Models, as #93 adds them:
kimi-k3,qwen3.6-35b-a3b,deepseek-v4.1-flash.The accuracy gates could not reuse §15's. That gate folds every dataset of a point into one sample-weighted score per metric; the agentic benchmarks gate three quantities aggregated three different ways, and folding them together gives a number with no meaning under either rule. So §15's gate stands down for a recognised agentic model and
agentic_targets.py+_check_agentic_accuracytake over:agentic-accuracy-inlineagentic_combinedagentic-accuracy-swebenchswe_benchagentic-osl-rangeresult_summary.jsonThresholds: Kimi K3 inline ≥ 58.32, OSL 425–520, SWE-bench ≥ 93.5. Qwen3.6-35B-A3B ≥ 55.86, 344–422, ≥ 69.
Two details worth a reviewer's eye:
output_sequence_lengths, which has the same shape and a different value. The README is emphatic about this, and a silent fallback would gate the wrong number — so an absent full-run block warns rather than substituting.A short SWE-bench set is still gated, on the mean of what is present, and flagged as short — three strong points passing unremarked seemed the wrong failure mode, and the missing results are
accuracy-coverage's report.Open items
agentic_inferenceas a load pattern but does not wire the distinction to any rule, so the two will conflict on_check_load_pattern. Either let fix(checker): accept native reports and agentic run configurations #93 keep the vocabulary change and rebase this on top as the wiring, or fold this into fix(checker): accept native reports and agentic run configurations #93.ConcurrencySchedulerpoints and says "single-pass agentic workloads are handled only by the ad-hoc diagnostic tool", butagentic_inferenceis a fixed-concurrency scheduler — whether an agentic point is "single-pass" is not something the load pattern answers, so exempting it would be a guess.data/seed_sets.yamlanddata/approved_drafters.yamlare the pattern; noted at the list rather than done here, since fix(checker): accept native reports and agentic run configurations #93 adds the names the same way.🤖 Generated with Claude Code