Skip to content

Determine the benchmark type from §6.1's load pattern - #94

Merged
arav-agarwal2 merged 5 commits into
mainfrom
v1.0-rules/agentic-determination
Sep 30, 2026
Merged

arav-agarwal2 merged 5 commits into
mainfrom
v1.0-rules/agentic-determination

Conversation

@arav-agarwal2

@arav-agarwal2 arav-agarwal2 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Four §9.1 rows apply differently to agentic benchmarks, but §8.3 carries no field naming the benchmark type. offline-point-present admitted as much in its own docstring and warned instead of erroring:

Skipped entirely for agentic benchmarks… This checker has no way to tell an agentic curve from a single-turn one yet — §3 does not surface the distinction in any file it reads — so the absent case is reported as a WARNING rather than an error.

§6.1's load pattern answers it. The reference implementation names its fixed-concurrency agentic scheduler agentic_inference, and that name is the only agentic signal in any file a submission carries:

runtime_settings:
  load_pattern: agentic_inference

The determination

PointConfig.is_agentic is the per-point signal. ModelContext.is_agentic is the curve-level answer and requires unanimity — §8.5 defines a result as one system, one benchmark model, one dataset, so a curve is one benchmark. A curve whose points disagree is reported by a new benchmark-type-consistency ERROR and read as single-turn: the stricter branch, so one mislabelled point cannot switch off §5.7.

What swings on it

Rule Single-turn Agentic
load-pattern concurrency agentic_inference — both accepted; poisson and friends stay out
point-count 8 with a dedicated Offline run, else 7 always 7; a dedicated declaration cannot lift it to 8
offline-point-present ERROR when absent (was WARN) ERROR when present
accuracy-coverage N=5 N=4

offline-point-present inverts because §5.7 is explicit:

An agentic submission neither requires nor may include an Offline point.

and §9.1 asks for "exactly one … for non-agentic benchmarks; none is present for agentic benchmarks".

accuracy-coverage needed no logic change — N falls out already, since the Offline row iterates offline_points, which an agentic curve has none of. Its pass message now names which N applied so a reviewer can see the reading. offline-ordering needed nothing either: it already no-ops without a dedicated point.

concurrency          agentic=False min_points=7 offline-point-present=ERROR  No point declares `offline`…
agentic_inference    agentic=True  min_points=7 offline-point-present=INFO   Agentic benchmark: no Offline point, as §5.7 requires

Fixture fallout

Flipping the WARN to an ERROR broke 9 tests, because no fixture declared offline at all — the hedge had been covering a corpus that is non-compliant under v1.0. Fixed at the source rather than per-fixture: regenerate_fixtures.py now elects the C_max point as the Offline result (§5.7.2 Option 2), which adds no run and leaves every point count unchanged. 12 fixtures gain one line each; --check is clean, so it stays idempotent.

The builder contract test needed its curve moved rather than patched: it topped out at 1000 against max_supported_concurrency: 1024, so no point existed to elect. The builder itself is untouched — it must not invent §8.3 disclosure (#72), so the declaration lives in the test's synthetic input.

Verification

1,016 tests pass (was 976) · mypy --strict clean · ruff clean · sphinx -W clean. 40 new tests covering the determination, the unanimity rule, both offline-point-present branches, both point minimums, and each agentic gate's aggregation.

Second commit: the agentic models and their accuracy gates

§3.2 makes the reference implementation the authority for per-model specifications, so the models and the thresholds both come from its Agentic Inference example. Models, as #93 adds them: kimi-k3, qwen3.6-35b-a3b, deepseek-v4.1-flash.

The accuracy gates could not reuse §15's. That gate folds every dataset of a point into one sample-weighted score per metric; the agentic benchmarks gate three quantities aggregated three different ways, and folding them together gives a number with no meaning under either rule. So §15's gate stands down for a recognised agentic model and agentic_targets.py + _check_agentic_accuracy take over:

Rule Quantity Source Aggregation
agentic-accuracy-inline Inline accuracy agentic_combined per point — "Every … submitted Pareto point must satisfy all of the model-specific accuracy thresholds"
agentic-accuracy-swebench SWE-bench swe_bench mean-of-4 — §4.3's multi-turn branch: "the arithmetic mean … MUST meet the quality threshold; individual results need not"
agentic-osl-range OSL per-turn mean result_summary.json per point, against a range

Thresholds: Kimi K3 inline ≥ 58.32, OSL 425–520, SWE-bench ≥ 93.5. Qwen3.6-35B-A3B ≥ 55.86, 344–422, ≥ 69.

Two details worth a reviewer's eye:

  • The OSL field is resolved explicitly and never falls back to the windowed output_sequence_lengths, which has the same shape and a different value. The README is emphatic about this, and a silent fallback would gate the wrong number — so an absent full-run block warns rather than substituting.
  • DSV4 is recognised but every threshold is TBD upstream, so it reports as ungateable rather than passing silently. That is a distinct report from an unrecognised model, which also warns.

A short SWE-bench set is still gated, on the mean of what is present, and flagged as short — three strong points passing unremarked seemed the wrong failure mode, and the missing results are accuracy-coverage's report.

Open items

🤖 Generated with Claude Code

Four §9.1 rows apply differently to agentic benchmarks — §5.3's point
minimum and accuracy count, §5.7's Offline point, and §9.1's "Offline
point present" — but §8.3 carries no field naming the benchmark type.
`offline-point-present` said so in its own docstring and warned instead
of erroring.

§6.1's load pattern answers it. The reference implementation names its
fixed-concurrency agentic scheduler `agentic_inference`, and that name
is the only agentic signal in any file a submission carries.

- `PointConfig.is_agentic` reads the pattern; `ModelContext.is_agentic`
  requires the curve's points to agree, since §8.5 makes one result one
  benchmark. A curve that disagrees is reported by a new
  `benchmark-type-consistency` rule and read as single-turn — the
  stricter branch, so one mislabelled point cannot switch off §5.7.
- `load-pattern` accepts both §6.1 patterns; `poisson` and friends stay
  out.
- `offline-point-present` inverts for agentic: §5.7 says an agentic
  submission "neither requires nor may include" an Offline point, so
  presence is the error. For single-turn, absence is now the ERROR §9.1
  specifies rather than a hedge.
- `point-count` holds agentic at 7 even where a point declares
  `dedicated`; that declaration is itself the defect.
- `accuracy-coverage` already yields N=4 for agentic, because the
  Offline row iterates `offline_points`. Its pass message now names
  which N applied.

No fixture declared `offline` at all, so the WARN had been covering a
corpus that is non-compliant under v1.0. `regenerate_fixtures.py` now
elects the C_max point as the Offline result (§5.7.2 Option 2), which
adds no run and leaves every point count unchanged.

The builder contract test needed its curve moved rather than patched:
it topped out at 1000 against `max_supported_concurrency: 1024`, so no
point existed to elect. The builder is untouched — it must not invent
§8.3 disclosure (#72), so the declaration lives in the test's input.

Overlaps #93, which also admits `agentic_inference` as a load pattern
but does not wire the distinction to any rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@arav-agarwal2
arav-agarwal2 marked this pull request as ready for review September 24, 2026 16:28
arav-agarwal2 and others added 2 commits September 24, 2026 12:35
§3.2 makes the reference implementation the authority for per-model
specifications — "canonical weights, dataset, chat template, server
parameters, accuracy target" — so the models and thresholds both come
from its Agentic Inference example.

Models, as #93 adds them: kimi-k3, qwen3.6-35b-a3b, deepseek-v4.1-flash.

The accuracy gates do not fit §15's. That gate folds every dataset of a
point into one sample-weighted score per metric and gates each point
against it; the agentic benchmarks gate three quantities aggregated
three different ways, so folding them together gives a number with no
meaning under either rule. New `agentic_targets.py` holds the table and
`_check_agentic_accuracy` gates them, with §15's gate standing down for
a recognised agentic model:

- `agentic-accuracy-inline` — per point. "Every Kimi K3 and
  Qwen3.6-35B-A3B submitted Pareto point must satisfy all of the
  model-specific accuracy thresholds."
- `agentic-accuracy-swebench` — mean-of-N across points, which is §4.3's
  multi-turn branch: "The arithmetic mean of the N required accuracy
  results MUST meet the quality threshold; individual results need not."
  A short set is still gated, on the mean of what is present, and said
  to be short.
- `agentic-osl-range` — per point, against a range, and read from
  `result_summary.json` rather than the accuracy results. The field is
  resolved explicitly and never falls back to the windowed
  `output_sequence_lengths`, which has the same shape and a different
  value; a silent fallback would gate the wrong number.

DSV4 is recognised but every one of its thresholds is TBD upstream, so
it reports as ungateable rather than passing silently — a different
report from an unrecognised model, which is also distinguished.

§3.2 says this list does not belong in a release: it is published "at
least 6 weeks before the submission round opens", so a new round should
not need a new checker. `data/seed_sets.yaml` and
`data/approved_drafters.yaml` are the pattern to follow when that is
worth doing; noted at the list rather than done here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread src/endpoints_submission_cli/submissions/builder.py
if not self.valid_points:
return self
patterns = sorted({c.runtime_settings.load_pattern for _, c in self.valid_points})
if len(patterns) > 1:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pointing out that in rules, it seems offline load pattern is accepted: https://github.com/mlcommons/endpoints_policies/blob/v1.0_rules_dev/endpoints_rules.md#61-load-pattern

Comment thread src/submission_checker/models/aggregate/context.py Outdated
Comment thread src/submission_checker/models/aggregate/context.py Outdated
Comment thread src/submission_checker/checker.py Outdated
"deepseek-r1",
"kimi-k3",
"qwen3.6-35b-a3b",
"deepseek-v4.1-flash",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've changed it from flash, but honestly yea. It looks like this may be something to make an issue on and check during rules today.

arav-agarwal2 and others added 2 commits September 30, 2026 10:44
- Benchmark type: a dedicated Offline run no longer votes. §6.1 gives
  it its own load pattern (`max_throughput` in the reference
  implementation), so a correct single-turn curve was failing
  `benchmark-type-consistency`. And an agentic curve that wrongly
  carried one was read as single-turn, which hid the §5.7 violation
  from `offline-point-present`. An elected C_max point is a
  fixed-concurrency run and still votes.
- SWE-bench mean-of-N takes one value per mandatory band. §5.3 puts
  the four required results at the mandatory points, and §4.3 asks
  for one run per region. Results outside the four bands are left
  out. The rules name no tiebreak for a band with several results, so
  those are averaged into one value, with a warning, rather than
  letting a submitter pick which one counts.
- Agentic scores are always read as fractions. Both reference scorers
  return [0, 1], and guessing the unit from the value would read a
  0.9% score reported as 0.9 as 90%. A value outside [0, 1] is an
  error.
- `deepseek-v4.1-flash` is now `deepseek-v4-pro`, the reference
  README's DeepSeek-V4-Pro (DSV4). Qwen is written canonically, as
  `qwen3_6-35b-a3b`, to match #96.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arav-agarwal2
arav-agarwal2 merged commit 6e1e592 into main Sep 30, 2026
4 of 7 checks passed
arav-agarwal2 added a commit that referenced this pull request Sep 30, 2026
Restore the canonical Llama name lost in #94's merge
arekay-nv pushed a commit to arekay-nv/endpoints-submission-cli that referenced this pull request Oct 1, 2026
mlcommons#94's merge resolved the `_ALLOWED_MODEL_NAMES` conflict by taking
mlcommons#94's side of the block, which predated mlcommons#96. That brought back
`llama3.1-8b` and the comment saying the name comes from
system_desc.json, so main rejects the canonical `llama3_1-8b` that its
own fixtures, and every submission the builder writes, declare.
`valid_standardized` fails `model-name-valid`, along with ten tests.

This restores mlcommons#96's name and comment, and keeps mlcommons#94's agentic entries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants