Skip to content

fix(checker): accept native reports and agentic run configurations - #93

Merged
arav-agarwal2 merged 2 commits into
mlcommons:mainfrom
arekay-nv:fix/native-agentic-submissions
Oct 3, 2026
Merged

arav-agarwal2 merged 2 commits into
mlcommons:mainfrom
arekay-nv:fix/native-agentic-submissions

Conversation

@arekay-nv

Copy link
Copy Markdown
Contributor

Recognize kimi-k3, qwen3.6-35b-a3b, and deepseek-v4.1-flash. Read native accuracy_scores lists and decimal percentile keys without rewriting measured files, and accept the power template's component field aliases.

Accept the agentic_inference fixed-concurrency load pattern. Treat stream_all_chunks as client IPC forwarding configuration, not server streaming state. Allow zero warmup concurrency only when duration and request counts are all zero, and skip warmup-log disclosure for those disabled warmups.

Preserve legacy formats and reject duplicate dataset names, conflicting sample counts, and contradictory percentile aliases. Document the accepted formats.

Validation: 1,031 tests passed; Ruff and mypy passed. Seven K3 points drop from 57 to 21 errors with measured inputs unchanged; remaining disclosures, metric totals, drafter approvals, and power evidence are still required. Zero-warmup declarations give the same results without the INT_MAX workaround.

Refs: #90
Refs: #91
Refs: #92

Recognize kimi-k3, qwen3.6-35b-a3b, and deepseek-v4.1-flash. Read native
accuracy_scores lists and decimal percentile keys without rewriting measured
files, and accept the power template's component field aliases.

Accept the agentic_inference fixed-concurrency load pattern. Treat
stream_all_chunks as client IPC forwarding configuration, not server streaming
state. Allow zero warmup concurrency only when duration and request counts are
all zero, and skip warmup-log disclosure for those disabled warmups.

Preserve legacy formats and reject duplicate dataset names, conflicting sample
counts, and contradictory percentile aliases. Document the accepted formats.

Validation: 1,031 tests passed; Ruff and mypy passed. Seven K3 points drop from
57 to 21 errors with measured inputs unchanged; remaining disclosures, metric
totals, drafter approvals, and power evidence are still required. Zero-warmup
declarations give the same results without the INT_MAX workaround.

Refs: mlcommons#90
Refs: mlcommons#91
Refs: mlcommons#92
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

arav-agarwal2 added a commit that referenced this pull request Sep 24, 2026
§3.2 makes the reference implementation the authority for per-model
specifications — "canonical weights, dataset, chat template, server
parameters, accuracy target" — so the models and thresholds both come
from its Agentic Inference example.

Models, as #93 adds them: kimi-k3, qwen3.6-35b-a3b, deepseek-v4.1-flash.

The accuracy gates do not fit §15's. That gate folds every dataset of a
point into one sample-weighted score per metric and gates each point
against it; the agentic benchmarks gate three quantities aggregated
three different ways, so folding them together gives a number with no
meaning under either rule. New `agentic_targets.py` holds the table and
`_check_agentic_accuracy` gates them, with §15's gate standing down for
a recognised agentic model:

- `agentic-accuracy-inline` — per point. "Every Kimi K3 and
  Qwen3.6-35B-A3B submitted Pareto point must satisfy all of the
  model-specific accuracy thresholds."
- `agentic-accuracy-swebench` — mean-of-N across points, which is §4.3's
  multi-turn branch: "The arithmetic mean of the N required accuracy
  results MUST meet the quality threshold; individual results need not."
  A short set is still gated, on the mean of what is present, and said
  to be short.
- `agentic-osl-range` — per point, against a range, and read from
  `result_summary.json` rather than the accuracy results. The field is
  resolved explicitly and never falls back to the windowed
  `output_sequence_lengths`, which has the same shape and a different
  value; a silent fallback would gate the wrong number.

DSV4 is recognised but every one of its thresholds is TBD upstream, so
it reports as ungateable rather than passing silently — a different
report from an unrecognised model, which is also distinguished.

§3.2 says this list does not belong in a release: it is published "at
least 6 weeks before the submission round opens", so a new round should
not need a new checker. `data/seed_sets.yaml` and
`data/approved_drafters.yaml` are the pattern to follow when that is
worth doing; noted at the list rather than done here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@arekay-nv
arekay-nv marked this pull request as ready for review September 25, 2026 13:59
Main adopted agentic_inference and the canonical model names, so its
versions win; the branch keeps its streaming-config semantics, *_watts
power aliases, and parametrized allowed-model test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
arekay-nv pushed a commit to arekay-nv/endpoints-submission-cli that referenced this pull request Oct 1, 2026
Four §9.1 rows apply differently to agentic benchmarks — §5.3's point
minimum and accuracy count, §5.7's Offline point, and §9.1's "Offline
point present" — but §8.3 carries no field naming the benchmark type.
`offline-point-present` said so in its own docstring and warned instead
of erroring.

§6.1's load pattern answers it. The reference implementation names its
fixed-concurrency agentic scheduler `agentic_inference`, and that name
is the only agentic signal in any file a submission carries.

- `PointConfig.is_agentic` reads the pattern; `ModelContext.is_agentic`
  requires the curve's points to agree, since §8.5 makes one result one
  benchmark. A curve that disagrees is reported by a new
  `benchmark-type-consistency` rule and read as single-turn — the
  stricter branch, so one mislabelled point cannot switch off §5.7.
- `load-pattern` accepts both §6.1 patterns; `poisson` and friends stay
  out.
- `offline-point-present` inverts for agentic: §5.7 says an agentic
  submission "neither requires nor may include" an Offline point, so
  presence is the error. For single-turn, absence is now the ERROR §9.1
  specifies rather than a hedge.
- `point-count` holds agentic at 7 even where a point declares
  `dedicated`; that declaration is itself the defect.
- `accuracy-coverage` already yields N=4 for agentic, because the
  Offline row iterates `offline_points`. Its pass message now names
  which N applied.

No fixture declared `offline` at all, so the WARN had been covering a
corpus that is non-compliant under v1.0. `regenerate_fixtures.py` now
elects the C_max point as the Offline result (§5.7.2 Option 2), which
adds no run and leaves every point count unchanged.

The builder contract test needed its curve moved rather than patched:
it topped out at 1000 against `max_supported_concurrency: 1024`, so no
point existed to elect. The builder is untouched — it must not invent
§8.3 disclosure (mlcommons#72), so the declaration lives in the test's input.

Overlaps mlcommons#93, which also admits `agentic_inference` as a load pattern
but does not wire the distinction to any rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@arav-agarwal2
arav-agarwal2 merged commit bdbb6af into mlcommons:main Oct 3, 2026
7 checks passed
arav-agarwal2 added a commit that referenced this pull request Oct 6, 2026
Two conflicts, and two consequences of main's changes that merged
cleanly but no longer held together.

- system_power.py: #93 added `_watts` aliases to the flat
  ComponentGroup schema, which Appendix E replaces with sourced values.
  This branch's side is kept.
- regenerate_fixtures.py: main's steps 8 (cooling) and 9 (accuracy
  shape) are kept, and maximal engagement becomes step 10.
- Step 8 now declares the TPU and GB300 fixtures liquid-cooled, while
  their Appendix E descriptors said air, which E.2 rejects. The
  generator now keeps a descriptor's `cooling` in step with the
  description; sub_c, sub_d and sub_j are regenerated.
- #93's two flat-template power tests in test_native_formats.py are
  removed. Appendix E has no `num_cpu`/`tdp_per_cpu_watts` or
  `provisioned_power_watts`, and test_pre_appendix_e_form_is_rejected
  covers the flat form.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants