Repository navigation
fix(checker): accept native reports and agentic run configurations - #93
Merged
arav-agarwal2 merged 2 commits intoOct 3, 2026
Merged
Conversation
Recognize kimi-k3, qwen3.6-35b-a3b, and deepseek-v4.1-flash. Read native accuracy_scores lists and decimal percentile keys without rewriting measured files, and accept the power template's component field aliases. Accept the agentic_inference fixed-concurrency load pattern. Treat stream_all_chunks as client IPC forwarding configuration, not server streaming state. Allow zero warmup concurrency only when duration and request counts are all zero, and skip warmup-log disclosure for those disabled warmups. Preserve legacy formats and reject duplicate dataset names, conflicting sample counts, and contradictory percentile aliases. Document the accepted formats. Validation: 1,031 tests passed; Ruff and mypy passed. Seven K3 points drop from 57 to 21 errors with measured inputs unchanged; remaining disclosures, metric totals, drafter approvals, and power evidence are still required. Zero-warmup declarations give the same results without the INT_MAX workaround. Refs: mlcommons#90 Refs: mlcommons#91 Refs: mlcommons#92
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
arav-agarwal2
added a commit
that referenced
this pull request
Sep 24, 2026
§3.2 makes the reference implementation the authority for per-model specifications — "canonical weights, dataset, chat template, server parameters, accuracy target" — so the models and thresholds both come from its Agentic Inference example. Models, as #93 adds them: kimi-k3, qwen3.6-35b-a3b, deepseek-v4.1-flash. The accuracy gates do not fit §15's. That gate folds every dataset of a point into one sample-weighted score per metric and gates each point against it; the agentic benchmarks gate three quantities aggregated three different ways, so folding them together gives a number with no meaning under either rule. New `agentic_targets.py` holds the table and `_check_agentic_accuracy` gates them, with §15's gate standing down for a recognised agentic model: - `agentic-accuracy-inline` — per point. "Every Kimi K3 and Qwen3.6-35B-A3B submitted Pareto point must satisfy all of the model-specific accuracy thresholds." - `agentic-accuracy-swebench` — mean-of-N across points, which is §4.3's multi-turn branch: "The arithmetic mean of the N required accuracy results MUST meet the quality threshold; individual results need not." A short set is still gated, on the mean of what is present, and said to be short. - `agentic-osl-range` — per point, against a range, and read from `result_summary.json` rather than the accuracy results. The field is resolved explicitly and never falls back to the windowed `output_sequence_lengths`, which has the same shape and a different value; a silent fallback would gate the wrong number. DSV4 is recognised but every one of its thresholds is TBD upstream, so it reports as ungateable rather than passing silently — a different report from an unrecognised model, which is also distinguished. §3.2 says this list does not belong in a release: it is published "at least 6 weeks before the submission round opens", so a new round should not need a new checker. `data/seed_sets.yaml` and `data/approved_drafters.yaml` are the pattern to follow when that is worth doing; noted at the list rather than done here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
arekay-nv
marked this pull request as ready for review
September 25, 2026 13:59
Main adopted agentic_inference and the canonical model names, so its versions win; the branch keeps its streaming-config semantics, *_watts power aliases, and parametrized allowed-model test. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
arekay-nv
pushed a commit
to arekay-nv/endpoints-submission-cli
that referenced
this pull request
Oct 1, 2026
Four §9.1 rows apply differently to agentic benchmarks — §5.3's point minimum and accuracy count, §5.7's Offline point, and §9.1's "Offline point present" — but §8.3 carries no field naming the benchmark type. `offline-point-present` said so in its own docstring and warned instead of erroring. §6.1's load pattern answers it. The reference implementation names its fixed-concurrency agentic scheduler `agentic_inference`, and that name is the only agentic signal in any file a submission carries. - `PointConfig.is_agentic` reads the pattern; `ModelContext.is_agentic` requires the curve's points to agree, since §8.5 makes one result one benchmark. A curve that disagrees is reported by a new `benchmark-type-consistency` rule and read as single-turn — the stricter branch, so one mislabelled point cannot switch off §5.7. - `load-pattern` accepts both §6.1 patterns; `poisson` and friends stay out. - `offline-point-present` inverts for agentic: §5.7 says an agentic submission "neither requires nor may include" an Offline point, so presence is the error. For single-turn, absence is now the ERROR §9.1 specifies rather than a hedge. - `point-count` holds agentic at 7 even where a point declares `dedicated`; that declaration is itself the defect. - `accuracy-coverage` already yields N=4 for agentic, because the Offline row iterates `offline_points`. Its pass message now names which N applied. No fixture declared `offline` at all, so the WARN had been covering a corpus that is non-compliant under v1.0. `regenerate_fixtures.py` now elects the C_max point as the Offline result (§5.7.2 Option 2), which adds no run and leaves every point count unchanged. The builder contract test needed its curve moved rather than patched: it topped out at 1000 against `max_supported_concurrency: 1024`, so no point existed to elect. The builder is untouched — it must not invent §8.3 disclosure (mlcommons#72), so the declaration lives in the test's input. Overlaps mlcommons#93, which also admits `agentic_inference` as a load pattern but does not wire the distinction to any rule. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
arav-agarwal2
approved these changes
Oct 3, 2026
arav-agarwal2
added a commit
that referenced
this pull request
Oct 6, 2026
Two conflicts, and two consequences of main's changes that merged cleanly but no longer held together. - system_power.py: #93 added `_watts` aliases to the flat ComponentGroup schema, which Appendix E replaces with sourced values. This branch's side is kept. - regenerate_fixtures.py: main's steps 8 (cooling) and 9 (accuracy shape) are kept, and maximal engagement becomes step 10. - Step 8 now declares the TPU and GB300 fixtures liquid-cooled, while their Appendix E descriptors said air, which E.2 rejects. The generator now keeps a descriptor's `cooling` in step with the description; sub_c, sub_d and sub_j are regenerated. - #93's two flat-template power tests in test_native_formats.py are removed. Appendix E has no `num_cpu`/`tdp_per_cpu_watts` or `provisioned_power_watts`, and test_pre_appendix_e_form_is_rejected covers the flat form. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Recognize kimi-k3, qwen3.6-35b-a3b, and deepseek-v4.1-flash. Read native accuracy_scores lists and decimal percentile keys without rewriting measured files, and accept the power template's component field aliases.
Accept the agentic_inference fixed-concurrency load pattern. Treat stream_all_chunks as client IPC forwarding configuration, not server streaming state. Allow zero warmup concurrency only when duration and request counts are all zero, and skip warmup-log disclosure for those disabled warmups.
Preserve legacy formats and reject duplicate dataset names, conflicting sample counts, and contradictory percentile aliases. Document the accepted formats.
Validation: 1,031 tests passed; Ruff and mypy passed. Seven K3 points drop from 57 to 21 errors with measured inputs unchanged; remaining disclosures, metric totals, drafter approvals, and power evidence are still required. Zero-warmup declarations give the same results without the INT_MAX workaround.
Refs: #90
Refs: #91
Refs: #92