Skip to content

Two measurement gaps beyond L1 and L2 #21

Description

@euyis1019

What L1 and L2 do not reach

L1 checks an engine against a protocol surface. L2 checks it against web platform semantics. Both work one page and one mechanism at a time. Neither answers the question an agent framework asks first: can this engine hold up across a long, stateful interaction?

The published run points at that gap. Three capabilities cover multi-step flows, a login followed by a filter, a filtered aggregate, a paginated cart. They are the only capabilities that Chrome passes and no candidate engine passes. They are also the smallest group in the map. The measurement that separates the engines best is the one there is least of.

A second gap sits past that one. An engine can pass every conformance check and still be the wrong browser to run rollouts on. It may be too slow once a real agent loop drives it, or wrong on pages no fixture author anticipated. Finding that out takes a real agent working on real pages, and no deterministic grader can score that work. It is worth measuring, and it cannot share a denominator with work that is deterministic.

What a layer means here

A layer in this repository is not a difficulty tier. Each layer carries its own evaluation axis and counts its own kind of unit. L1 counts task attempts. L2 counts capability attempts. So a layer is one score unit together with one object of measurement.

That definition splits the two gaps into two layers rather than one.

Whole interactive environments, served locally and graded deterministically, still measure the engine and still count tasks. They belong beside L1 and L2.

A real agent on live pages measures a pair, the agent together with the engine, and reports a success rate under a policy that never repeats exactly. Different object, different unit, different class of evidence.

The definition also settles a question that would otherwise get decided by accident. If the live agent work sat in the same layer as the deterministic environments, the whole layer would have to be marked as not formally scored, and the deterministic subsets would lose their standing for no reason of their own. The Kitesurf lane already shows what a lane outside formal scoring looks like: its own report generator, its own branch, no row in the headline table. One of these two gaps needs that treatment. The other must never need it.

The picture that results

Layer What it measures Unit Determinism Formally scored
L1 protocol and driver compatibility the engine task full yes
L2 web platform semantics the engine capability full yes
L3 frozen interactive environments the engine task or episode high yes
L4 end-to-end agent on live pages agent and engine together success rate low no, runs as a lane

Determinism falls and realism rises across the four. That is also why L4 sits outside the headline table: it measures a different object.

Why the boundary is worth drawing before the tasks exist

Once a task carries a layer id, moving it changes what the dataset contains. That is a major version bump, and it breaks comparison with every published run. Drawing the boundary after the tasks exist is the expensive order.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

architectureStructure of the benchmarkdocsDocumentationscoringScoring rules and denominators

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions