Skip to content

test(conformance): finish ACTION historical validity-window replay cases #36

Description

@carloshvp

Context

The A2A delegation-link field has landed in agentrust-io/trace-spec#80, and the public delegation verifier API has landed in agentrust-io/agent-manifest#218. cA2A also now has a conformance suite from #28.

A useful follow-up would be fixture-style conformance coverage that connects these pieces to offline-verifiable action evidence:

delegation block -> public delegation verifier -> TRACE/action receipt evidence

This would line up with the action receipt discussion in agentrust-io/trace-spec#66 and the embodied action receipt example in agentrust-io/examples#36.

Proposed fixture cases

MUST-level cases:

  • valid root -> delegated child TRACE record, with delegation.parent_record_hash matching the canonical parent record hash
  • valid delegation.credential_id that resolves through the public agent-manifest delegation verifier
  • parent record present but canonical hash mismatch
  • missing parent record for a non-root delegated hop
  • delegation credential id unknown to the verifier
  • delegation credential signature invalid
  • delegation credential expired or not yet valid
  • delegatee/session/channel binding mismatch
  • requested action outside the effective delegated scope
  • valid delegation chain with local policy denial, reported as valid provenance plus authorization/policy denial rather than malformed evidence

SHOULD-level cases:

  • multi-hop attenuation where each hop narrows scope
  • attempted scope widening at an intermediate hop
  • valid negative outcome, such as delegated action rejected by the controller, treated as useful evidence rather than verifier failure
  • external subject identifier/digest present for cross-system resolution, but not dereferenced by base TRACE verification

Boundary

The verifier should distinguish three classes of result:

  1. provenance invalid: malformed delegation block, bad hash, unknown credential, invalid signature, broken binding
  2. authorization invalid: valid delegation evidence, but requested action is outside delegated scope
  3. valid negative outcome: delegation/action was well evidenced, but local policy/controller denied or rejected the action

That boundary keeps cA2A compatible with embodied/action evidence: the system can prove what was authorized and attempted without claiming that the physical or business-world outcome succeeded.

I can help with the fixture shape or a follow-up PR once the preferred fixture location/API surface is clear.

Activity

  1. carloshvp commented on Jul 29, 2026

    @carloshvp
    MemberAuthor

    Reconciliation after #37 merged:

    Covered by ACTION-001 through ACTION-007:

    • a valid parent-linked delegated action through ca2a_verify.verify_delegation_chain()
    • parent hash mismatch, missing parent, and unknown credential ID as provenance-invalid
    • out-of-scope actions and local-policy denials as authorization-invalid
    • controller rejection as valid negative evidence

    Residual checklist:

    • Exercise the public agent-manifest delegation verifier, or explicitly revise that acceptance criterion to the cA2A ca2a_verify API used by ACTION-001.
    • Add an ACTION-path invalid-signature case. DELEG-001 covers the base credential check, but the ACTION group does not assert the resulting provenance classification.
    • Define credential validity-window fields before adding expired and not-yet-valid cases. DelegationCredential currently has no time bounds.
    • Add an explicit delegatee-mismatch ACTION case and define session or channel binding before adding those mismatch cases. The current action evidence binds record hash, credential ID, and requested capability; base provenance separately cross-checks the credential subject.
    • Add explicit ACTION-path coverage for a longer attenuating chain and an intermediate scope-widening attempt. The generic DELEG group covers these chain rules, but not their action-evidence classification.
    • Decide how to represent an external subject identifier or digest, then add the SHOULD-level case that verifies it without dereferencing it.

    The core provenance, authorization, and negative-outcome split is implemented. Keeping this issue open for the checklist above.

  2. bytebackllc commented on Aug 13, 2026

    @bytebackllc
    Contributor

    Checklist state as of today, then proposals for the two items that are design-blocked rather than code-blocked.

    Where the residual checklist stands

    Proposal: channel binding for action evidence

    The gap: since #107, the holder proof commits to the channel pair (audience, caller_channel_key), but proofs deliberately leave no trace in the emitted record — a HOLD refusal emits no provenance, and offline replay cannot reconstruct the proof. So an auditor of action evidence today cannot tell which channel an action was bound to, and a "channel mismatch" case has nothing to check.

    Proposal: record the channel identity the callee actually verified, as two optional fields on DelegationRecord:

    audience_channel_key: str | None      # the callee channel key the proof named
    caller_channel_key:   str | None      # the caller's offered channel key, when it made one
    
    • Omit-when-absent in the hashed body, like the denial fields — records emitted before the fields existed keep their hashes (the caller_attestation precedent of always-hashing is the exception for a reason that doesn't apply here: "not offered" vs "not checked" ambiguity doesn't arise, because a record with no channel fields is simply a pre-field record).
    • P-4a already requires every request field that reaches the record to be committed in the proof, and both of these are in the proof body since fix(delegation): commit the parent link, and drop the proof replay cache #107 — so the proof and the record close over the same channel pair, and in-flight substitution of either key is already HOLDER_PROOF_INVALID on the live path.
    • New MUST cases: action evidence naming a channel key that differs from the record's → provenance_invalid / PROVENANCE_LINK_BROKEN; evidence against a pre-field record (no channel fields) → verifiable as today, channel binding reported as not-asserted rather than failed.

    Alternative considered: recording only audience_channel_key (session = which callee), leaving the caller key out. Rejected because HOLD-006 is precisely about the caller half — an attested caller using another party's chain — and an audit-time analogue of it needs the caller key on the record.

    Proposal: external subject identifier

    For the SHOULD-level case ("present for cross-system resolution, but not dereferenced"): one optional field on the action evidence,

    external_subject: {"scheme": "<namespace, e.g. urn|https|did>", "id": "<identifier>", "digest": "sha256:<64 lowercase hex>"} | None
    
    • Strict wire rules, matching the credential's ethos: exact field set, non-empty strings, digest anchored to sha256: + lowercase hex. Malformed shape → provenance_invalid.
    • Committed into the evidence (and record hash when it reaches the record), so it cannot be swapped under verified evidence.
    • Base TRACE verification never dereferences id — no scheme handlers, no network. The verifier surfaces it as carried, not resolved; asserting anything about the external system is explicitly out of scope of the base verdict, which keeps the provenance / authorization / outcome boundary intact.
    • SHOULD case: well-formed external subject verifies offline and is reported; a digest-format violation is classified as malformed evidence, not as an external-system failure.

    Happy to implement either or both exactly as decided — same shape as #110: model + conformance IDs + README rows + spec docs in one PR.

  3. Evalbound commented on Aug 21, 2026

    @Evalbound

    I reviewed the current ACTION coverage and residual checklist against cA2A main 52141e84e6bf8bd26f08d8e27e73ae707a9f95a3. The three-way boundary here is useful and I would keep it narrow: these fixtures establish whether delegation/action evidence was valid at the relevant decision time; they should not also be asked to determine whether a later authoritative correction superseded the intent or authority that the evidence depended on.

    I opened #129 to track that separate lifecycle question. The proposed model keeps historical ACTION evidence unchanged and adds a separate correction/applicability axis (superseded / reappraisal_required / post-commit) rather than overloading provenance-invalid, authorization-invalid, or valid-negative-outcome.

    So I am not proposing another #36 acceptance criterion. The contribution here is the boundary clarification: temporal credential validity and ACTION evidence validity are necessary, but not sufficient to answer present applicability after a principal correction. #129 is where I will keep any schema/verifier discussion so this issue can close on its existing checklist.

  4. Bashirloyan commented on Aug 24, 2026

    @Bashirloyan
    Contributor

    For reliable assurance, fixtures should separate provenance, authorization, and outcome.
    Each fixture should record two independent results:
    • evidence_status: whether the delegation chain, credential, signature, bindings, and parent hash are valid.
    • decision_status: whether the requested action falls within delegated scope and whether local policy permits it.
    This prevents valid policy denials from appearing as evidence failures and supports audit, control testing, and incident review.
    For multi-hop delegation, fixtures should also record effective scope at each hop to show attenuation or attempted scope widening.
    Should expected results live inside each fixture or in a separate conformance manifest? Once that and the API surface are confirmed, the next step is to define the fixture shape.

  5. K611-dot commented on Aug 25, 2026

    @K611-dot

    Hi @carloshvp — interested in picking this up as part of the AgenTrust Fellowship. Looking at the current layout (tests/conformance/test_profile_conformance.py + tests/fixtures/), would you want the MUST-level cases as a new module under tests/fixtures/delegation_action_evidence/ with parametrized entries wired into test_profile_conformance.py, or somewhere else? Also tracking trace-spec#80 and agent-manifest#218 since the delegation-link field and verifier API are still landing — happy to start on the cases that don't depend on either (hash mismatch, missing parent record, invalid signature) while those stabilise. Will open a draft PR once the fixture shape is confirmed.

  6. ams-belal commented on Aug 25, 2026

    @ams-belal

    I reviewed #36 against current main and the merged ACTION work (#37, #76, #80, #110, #134).

    One narrower gap I noticed is historical validity-window replay. verify_delegation_chain(..., at_time=...) already supports evaluating a chain at a supplied time, and the delegation spec describes audit replay at the action decision time. But _ActionEvidence does not currently carry that time, so ACTION-012/013 effectively evaluate credential validity against the verifier’s current clock.

    Would a small test-only contribution around that boundary be useful?

    Concretely, I was thinking of:

    • adding an evidenced decision time to _ActionEvidence;
    • passing it through verify_delegation_chain(at_time=...);
    • making ACTION-012/013 deterministic relative to that decision time; and
    • adding a positive historical-replay case where the credential was valid when the action was decided but has expired by the later audit time.

    I would keep this limited to the conformance tests/README, with no production or public API changes.

  7. athena-kanellatou commented on Aug 25, 2026

    @athena-kanellatou

    I’ve been working through the current conformance implementation against the cases proposed here. It looks like most of the original boundary is now represented explicitly by ACTION-001–ACTION-013, including the distinction between provenance-invalid evidence, authorization-invalid evidence, and a valid negative controller outcome.

    Two residual questions stood out to me before attempting a PR.

    First, ACTION-004 currently checks that the evidence credential_id matches the verified leaf credential after verify_delegation_chain(chain, ...). Is the intended #36 requirement satisfied by this, or should conformance exercise actual identifier-based resolution through the public Agent Manifest verification surface? The latter seems slightly stronger: it tests not only correspondence with an already supplied chain, but whether the delegation reference can be independently resolved and verified as an external auditor would encounter it.

    Second, the SHOULD case for an external subject identifier/digest appears not to have an ACTION-* fixture yet. I think there is value in making the non-dereference boundary executable: a verifier should preserve the external identifier/digest as resolution material, while base TRACE/cA2A verification should neither require network resolution nor infer identity claims from it.

    If that interpretation matches the intended boundary, I’d be happy to contribute a small follow-up PR covering the external-subject fixture and, depending on the preferred API boundary, an explicit credential-resolution conformance case.

  8. Ahmedibrahim222 commented on Aug 25, 2026

    @Ahmedibrahim222

    A small self-contained fixture bundle for each case could make the offline-verification requirement explicit. For example: parent_trace.json, child_trace.json, delegation_credential.json, verifier key material, and expected_result.json.

    For the expected result, would it make sense to assert separate fields for provenance, authorization, and action_outcome, along with a stable reason code? That would prevent a valid delegation followed by local policy denial from being reported as malformed evidence.

    I’d be interested in prototyping the valid root-to-child case, canonical parent-hash mismatch, and valid-provenance/policy-denial case. Is there a preferred fixture directory, and should the conformance runner call the public agent-manifest verifier directly or through a cA2A adapter?

  9. imran-siddique commented on Aug 25, 2026

    @imran-siddique
    Member

    Maintainer here, and I owe this thread a decision rather than another opinion. Six people have written on it in five days and none of you has been told where the line is.

    The original scope is closed

    @athena-kanellatou is right. Checked at pinned c6b040bd: tests/conformance/README.md documents ACTION-001 through ACTION-013, all MUST-level, and they cover every case @carloshvp proposed in July.

    proposed in #36 now
    valid root to delegated child, parent hash matching ACTION-001
    parent_record_hash mismatch ACTION-002
    parent record absent ACTION-003
    credential_id resolving through the public verifier ACTION-004

    Plus nine the issue did not ask for: scope-not-permitted against policy-denied as distinct classes (005, 006), negative outcome as valid evidence (007), invalid credential signature ordered before authorization (008), multi-hop attenuation (009), mid-chain scope widening (010), delegatee/subject mismatch (011), expired and not-yet-valid credentials (012, 013).

    So this issue is done as written. I am leaving it open only until the piece below has a home.

    The one gap I can confirm is real

    @ams-belal named it: historical validity-window replay. src/ca2a_verify/verify.py:33 already takes it:

    def verify_delegation_chain(
        chain: list[DelegationCredential],
        *,
        trusted_root_issuers: Collection[str],
        max_depth: int = 8,
        at_time: int | None = None,
    ) -> ChainResult:

    and its own docstring says an auditor replaying recorded evidence passes it. ACTION-012 and ACTION-013 test expired and not-yet-valid at verification time. Nothing tests that a credential which was valid when the action happened still verifies when replayed at that historical time, or that one which was not yet valid then is rejected even though it is valid now.

    That distinction is the whole point of an audit trail. Evidence that only verifies while the credential happens to still be live is not evidence, and a parameter with no conformance case behind it is a parameter that will drift.

    @K611-dot, this is yours if you want it

    You asked about layout. Cases go in tests/conformance/test_profile_conformance.py alongside the existing ACTION set, documented in tests/conformance/README.md in the same table format, with fixtures under tests/fixtures/. Follow ACTION-012 and ACTION-013 as the closest models, since they are the validity-window cases and yours are their replay counterparts.

    Minimum I would want, as ACTION-014 and ACTION-015:

    • a credential valid at the action's timestamp and expired now, verifying when at_time is the action's timestamp;
    • a credential not yet valid at the action's timestamp and valid now, rejected when at_time is the action's timestamp.

    Both should fail if at_time is ignored, which is the property worth asserting. Please prove that rather than asserting it: run them with the at_time argument dropped and show both go red.

    The other two proposals

    @Bashirloyan, separating evidence_status from outcome into two independent results is a design change to what a conformance run reports, not a fixture addition. It is a reasonable idea and it needs its own issue, because it changes the shape of every existing ACTION case rather than adding to them. Please open one if you want to pursue it.

    @Ahmedibrahim222, self-contained per-case bundles (parent_trace.json, child_trace.json, delegation_credential.json, key material) would make the offline-verification requirement concrete, and that is worth having. It is also a refactor of the existing fixture layout rather than new coverage, so it should not block ACTION-014/015. Open a separate issue and I will scope it.

    @moneyparking, your three-way boundary framing is what the ACTION classification codes already encode, and keeping it narrow was the right instinct.

    @carloshvp, thank you for filing this and for the reconciliation pass after #37. The checklist you wrote in July is the reason the coverage got to thirteen cases.

  10. ams-belal commented on Aug 26, 2026

    @ams-belal

    @imran-siddique Thanks for the clarification and for confirming the historical replay gap. The distinction between validity at action time and validity at later audit time was the boundary I was trying to isolate.

    ACTION-014/015 make that invariant concrete. I’ll use this clarification as the baseline and look for another narrowly scoped contribution that doesn’t overlap the work already underway.

  11. imran-siddique commented on Aug 30, 2026

    @imran-siddique
    Member

    @ams-belal Before you go looking, here is a concrete one, because you should not have to hunt for it after finding the gap that produced ACTION-014/015.

    agentrust-io/trace-spec#247, opened this morning, needs a schema-to-model parity test and nobody is on it.

    The problem it tracks is that TRACE states its rules across four surfaces (the normative spec, the JSON Schema, the reference model, and the docs) and nothing says which wins. Five instances turned up in three days. The one that makes the test obvious is #244: the schema's subject pattern is ^(spiffe://|did:), a bare prefix test, while models.py:343 requires ^(spiffe://[^/]+/.+|did:[a-z0-9]+:.+)$. So spiffe://bernstein.run is schema-conformant and refused by model_validate, and a producer validating against the vendored schema learns nothing about whether the reference implementation will accept the same record.

    That was found by a producer hitting it in production rather than by CI. A parity test would have caught it before it left the repository.

    Why I think it suits you specifically:

    • It does not overlap anything underway. ACTION-014/015 are with @K611-dot, and this is a different repository.
    • It is narrow. One test module comparing the two surfaces field by field, failing on any constraint present in one and absent or weaker in the other.
    • It is the same discipline you were already applying here. You isolated validity at action time from validity at later audit time by asking what a test would have to assert to tell them apart. This is that question asked of two schemas instead of two timestamps.
    • The interesting part is the design, not the code. "Weaker" needs defining. A regex that is a strict prefix of another is comparable; two unrelated regexes are not. Deciding what the test can honestly check, and saying plainly what it cannot, is most of the value, and getting that wrong in the permissive direction gives everyone a green check that means nothing.

    trace-spec#236 added tests/test_the_schema_and_the_models_agree.py yesterday in roughly that direction. Worth reading first: it may already cover part of this, in which case the useful contribution is extending it and saying which cases it does not reach.

    Two other pieces from this thread are also unclaimed, if you would rather stay in ca2a. @Bashirloyan's separation of evidence_status from outcome, and @Ahmedibrahim222's self-contained per-case bundles, were both invited to become their own issues and neither has been opened. I would rather the originators filed them, so ask them first if either appeals.

    Thanks for taking the clarification cleanly rather than arguing the boundary. Naming the historical replay gap is what produced two conformance cases that did not exist before.

  12. athena-kanellatou commented on Aug 30, 2026

    @athena-kanellatou

    Thank you @imransiddique, I appreciate the clarification and the pointer. Glad the replay distinction turned out to be useful for the conformance cases.

  13. imran-siddique commented on Sep 7, 2026

    @imran-siddique
    Member

    The original ACTION-001 through ACTION-013 scope is complete per the maintainer reconciliation. The remaining assigned work is ACTION-014/015 historical validity-window replay, including controls that fail when at_time is ignored. Reporting redesign is tracked separately in #144.

  14. changed the title [-]test(conformance): add delegation-link action evidence fixtures[/-] [+]test(conformance): finish ACTION historical validity-window replay cases[/+] on Sep 7, 2026
  15. ozereray commented on Sep 13, 2026

    @ozereray

    The three-way result boundary here is particularly valuable: provenance-invalid, authorization-invalid, and a valid-but-denied action should remain distinguishable. I’d make the final action receipt explicitly bind the delegated identity/credential, effective scope, requested action, local policy decision, and outcome, so a denial is still useful evidence rather than looking like a malformed delegation. That is close to the execution-evidence model we use in Aegisora: prove what was delegated, what policy decided, and what actually reached the execution boundary.

  16. imran-siddique commented on Sep 24, 2026

    @imran-siddique
    Member

    @K611-dot, ACTION-014/015 remain reserved, and I found no open PR for them. Are you still taking this work?

    The scope remains two historical-replay cases: valid then but expired now, and not yet valid then but valid now. Please include controls showing both tests fail when at_time is ignored. A draft PR is enough to make progress reviewable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

help wantedExtra attention is needed

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions