Skip to content

Add shared AAIF authority comparison evidence runner - #288

Open
wlad232 wants to merge 3 commits into
agentrust-io:mainfrom
wlad232:codex/aaif-authority-comparison
Open

wlad232 wants to merge 3 commits into
agentrust-io:mainfrom
wlad232:codex/aaif-authority-comparison

Conversation

@wlad232

@wlad232 wlad232 commented Oct 7, 2026

Copy link
Copy Markdown

Closes #287.

Adds pinned public-evidence adapters for MintID and Proofable, an executable synthetic executor fixture, and released TRACE reference shape validation. Reports keep publisher observations, absent fields, unperformed appraisals and actual execution provenance distinct; external evidence does not become attestation or an independently verified effect.

Validation: Python 3.12.15, 27 tests including valid control, forged dispatch, action tampering, duplicate suppression, deadline closure, manifest coverage and missing-case rejection. MintID operator-record hashes pass. The pinned Proofable trace intentionally fails the strict publisher digest check; the diagnostic identifies LF/CRLF equivalence without altering expected hashes. Native client runs reproduced issuer/control and remaining revocation observations, but retain an initial HTTP 503 abort and a failed positive baseline in the bounded retry. Native reproduction scope and artifact hashes are recorded in RESULTS.md.

No Verified tier or conformance level is requested. Original vendor sources/evidence remain in their repositories. GitHub maintainer is the authenticated human submitting account wlad232; the source/evidence and author-run limitations remain explicit.

@carloshvp carloshvp left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 084571b in an isolated Python 3.12 checkout. Two items need correction before approval:

  1. [P2] Enforce the complete comparison contract in integrations/alakris-authority-comparison/runner.py:59-73. validate_report validates only fields already present and supplies defaults for seven fields. It never requires authority, decision, dispatch, committed_effect, task_outcome, receipt, or policy_digest, despite comparison-contract.json declaring them required. Reproducer: generate the valid MintID testnet report, delete authority and decision from every case, then call validate_report(report). It returns successfully with no explicit missing-evidence states. This lets an incomplete adapter report pass while dropping the evidence gaps the contract is meant to preserve. Require every contract field (or explicitly populate missing fields with the appropriate state and reason), and add regression coverage for omitted core fields while retaining a valid control.

  2. [P2] Fix the new code's existing CI lint gate. The workflow rule set (ruff check integrations/alakris-authority-comparison --target-version py39 --select E4,E7,E9,F) reproduces 46 errors, including E701/E702, E741, and runner.py:130's unused trace variable. Please make the added files pass the repository's existing rule set; no broader style-rule expansion is requested.

Validation otherwise succeeded: hash-locked released dependencies installed; check_package.py --fetch passed; all 27 tests ran without skips; 12 TRACE Reference shapes validated; MintID publisher bytes passed; the known Proofable LF/CRLF checksum failure remained an explicit integrity FAIL. All 15 retained native-run collector hashes matched. I inspected the retained native artifacts but did not rerun vendor sandbox execution. The separation of author observations, synthetic effects, pointer shape, and native-run qualifications is appropriate; this review makes no Verified-tier or conformance claim.

@imran-siddique imran-siddique left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@wlad232 the repo-wide ruff gate fails on this package: 46 errors across create_references.py, runner.py and test_runner.py (E701/E702 one-line statements, E731 assigned lambdas, F841 an unused trace at runner.py:130). The rest of CI passes. To reproduce what CI runs:

ruff check integrations/alakris-authority-comparison --target-version py39 --select E4,E7,E9,F

Please push the fix; the package review follows once CI is green.

@chrisjleal

chrisjleal commented Oct 9, 2026 •

Copy link
Copy Markdown

Proofable update: frozen packet repinned, independent reader PR open

@wlad232 — Proofable's corrected Oct 8 package is frozen at proofable/docs commit
9851059900e29ba701bb0f4df2bf89a2107f430c, path
public/evidence/aaif/authority-at-dispatch/2026-10-08/, run aaif-e1217327a-20261008
(deployed revision e1217327a1568b39d5153686b8085471b5329cde).

This resolves the two blockers your review raised for Proofable:

  • every published digest is computed over the published LF bytes, so the strict byte check on
    trace.jsonl now passes (the earlier digest resolved via LF-to-CRLF conversion);
  • the disclosed CAIP-380 portable envelopes are included and verify from the published bytes
    alone (11 envelopes; qHash recomputed, EIP-191 signer recovered, independent of Proofable's
    runtime and SDK).

On the binding-veto case, the machine-readable results and the signed authority receipt distinguish
the two levels: authority-effect-results.json carries job_admission: ALLOW and
protected_action_decision: DENY, and the signed receipt for the case is action "job.dispatch"
with outcome "ALLOW". Only the protected run_command action was denied, before executor
assignment. The reader uses those structured fields. The frozen Markdown summary retains the older
aggregate label ("job dispatch DENY; run_command DENY") and has not been rewritten.

Independent code-based appraisal of this packet, from the ProbityAI reader, is open as a PR:
probityai/agent-evidence-observer#88 — repinned to 9851059, 0 failures
(14 checks), 38 tests passed, including a check that performs the envelope verification instead of
skipping it when no envelope is disclosed.

If you repin the shared runner's Proofable adapter to 9851059, the strict integrity check should
move from the historical FAIL to PASS, and the envelope check can be performed rather than marked
not-performed. The regenerated alakris.patch lives in the reviewer handoff on docs main:
https://github.com/proofable/docs/tree/3d688f4cd2e3fef940cb9c8e04dcf4f0d80ea57c/public/evidence/aaif/authority-at-dispatch/reviewer-handoff-2026-10-08
It applies cleanly to 8598d1028bf6c0986c676829fe906f72d63363e0; your PR #288 is at 084571bd, so
bring your newer source fixes into the PR first, then apply or port it.

The old 3a45f026 package keeps its historical failure. The corrected package is 9851059.

Scope is unchanged: author-operated (SELF); independent code-based appraisal of published evidence,
not an independent third-party execution, not AAIF conformance, not an AARM rank and not target-side
effect custody. Signatures prove an envelope is internally consistent and bound to its signer, not
that any off-chain fact behind it is true.

Copy link
Copy Markdown
Contributor

Proofable follow-up for the shared runner at 084571bd337ab605219a74ac183604a74f0c3eb5.

The corrected frozen packet is proofable/docs commit 9851059900e29ba701bb0f4df2bf89a2107f430c, path public/evidence/aaif/authority-at-dispatch/2026-10-08/. I rehashed its exact published bytes: all seven SHA256SUMS entries pass. Running this PR's unchanged proofable() adapter on that packet reports integrity PASS, with six consistent trace records.

The related reader now verifies all 11 envelopes and binds their qHashes to the trace and manifest. My current-head rerun results are recorded at probityai/agent-evidence-observer#88 (comment) (reader head 2b54ee47491f0d1a86c7d9f47d434f8f02622043; 14 packet checks pass, 46 implementation-record tests pass, full profile 103 pass / 1 source-dependent skip).

Suggested bounded integration:

  1. Repin fetch_fixtures.py to the Oct 8 commit/path and retrieve portable-proofs.json and verify-portable-proofs.mjs alongside the existing artifacts.
  2. Retain the Oct 6 checksum FAIL as historical evidence; give the corrected packet its own result and update the active-fixture test expectation.
  3. Replace the adapter's stale statement that public signature material is absent. Material presence and signature verification remain separate: importing the packet alone must not promote receipt_signature_verification to PASS.
  4. If integrating offline verification, pin the reviewed reader revision and retain its DID/chain/signature, exact qHash-set and ordered-event negative controls.

The initial pin/material-only patch has been exercised locally for checksum PASS, measured material presence, unperformed-signature state, and trace/envelope byte-tamper rejection; I have not run this shared runner's entire suite or changed this PR branch. Receipt verification remains unperformed in this shared adapter until the verifier is actually integrated and run.

No change to SELF custody, independent-live-run status, target-side effect limitations, conformance claims or unmeasured revocation latency is warranted.

dinakarjs commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Published the bounded Proofable follow-up for integration: dinakarjs#1 — now ready for review, with 12 changed files against this PR's unchanged head 084571bd337ab605219a74ac183604a74f0c3eb5. Contribution head: 70e386d609db7ba520a09c113990ba1bd080c296 on dinakarjs/integrations:fix/proofable-oct8-offline-appraisal.

It repins the corrected Oct 8 packet, retains the historical Oct 6 checksum FAIL separately, vendors the reviewed offline verifier with license/attribution, and checks qHash/signature/DID/chain plus exact trace/manifest/envelope qHash-set binding. The lint follow-up fixes the shared-runner compound statements, assigned lambdas, ambiguous test variable and unused local while retaining JSONL parsing; the vendored verifier remains unchanged.

Validation at the current head: all eight triggered GitHub workflows succeeded, including Lint and Alakris authority comparison evidence checks. python check_package.py --fetch fetched 20 immutable public artifacts, validated 20 reference shapes, and reported package checks PASS, MintID author bytes PASS, corrected Oct 8 Proofable bytes/envelopes PASS, with the known historical Proofable byte failure preserved. Evidence-check run: https://github.com/dinakarjs/integrations/actions/runs/37960938758. This supersedes the earlier incomplete external-fetch limitation. Local unittest rerun: 28 passed / one absent MintID local-outage fixture skip.

Producer SELF custody and evidence boundaries remain unchanged: no independent live comparison, current-freshness, target-side effect-custody, conformance or measured revocation-latency claim.

The connector could not open a direct PR against wlad232/integrations (HTTP 403), so this contributor-fork PR provides the exact review diff for the shared-runner author to integrate. No upstream branch was changed or PR merged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

AAIF #13: shared authority-at-dispatch runner, adapters and reproducible evidence fixtures

5 participants