diff --git a/CHANGELOG.md b/CHANGELOG.md index b8b200d..ada270f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,34 @@ This repository uses a **governance-first change record**. --- +## Unreleased — DEAS applicability boundary + +### Added + +- Decision Evidence Applicability Specification (DEAS) v0.2.0 as the canonical cross-regime working specification +- Seven bounded determination states for evidence-to-requirement analysis +- Explicit applicability, local-evidence, sufficiency-boundary, and non-equivalence fields +- Adjacent-system boundaries against Microsoft Agent Governance Toolkit, ScopeBlind/Acta, and Credo AI +- DEAS migration note preserving earlier DEPS links without preserving the portability claim + +### Changed + +- Renamed Decision Evidence Portability Specification (DEPS) to Decision Evidence Applicability Specification (DEAS) +- Replaced the hypothesis that evidence can “travel” across regimes with a bounded claim about applicability, assurance use, local evidence conditions, insufficiency, and non-equivalence +- Updated the GDI README, research claim register, framework mapping, related-work analysis, limitations, and working-specification status language +- Clarified that runtime governors produce enforcement evidence, signed-receipt systems establish defined integrity properties, Credo AI provides policy-pack and control infrastructure, GDI records the decision, HIT evaluates practical human influence, and DEAS qualifies evidence use under a specified requirement + +### Deprecated + +- `working-specifications/decision-evidence-portability-v0.1.md` as a current specification; the file remains as a migration pointer +- Unqualified claims that governance, controls, findings, evidence, or legal conclusions are portable across regimes + +### Research boundary + +DEAS is a working specification. Its schema, authoritative mappings, overlays, cases, reviewer protocol, and validation suite remain in development. The rename does not establish legal compliance, standards conformity, certification, admissibility, evidence sufficiency, or independent review. + +--- + ## v2.1.0 — Interoperability and examples expansion **Date:** April 2026 diff --git a/README.md b/README.md index 5324f8c..c5fc546 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ [![Specification](https://img.shields.io/badge/specification-v3.0-blue)](spec/GDI_v3_The_Decision_Architecture_for_Governed_AI.pdf) [![Schema](https://img.shields.io/badge/GDR%20schema-v2.0-4c8bf5)](schema/gdr.schema.json) [![Repository release](https://img.shields.io/badge/repository-v2.1.0-6f42c1)](CHANGELOG.md) -[![DEPS](https://img.shields.io/badge/DEPS-v0.1.1%20working-orange)](working-specifications/decision-evidence-portability-v0.1.md) +[![DEAS](https://img.shields.io/badge/DEAS-v0.2.0%20working-orange)](working-specifications/decision-evidence-applicability-v0.2.md) [![DOI](https://img.shields.io/badge/DOI-10.5281%2Fzenodo.20244601-blue)](https://doi.org/10.5281/zenodo.20244601) [![License: Apache 2.0](https://img.shields.io/badge/license-Apache%202.0-green)](LICENSE) [![ORCID](https://img.shields.io/badge/ORCID-0009--0001--8121--2878-brightgreen)](https://orcid.org/0009-0001-8121-2878) @@ -16,7 +16,7 @@ Governed Decision Intelligence is an open specification and Python reference implementation for recording how a consequential AI-assisted decision was framed, authorized, escalated, and preserved before execution. Its primary artifact is the Governed Decision Record (GDR), a schema-validated record of the decision question, system output, evidence, alternatives, risk posture, authority, gate classification, and downstream obligations. -The repository studies a bounded problem: governance programs often retain policies, model inventories, and system logs while leaving the individual decision difficult to reconstruct. GDI tests whether a contemporaneous decision record can make authorization, evidence access, uncertainty, intervention, and escalation inspectable at the point where an AI output may become an institutional action. +The repository studies a bounded problem: governance programs often retain policies, model inventories, runtime events, and system logs while leaving the individual institutional decision difficult to reconstruct. GDI tests whether a contemporaneous decision record can make authorization, evidence access, uncertainty, intervention, and escalation inspectable at the point where an AI output may become an institutional action. A literature search completed in May 2026 did not identify an open specification combining these elements in one decision record. That finding is provisional. The search method, adjacent work, limits, and update conditions are documented in [RESEARCH.md](RESEARCH.md). @@ -35,7 +35,7 @@ GDI approaches this question through artifact construction, schema validation, d | Gate classifier | Reference implementation | 114 deterministic internal tests | Test coverage does not establish domain calibration | | Interoperability driver | External PR under review | Local signed-receipt checks and submitted conformance driver | External conformance has not been accepted | | Framework mappings | Interpretive analysis | Source-linked evidence mapping | No legal compliance or certification claim | -| Decision Evidence Portability Specification | v0.1.1 working specification | Public scope, ontology, and acceptance conditions | Schema, overlays, and case studies remain in development | +| Decision Evidence Applicability Specification | v0.2.0 working specification | Revised thesis, determination states, adjacent-system boundary, and acceptance conditions | Schema, authoritative mappings, overlays, cases, and reviews remain in development | | Field validation | Planned | No deployment study published | External validity remains open | The status table is authoritative for current maturity claims. The [changelog](CHANGELOG.md) records repository releases. The PDF version, schema contract version, repository release, and working-specification version are separate identifiers. @@ -49,7 +49,9 @@ GDI contributes four connected artifacts: 3. A confidence-assessment input that can route uncertain decisions toward review or escalation. 4. A reference implementation and interoperability pattern for producing tamper-evident records. -The working [Decision Evidence Portability Specification](working-specifications/decision-evidence-portability-v0.1.md) extends this program by asking which evidence objects can support assurance across several governance regimes, which require local overlays, and where apparent equivalence fails. +The working [Decision Evidence Applicability Specification](working-specifications/decision-evidence-applicability-v0.2.md) extends this program by asking whether a defined evidence artifact is applicable to a specified governance requirement, what bounded assurance proposition it may support, what additional local evidence is required, where it is insufficient, and where apparent equivalence fails. + +DEAS does not claim that governance, controls, findings, evidence, or legal conclusions are portable across regimes. ## What GDI records @@ -91,7 +93,21 @@ external obligation or institutional policy monitoring, appeal, repair, and reassessment ``` -Runtime governors decide whether an action may proceed under configured policy. GDI records the decision evidence surrounding that control event. The distinction allows interoperability with projects such as Microsoft Agent Governance Toolkit, Credo AI Agent Governor, and signed-receipt standards without claiming to replace them. +Runtime governors decide whether an action may proceed under configured policy. GDI records the institutional decision evidence surrounding that control event. Signed-receipt systems may seal the GDR or related events. HIT may assess whether the documented human authority retained practical force. DEAS may evaluate what the resulting evidence can support under a specified governance requirement. + +The layers can interoperate without collapsing their claims. + +## Adjacent-system boundary + +| System | Primary layer | Boundary from GDI and DEAS | +|---|---|---| +| Microsoft Agent Governance Toolkit | Runtime action governance | Evaluates policy, identity, capability, approval, sandbox, allow or deny, and audit events. GDI may record those events but does not replace runtime enforcement. | +| ScopeBlind / Acta | Signed-receipt interoperability | Verifies schema, canonicalization, signatures, attribution, ordering, and chain linkage. GDI can supply payload semantics; DEAS can evaluate assurance use. Neither defines the receipt protocol. | +| Credo AI | Governance platform, policy packs, control mapping, evidence workflows, and runtime governance | GDI does not provide policy packs or compliance automation. DEAS does not provide a universal harmonized-control model. | +| Human Influence Telemetry | Documentary human-influence assessment | HIT evaluates whether formal human authority retained practical force. A complete GDR does not prove that result. | +| DEAS | Evidence-to-requirement qualification | Evaluates applicability, local evidence conditions, sufficiency boundaries, and non-equivalence. It does not make evidence portable. | + +See [Related work and scope boundary](docs/related-work.md). ## Quick start @@ -106,7 +122,7 @@ python -m pip install -r requirements-dev.txt python scripts/validate_repository.py ``` -The validation script checks the JSON Schema, the populated GDR example, top-level JSON and YAML schema contract parity, the gate-classifier test suite, deterministic interoperability fixtures, signed-receipt integrity, and citation metadata. +The validation script checks the JSON Schema, populated GDR example, top-level JSON and YAML schema contract parity, gate-classifier test suite, deterministic interoperability fixtures, signed-receipt integrity, and citation metadata. A successful run ends with: @@ -114,7 +130,7 @@ A successful run ends with: repository validation: PASS ``` -See [Validation and reproducibility](docs/validation-and-reproducibility.md) for the environment, commands, expected results, and the limits of each test. +See [Validation and reproducibility](docs/validation-and-reproducibility.md) for the environment, commands, expected results, and limits of each test. ## Minimal Python example @@ -152,9 +168,10 @@ The repository separates demonstrated properties from research hypotheses. | A populated GDR can be validated against the published schema | Schema validation in CI | Demonstrated for included fixtures | | Gate classification is deterministic for the included rules | 114 internal tests | Demonstrated for the reference implementation | | Gate records detect modification to sealed fields | Integrity tests | Demonstrated for the current hash profile | -| GDI receipts can interoperate with the external ACTA test-vector suite | Local checks and open external PR | External review pending | +| GDI receipts can interoperate with the external Acta test-vector suite | Local checks and open external PR | External review pending | | GDR fields can support evidence requests under NIST, ISO, and EU governance instruments | Source-linked interpretive mapping | Requires implementation and qualified review | | A GDR preserves substantive human judgment | No field study yet | Unresolved; Human Influence Telemetry addresses this test separately | +| DEAS can determine evidence applicability and non-equivalence across defined requirements | Working specification only | Unresolved until schema, mappings, cases, and review exist | | GDI improves outcomes in deployed institutions | No comparative deployment data | Unresolved | Full claim definitions, evidence classes, update conditions, and threats to validity appear in [RESEARCH.md](RESEARCH.md). @@ -178,16 +195,20 @@ GDI sits beside several established and emerging approaches: - Governance frameworks define organizational duties and risk-management processes. - Runtime governance systems enforce permissions, approvals, and action boundaries. -- Cryptographic receipt standards establish integrity, attribution, and ordering. -- GDI defines the decision record and the evidence needed to reconstruct institutional authorization. -- Human Influence Telemetry tests whether formal human authority retained practical force inside the workflow. +- Cryptographic receipt standards establish integrity, attribution, ordering, and chain linkage. +- Governance platforms translate policy into controls, workflows, and evidence requests. +- GDI defines the decision record and evidence needed to reconstruct institutional authorization. +- Human Influence Telemetry evaluates whether formal human authority retained practical force inside the workflow. +- DEAS evaluates the bounded assurance use of a defined evidence artifact under one specified requirement. -The detailed comparison, including Microsoft Agent Governance Toolkit, Credo AI Agent Governor, ACTA signed receipts, accountability theory, and meaningful human control, appears in [Related work](docs/related-work.md). +The detailed comparison appears in [Related work](docs/related-work.md). ## Limits GDI does not establish that a model output is true, a confidence score is calibrated, a reviewer understood the evidence, or a decision was lawful. A hash can detect modification to sealed fields; it cannot establish that the original fields were accurate. A complete record can document ceremonial review as easily as substantive review unless the workflow also measures evidence access, override ability, independent reasoning, and repair authority. +DEAS does not establish evidence portability, legal sufficiency, conformity, certification, or admissibility. Applicability means relevance to a bounded requirement, not proof that the requirement has been satisfied. + The specification also does not solve deceptive alignment, reward hacking, compromised telemetry, collusive agents, or failures that remain invisible to the governance layer. See [Limitations and threats to validity](docs/limitations.md). ## Repository map @@ -224,7 +245,8 @@ governed-decision-intelligence/ │ ├── confidence-threshold-model.md │ └── bilateral-pattern.md └── working-specifications/ - └── decision-evidence-portability-v0.1.md + ├── decision-evidence-applicability-v0.2.md + └── decision-evidence-portability-v0.1.md # deprecated migration pointer ``` ## Release and review policy diff --git a/RESEARCH.md b/RESEARCH.md index 949af78..38f6439 100644 --- a/RESEARCH.md +++ b/RESEARCH.md @@ -23,7 +23,7 @@ The repository contributes an open decision-record specification containing: - a four-level deliberation gate taxonomy; - a deterministic Python reference classifier; - a signed-receipt interoperability experiment; -- a working specification for cross-regime decision-evidence portability. +- a working specification for cross-regime decision-evidence applicability, sufficiency boundaries, and non-equivalence. A literature search completed in May 2026 did not identify an open specification that combined the same record objects, validation rules, gate model, and pre-execution evidence pattern. Later or missed work may narrow that finding. The review must remain revisable. @@ -39,6 +39,15 @@ The specification defines governance objects and their relationships. The JSON S The review covers scholarly publications, official governance instruments, open technical specifications, and vendor documentation that defines adjacent runtime or assurance layers. Novelty statements use the form “the review did not identify” because absence cannot be proven exhaustively. +The current adjacent-system boundary distinguishes: + +- Microsoft Agent Governance Toolkit as runtime action governance and policy enforcement; +- ScopeBlind/Acta as signed-receipt interoperability and cryptographic verification; +- Credo AI as policy-pack, control-mapping, workflow, evidence-collection, and runtime-governance infrastructure; +- GDI as the governed decision-record layer; +- Human Influence Telemetry as documentary assessment of practical human influence; +- Decision Evidence Applicability Specification as bounded evidence-to-requirement qualification. + ### Technical validation Repository validation checks: @@ -56,7 +65,11 @@ Passing tests establish behavior for the included artifacts and fixtures. They d ### Comparative evidence mapping -Framework mappings connect GDR fields to evidence that may be relevant under external instruments. Every mapping must identify source authority, source type, legal or normative force, regulated actor, governed object, GDI evidence object, rationale, limitation, confidence, and review date. The mapping output states evidence relevance. It never declares legal compliance. +Framework mappings connect GDR fields to evidence that may be relevant under external instruments. Every mapping must identify source authority, source type, legal or normative force, regulated actor, governed object, GDI evidence object, rationale, limitation, confidence, and review date. + +The mapping output states evidence relevance. It never declares legal compliance. + +DEAS strengthens this method by requiring one bounded assurance proposition, an applicability determination, local evidence conditions, a sufficiency boundary, and a point of non-equivalence. Applicability does not mean sufficiency, and structural reuse does not mean legal or normative portability. ## Evidence classes @@ -81,7 +94,8 @@ Framework mappings connect GDR fields to evidence that may be relevant under ext | C7 | A completed GDR proves substantive human judgment | No direct evidence | Unsupported | Requires Human Influence Telemetry and field observation | | C8 | GDI improves institutional outcomes | No comparative deployment study | Unresolved | Requires pre-registered field evaluation or credible comparative study | | C9 | GDI is the first specification of its kind | May 2026 literature review | Provisional novelty claim | Narrow when prior or later work covers the same contribution | -| C10 | Decision evidence can travel across governance regimes | DEPS v0.1.1 | Working hypothesis | Requires schema, overlays, cases, non-equivalence findings, and review | +| C10 | A defined decision-evidence artifact can be evaluated for applicability, local evidence conditions, sufficiency boundaries, and non-equivalence under a specified governance requirement | DEAS v0.2.0 working specification | Working hypothesis | Requires schema, authoritative mappings, overlays, cases, reviewer protocol, and external review | +| C11 | Decision evidence is legally or normatively portable across governance regimes | No valid general evidence | Unsupported | Would require a bounded regime pair and proof that authority, actor, scope, threshold, procedure, and consequence remain equivalent | ## Construct definitions @@ -97,6 +111,12 @@ Human authority means a named person or institutional role holds the power to au Decision evidence is the set of contemporaneous records used to assess how an AI-informed action was framed and authorized. It includes source provenance, decision context, control events, review actions, recorded reasons, and integrity metadata. +### Evidence applicability + +Evidence applicability is the relationship between one defined artifact and the actor, object, stage, scope, and assurance proposition governed by one identified external requirement. + +Applicability means relevance. It does not establish sufficiency, conformity, certification, admissibility, or legal compliance. + ### Gate classification Gate classification is an institutional rule assigning a required level of deliberation to a proposed action. It is independent of the model being governed. Model confidence may enter the rule as one input. @@ -105,16 +125,22 @@ Gate classification is an institutional rule assigning a required level of delib The novelty review can miss unpublished work, proprietary systems, non-English sources, newly released material, or work using different terminology. A complete record may capture ceremonial oversight. Model-reported confidence may be absent or poorly calibrated. A valid hash proves consistency of sealed fields, not truth. Included examples do not establish performance in other domains. Framework mappings require qualified review. Compromised telemetry or governance-layer bypass can defeat the record. +Cross-regime evidence work adds further threats: shared terms may hide different actor duties or legal force; vendor taxonomies may be mistaken for authoritative requirements; cryptographic integrity may be mistaken for substantive sufficiency; policy-pack reuse may be mistaken for legal equivalence; and an evidence relationship accepted in one jurisdiction may fail under another. + ## Review protocol Internal validation supports “demonstrated in the included implementation.” External replication supports “independently reproduced.” Accepted conformance supports “conforming to the tested profile.” Qualified legal or standards review supports a bounded conformity interpretation. Comparative field evidence supports outcome claims. +DEAS determinations require review competence appropriate to the source authority. A technical reviewer may evaluate schema and provenance behavior. Legal sufficiency requires qualified legal review. Standards conformity requires access to the authoritative standard and qualified interpretation. + Repository language must be downgraded when the supporting condition no longer holds. ## Release criteria A stable research release requires all validation checks to pass, synchronized schemas and examples, migration notes for normative changes, an updated claim register and limitations, accurate external-review status, matching citation metadata, and a DOI pointing to the exact released artifact. +A release containing DEAS work must also preserve the adjacent-system boundaries, prohibit unqualified portability claims, and distinguish evidence relevance from assurance sufficiency. + ## Update cadence -Adjacent-work review is quarterly. Framework-source review occurs at least annually and after material legal or standards changes. The claim register is reviewed for every numbered release. \ No newline at end of file +Adjacent-work review is quarterly. Framework-source review occurs at least annually and after material legal or standards changes. The claim register is reviewed for every numbered release. diff --git a/docs/framework-compatibility.md b/docs/framework-compatibility.md index fcd081e..4dbefbb 100644 --- a/docs/framework-compatibility.md +++ b/docs/framework-compatibility.md @@ -15,6 +15,8 @@ Each mapping asks: which source creates the requirement, which actor and context | Hypothesis | The relationship lacks sufficient review or testing | | Outside scope | GDI does not address the requirement | +These statuses describe evidence relevance only. They do not establish sufficiency. + ## NIST AI Risk Management Framework 1.0 Official source: https://www.nist.gov/itl/ai-risk-management-framework @@ -70,15 +72,49 @@ Official source: https://oecd.ai/en/ai-principles | Robustness, security, and safety | Risk posture, escalation, stopping, monitoring record | Principle-level relationship | GDI does not conduct security or safety evaluation | | Accountability | Named authority, review roles, audit trail, repair ownership where implemented | Direct conceptual relationship | Accountability also requires a forum, consequences, and institutional practice | -## Decision Evidence Portability +## Decision Evidence Applicability Specification + +The working Decision Evidence Applicability Specification develops a stricter evidence-to-requirement record. + +Every DEAS determination must record: -The working Decision Evidence Portability Specification develops a stricter mapping record. Every cross-regime claim must record source authority, source type, force, regulated actor, governed object, required evidence, assurance test, local condition, non-equivalence, and confidence. +- one identified evidence artifact; +- one authoritative requirement; +- the source authority, source type, and legal or normative force; +- the regulated actor and governed object; +- the decision or lifecycle stage; +- one bounded assurance proposition; +- the applicability rationale; +- required and supplied evidence; +- the assurance test; +- a determination state; +- local evidence conditions; +- the sufficiency boundary; +- at least one point of non-equivalence where applicable; +- mapping confidence, reviewer competence, and update conditions. -That work remains in development. It cannot be cited as a completed cross-regime assurance method until its schema, overlays, cases, and review conditions are satisfied. +DEAS does not make evidence portable. It may conclude that an artifact is applicable, applicable only with local evidence, insufficient, non-equivalent, not applicable, or indeterminate. + +That work remains in development. It cannot be cited as a completed cross-regime assurance method until its schema, authoritative mappings, overlays, cases, reviewer protocol, and review conditions are satisfied. ## Non-equivalence rules -A mapping fails when shared terminology hides a material difference in legal force, regulated actor, governed object, timing, evidence burden, assurance method, enforcement consequence, or remedy. Every mature mapping must record at least one point of divergence. +A mapping fails when shared terminology hides a material difference in legal force, regulated actor, governed object, timing, evidence burden, assurance method, procedural posture, enforcement consequence, or remedy. + +Structural reuse is not legal or normative equivalence. Cryptographic integrity is not assurance sufficiency. A policy-pack control mapping is not automatically an authoritative legal interpretation. + +Every mature mapping must record at least one point of divergence or an evidenced explanation for why no material divergence applies. + +## Adjacent-system boundary + +- Microsoft Agent Governance Toolkit produces runtime policy, approval, identity, capability, enforcement, and audit events. +- ScopeBlind/Acta establishes signed-receipt schema, canonicalization, signature, attribution, ordering, and chain-verification properties. +- Credo AI provides policy packs, control mappings, governance workflows, evidence collection, compliance mapping, monitoring, and runtime governance. +- GDI structures one consequential decision record. +- HIT evaluates practical human influence. +- DEAS evaluates the bounded assurance use of a defined evidence artifact under a specified requirement. + +No layer inherits another layer's claims merely because artifacts are linked. ## Review status @@ -88,5 +124,8 @@ A mapping fails when shared terminology hides a material difference in legal for | ISO/IEC 42001:2023 | 2026-07-16 | Public-summary review; licensed-standard review pending | | EU AI Act | 2026-07-16 | Author review; qualified legal review pending | | OECD AI Principles | 2026-07-16 | Author review; external review pending | +| Microsoft Agent Governance Toolkit | 2026-07-17 | Public repository review; external confirmation pending | +| ScopeBlind/Acta | 2026-07-17 | Public repository and interoperability review | +| Credo AI | 2026-07-17 | Public product and Policy Pack documentation review | -Part of the Governed Decision Intelligence research repository. Apache License 2.0. \ No newline at end of file +Part of the Governed Decision Intelligence research repository. Apache License 2.0. diff --git a/docs/limitations.md b/docs/limitations.md index 467c979..990a1a3 100644 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -12,6 +12,8 @@ Human Influence Telemetry addresses this gap by testing evidence access, indepen A record hash can detect later modification to fields included in the hash profile. It cannot establish that the original record was accurate, complete, timely, or honestly produced. Integrity depends on canonicalization, field selection, key custody where signatures are used, and protection of the record-generation path. +A cryptographically valid ScopeBlind/Acta receipt can strengthen attribution, integrity, ordering, and chain evidence. It does not establish payload truth, governance quality, legal sufficiency, or practical human influence. + ## Confidence scores require calibration The reference implementation accepts confidence values between 0 and 1 and compares them with institutional thresholds. These values may represent different quantities across models and tasks. Some systems provide no meaningful confidence estimate. Deployment requires domain-specific calibration, error-cost analysis, drift monitoring, and a rule for absent or invalid confidence. @@ -20,6 +22,34 @@ The reference implementation accepts confidence values between 0 and 1 and compa Mappings to NIST AI RMF, ISO/IEC 42001, the European Union Artificial Intelligence Act, and OECD principles identify potentially relevant evidence. They do not establish applicability, conformity, certification, or legal sufficiency. ISO clause-level conclusions require access to the licensed standard. Legal conclusions require qualified review. +## Applicability is not portability + +Decision Evidence Applicability Specification evaluates one defined evidence artifact against one identified governance requirement. + +An applicability determination does not establish that the evidence, control, finding, or legal conclusion can be transferred across regimes without loss of meaning or force. + +The same artifact may be: + +- relevant under several regimes; +- sufficient under one and insufficient under another; +- structurally reusable but legally non-equivalent; +- acceptable only when combined with local evidence; +- cryptographically authentic but substantively weak; +- technically complete but outside the governed actor or lifecycle stage. + +DEAS remains a working specification. Its schema, authoritative mappings, overlays, cases, reviewer protocol, and validation suite are incomplete. + +## Adjacent-system boundaries can change + +Microsoft Agent Governance Toolkit, ScopeBlind/Acta, Credo AI, and other adjacent systems may change scope, terminology, or implementation. Repository comparisons are dated and provisional. + +GDI and DEAS must not claim functions performed by those systems without an implemented and tested capability. In particular, the repository must separate: + +- runtime action enforcement from decision reconstruction; +- signed-receipt verification from payload semantics; +- policy packs and harmonized control mappings from evidence-applicability determinations; +- documentary evidence relevance from legal or standards sufficiency. + ## Examples have limited external validity The included insurance and agent-tool examples are synthetic. They test representation and control logic. They do not establish effects in healthcare, employment, public benefits, finance, biopharma, public administration, or national-security settings. @@ -38,8 +68,8 @@ Decision records can contain personal data, sensitive evidence, model inputs, an ## Independent validation remains pending -The reference implementation and mappings were developed by the project author. External review, replication, adversarial testing, and comparative deployment studies remain open research needs. +The reference implementation and mappings were developed by the project author. External review, replication, adversarial testing, legal or standards review, and comparative deployment studies remain open research needs. ## Update condition -These limitations must be revised whenever the schema, hash profile, framework mappings, reference implementation, or empirical evidence changes. \ No newline at end of file +These limitations must be revised whenever the schema, hash profile, framework mappings, reference implementation, adjacent-system scope, DEAS determination model, or empirical evidence changes. diff --git a/docs/related-work.md b/docs/related-work.md index 4b89229..02c8469 100644 --- a/docs/related-work.md +++ b/docs/related-work.md @@ -1,34 +1,44 @@ # Related Work and Scope Boundary -GDI occupies the decision-evidence layer between runtime controls and institutional assurance. Adjacent systems may enforce policy, evaluate models, establish cryptographic integrity, or define organizational duties. GDI records what the institution knew, authorized, and required around one consequential decision. +GDI occupies the decision-evidence layer between runtime controls and institutional assurance. Adjacent systems may enforce policy, evaluate models, establish cryptographic integrity, harmonize governance requirements, or define organizational duties. GDI records what the institution knew, authorized, and required around one consequential decision. ## Organizational governance frameworks -NIST AI RMF, ISO/IEC 42001, the OECD AI Principles, and the European Union Artificial Intelligence Act define governance functions, management processes, actor duties, or legal requirements. They provide the authority and assessment context. GDI is one possible source of operational evidence within those programs. +NIST AI RMF, ISO/IEC 42001, the OECD AI Principles, and the European Union Artificial Intelligence Act define governance functions, management processes, actor duties, or legal requirements. They provide authority and assessment context. GDI is one possible source of operational evidence within those programs. ## Runtime agent governance -Microsoft Agent Governance Toolkit and Credo AI Agent Governor address harness-level controls such as permissions, policy checks, approval points, action boundaries, and enforcement events. Their central question is whether an agent action is permitted under configured policy. +Microsoft Agent Governance Toolkit addresses runtime action governance through policy evaluation, identity and capability checks, approval controls, action interception, sandboxing, allow or deny decisions, and audit events. -GDI asks a related question: what decision was being made, which evidence and uncertainty shaped it, what human authority applied, which deliberation gate fired, and what record can later support audit, appeal, or repair? +Credo AI provides policy packs, policy-to-code translation, risk and control mappings, governance workflows, evidence collection, compliance mapping, monitoring, and runtime agent-governance capabilities. + +Their central runtime question is whether an agent action is permitted, constrained, escalated, or sanctioned under configured governance. + +GDI asks a different question: what consequential decision was being made, which evidence and uncertainty shaped it, what human or institutional authority applied, which deliberation gate fired, and what record can later support audit, appeal, repair, or reassessment? The systems can interoperate through this sequence: ```text -external obligation or policy +external requirement or institutional policy → configured runtime control → policy or enforcement event → Governed Decision Record - → assurance, audit, appeal, or repair + → human-influence assessment + → evidence-applicability determination + → assurance, audit, appeal, repair, or reassessment ``` -GDI does not compile policy into an agent harness or prescribe a universal runtime governor. +GDI does not compile policy into an agent harness, provide policy packs, prescribe a universal runtime governor, or automate compliance. ## Signed receipts and cryptographic evidence -ACTA signed receipts and related receipt standards establish integrity, attribution, ordering, and chain linkage. They answer whether a record was altered and which issuer produced it. A GDR supplies the decision semantics placed inside or referenced by the receipt. +ScopeBlind agent-governance test vectors and Acta signed receipts address schema, canonicalization, signature validity, attribution, ordering, and chain linkage across implementations. + +A GDR can supply decision semantics inside or referenced by a receipt. The receipt layer can attest to defined integrity properties of the payload or event sequence. -Cryptographic integrity and governance quality remain distinct. A valid signature can attest to a poor or inaccurate decision record. The interoperability experiment tests whether the two evidence layers can be combined. +Cryptographic integrity and governance quality remain distinct. A valid signature can attest to a poor, incomplete, or inaccurate decision record. Receipt interoperability does not establish that the payload satisfies a legal, institutional, or assurance requirement. + +The GDI interoperability experiment tests whether the two evidence layers can be composed. It does not claim that GDI defines the receipt protocol. ## Model and system documentation @@ -42,18 +52,37 @@ Bovens' accountability model treats accountability as a relationship in which an Meaningful human control research asks whether human reasons and responsibility remain connected to system behavior. Human Influence Telemetry extends this concern into observable workflow evidence: evidence access, independent reasoning, override capability, appeal ownership, repair responsibility, and system-change authority. -GDI records the declared authority structure. HIT tests whether that authority had practical force. +GDI records the declared authority structure. HIT evaluates whether that authority retained practical force. ## Decision science and phase-gate methods GDI draws on structured decision analysis and phase-gate practices that make alternatives, evidence, uncertainty, conditions, stopping, and approval explicit. The reference implementation adapts these ideas to AI-assisted and agent-mediated decisions. Domain calibration remains an institutional task. -## Decision Evidence Portability Specification +## Decision Evidence Applicability Specification + +The working Decision Evidence Applicability Specification asks whether a defined evidence artifact is relevant to one identified governance requirement, what bounded assurance proposition it may support, what additional local evidence is required, where it is insufficient, and where apparent equivalence fails. -The working Decision Evidence Portability Specification asks which decision-evidence objects can support assurance across several governance regimes, where local overlays are required, and where apparent equivalence fails. Its schema, overlays, case studies, and review conditions remain in development. +DEAS does not claim that governance, controls, findings, evidence, or legal conclusions are portable across regimes. + +DEAS also does not replace: + +- Microsoft AGT runtime enforcement; +- ScopeBlind/Acta receipt interoperability and cryptographic verification; +- Credo AI policy packs, control mappings, governance workflows, or runtime capabilities; +- GDI decision-record construction; +- HIT assessment of practical human influence. + +Its schema, authoritative mappings, overlays, cases, reviewer protocol, and validation suite remain in development. ## Contribution boundary -The project claims an integrated open artifact composed of a decision-record schema, gate taxonomy, reference classifier, and evidence-preservation pattern. It does not claim ownership of accountability theory, human oversight, runtime policy enforcement, cryptographic receipts, model documentation, or phase-gate decision methods. +The project claims an integrated open artifact composed of a decision-record schema, gate taxonomy, reference classifier, evidence-preservation pattern, and working evidence-applicability method. It does not claim ownership of accountability theory, human oversight, runtime policy enforcement, cryptographic receipts, policy-pack governance, model documentation, phase-gate decision methods, or legal interpretation. + +The dated literature review did not identify an open specification combining the same elements. This finding remains provisional and must narrow when overlapping prior work is found. + +## Public references reviewed -The dated literature review did not identify an open specification combining the same elements. This finding remains provisional and must narrow when overlapping prior work is found. \ No newline at end of file +- Microsoft Agent Governance Toolkit: https://github.com/microsoft/agent-governance-toolkit +- ScopeBlind agent-governance test vectors: https://github.com/ScopeBlind/agent-governance-testvectors +- Credo AI: https://www.credo.ai/ +- Credo AI Policy Packs: https://www.credo.ai/glossary/credo-ai-policy-pack diff --git a/working-specifications/decision-evidence-applicability-v0.2.md b/working-specifications/decision-evidence-applicability-v0.2.md new file mode 100644 index 0000000..7cc8bc8 --- /dev/null +++ b/working-specifications/decision-evidence-applicability-v0.2.md @@ -0,0 +1,241 @@ +# Decision Evidence Applicability Specification + +**Acronym:** DEAS +**Subtitle:** A cross-regime assurance specification for evaluating the applicability, sufficiency, and non-equivalence of evidence generated by AI-assisted decisions and agentic workflows +**Status:** Working specification +**Version:** 0.2.0 +**Date:** 2026-07-17 +**Author:** Mark Julius Banasihan +**ORCID:** 0009-0001-8121-2878 + +## 1. Bounded thesis + +A recurring set of decision-evidence objects may be relevant under multiple regulatory, standards, contractual, and institutional governance regimes. + +DEAS evaluates: + +1. whether a defined evidence artifact is applicable to a specified governance requirement; +2. what bounded assurance claim that artifact may support; +3. what additional local evidence is required; +4. whether the artifact is insufficient under the governing authority, actor, scope, threshold, timing, or procedural conditions; +5. where apparently similar requirements are non-equivalent. + +DEAS does not make governance, controls, findings, evidence, or legal conclusions portable across regimes. + +## 2. Problem + +Crosswalks commonly compare terms, principles, controls, or framework categories. Those mappings can conceal differences in: + +- legal or normative force; +- regulated actor; +- governed object; +- jurisdiction and territorial scope; +- decision stage; +- evidence burden; +- assurance method; +- procedural posture; +- enforcement consequence; +- remedy. + +The same artifact may be relevant in several regimes while remaining sufficient in none, sufficient in one, or acceptable only when combined with regime-specific evidence. + +The missing layer is a machine-readable determination of what a specific evidence artifact can support under a specific requirement. + +## 3. Unit of analysis + +The unit of analysis is one relationship between: + +- one identified decision-evidence artifact or artifact set; and +- one identified external governance requirement. + +A DEAS determination is invalid when either side is left as an unbounded framework category. + +## 4. Scope boundary + +DEAS does not: + +- define an agent runtime governor; +- intercept, permit, deny, sandbox, or terminate agent actions; +- compile policy into harness or runtime configuration; +- provide a universal control library or policy pack; +- create or verify cryptographic receipts; +- define signed-receipt canonicalization, signatures, or chain rules; +- establish cross-implementation receipt conformance; +- certify legal compliance or standards conformity; +- declare evidence legally admissible or sufficient; +- assume that evidence accepted in one regime carries the same force in another. + +DEAS begins with an identified evidence artifact and an identified external requirement. + +The traceability path is: + +**external requirement → required assurance proposition → evidence artifact → applicability test → local evidence condition → sufficiency boundary → non-equivalence finding** + +Runtime controls and signed receipts are evidence-producing or evidence-integrity mechanisms. They are inputs to DEAS rather than functions that DEAS performs. + +## 5. Relationship to the portfolio + +- **Governed Decision Intelligence (GDI)** structures the decision question, evidence, alternatives, uncertainty, authority, outcome, and obligations in a Governed Decision Record. +- **Human Influence Telemetry (HIT)** evaluates whether documented human access, judgment, authority, correction, repair, and reform retained practical force. +- **DEAS** evaluates what a defined GDI, HIT, runtime-control, monitoring, or receipt artifact may support under a specified governance requirement. + +The artifacts remain separable. A GDR can exist without HIT. HIT can assess records other than GDRs. DEAS can evaluate evidence artifacts produced by either system or by an external system. + +## 6. Boundary against adjacent systems + +### 6.1 Microsoft Agent Governance Toolkit + +Microsoft AGT operates at the runtime action-governance layer. It evaluates configured policy, identity, capability, approval, sandbox, and execution conditions and records allow, deny, approval, and audit outcomes. + +DEAS does not perform those functions. An AGT policy or enforcement event may be evaluated as evidence under DEAS. + +### 6.2 ScopeBlind / Acta signed receipts + +ScopeBlind test vectors and Acta receipt implementations evaluate signed-receipt schema, canonicalization, signature validity, attribution, ordering, and chain linkage across implementations. + +DEAS does not provide receipt interoperability or cryptographic conformance. A verified receipt may strengthen provenance or integrity evidence while leaving payload truth, governance quality, applicability, and sufficiency unresolved. + +### 6.3 Credo AI + +Credo AI provides policy packs, policy-to-code translation, risk and control mappings, governance workflows, compliance mapping, evidence collection, monitoring, and runtime agent-governance capabilities. + +DEAS does not replace those platform functions or provide a universal harmonized-control model. It evaluates the bounded assurance use of a specific evidence artifact under a specific requirement and must preserve local conditions and non-equivalence. + +## 7. Core terms + +### Evidence artifact + +A defined record, event, log, assessment, receipt, decision record, approval record, review note, model or system artifact, monitoring result, appeal record, repair record, or reform record with identifiable provenance. + +### External governance requirement + +A requirement created by an identified law, regulation, standard, contract, policy, regulatory instrument, governance framework, or assurance procedure. + +### Assurance proposition + +The bounded statement the evidence is offered to support. Examples include that a named authority was assigned, a review occurred, an intervention mechanism existed, a control executed, an appeal was available, or a record retained integrity properties. + +### Applicability + +The relationship between an evidence artifact and the actor, object, stage, scope, and proposition governed by the external requirement. + +Applicability means relevance. It does not mean sufficiency. + +### Sufficiency boundary + +The point beyond which the artifact cannot support the assurance proposition without additional evidence, qualified interpretation, procedural testing, or legal or standards review. + +### Local evidence condition + +An additional requirement imposed by the specific regime, jurisdiction, actor role, sector, decision stage, or assurance procedure. + +### Non-equivalence + +A material difference that prevents two requirements, controls, findings, or evidence requests from being treated as interchangeable. + +## 8. Determination states + +Each mapping must return one primary determination: + +- `applicable_but_unreviewed`: the artifact is relevant, but assurance sufficiency has not received the required review; +- `applicable_with_local_evidence`: the artifact is relevant only when combined with identified local evidence; +- `applicable_and_bounded`: the artifact supports the stated assurance proposition within declared conditions and limitations; +- `insufficient`: the artifact is relevant but cannot support the proposition under the stated threshold; +- `non_equivalent`: a material difference prevents reuse of the source relationship as an equivalent target relationship; +- `not_applicable`: the artifact does not address the governed actor, object, stage, or proposition; +- `indeterminate`: the available authority or evidence cannot resolve the relationship. + +No determination state means legal compliance, certification, conformity, admissibility, or universal reuse. + +## 9. Mapping record + +Each DEAS mapping must record: + +- mapping identifier and specification version; +- evidence artifact identifier, type, version, producer, provenance, and integrity status; +- source authority and authoritative citation; +- source type and legal or normative force; +- requirement identifier and text or bounded paraphrase; +- regulated actor; +- governed object; +- decision or lifecycle stage; +- jurisdiction or institutional scope; +- assurance proposition; +- applicability rationale; +- required evidence; +- evidence supplied; +- assurance test; +- determination state; +- local evidence condition; +- sufficiency boundary; +- point of non-equivalence; +- mapping confidence; +- reviewer competence and conflict disclosure; +- review date and update condition. + +## 10. Evidence classes + +DEAS must distinguish at least: + +- authoritative requirement evidence; +- implementation evidence; +- runtime control or enforcement evidence; +- decision-record evidence; +- human-influence assessment evidence; +- monitoring and incident evidence; +- cryptographic integrity evidence; +- appeal, repair, and reform evidence; +- interpretive mapping; +- unresolved hypothesis. + +Evidence classes must not be collapsed merely because they share a filename, control label, or framework term. + +## 11. Initial scope + +Version 0.2 will test one agentic workflow in which an AI system retrieves evidence, recommends a decision, prepares an action, and requires human authorization before execution. + +The initial comparison set is: + +- NIST AI Risk Management Framework; +- ISO/IEC 42001 and selected supporting standards; +- European Union Artificial Intelligence Act; +- selected Chinese agent-governance requirements as a comparative annex; +- one institutional policy requirement; +- one runtime-control event; +- one signed-receipt integrity result; +- one Governed Decision Record; +- one Human Influence Telemetry assessment record. + +## 12. Acceptance conditions + +DEAS will advance beyond working-specification status only when: + +1. every mapping links to an authoritative source; +2. each mapping identifies one bounded assurance proposition; +3. each regime contains at least one explicit point of non-equivalence or an evidenced reason why none applies; +4. one core decision record validates under at least two jurisdictional or institutional overlays without changing historical facts; +5. each overlay may require additional local evidence; +6. missing authority, intervention, logging, provenance, or integrity evidence produces a defined determination; +7. external requirements map to identifiable evidence artifacts without claiming universal control equivalence; +8. outputs state applicability, local evidence conditions, sufficiency boundaries, and non-equivalence; +9. outputs never declare legal compliance, certification, conformity, or admissibility; +10. adjacent-system claims are reviewed against Microsoft AGT, ScopeBlind/Acta, and Credo AI; +11. the draft receives legal or standards review and technical assurance review. + +## 13. Migration from DEPS 0.1.1 + +The prior working title **Decision Evidence Portability Specification (DEPS)** is deprecated. + +The rename to **Decision Evidence Applicability Specification (DEAS)** corrects an overbroad implication. The project does not claim that evidence retains the same legal meaning, assurance weight, sufficiency, or consequence when moved across regimes. + +The ontology remains under development. Existing notes may be migrated when they are rewritten as bounded evidence-to-requirement determinations under this specification. + +## 14. Current status + +This document establishes the revised name, bounded thesis, adjacent-system boundary, determination states, and acceptance conditions. + +The machine-readable schema, overlays, authoritative mappings, worked workflow, cases, reviewer protocol, and validation suite remain in development. + +## 15. Citation + +Banasihan, Mark Julius. “Decision Evidence Applicability Specification.” Working specification, version 0.2.0, July 17, 2026. diff --git a/working-specifications/decision-evidence-portability-v0.1.md b/working-specifications/decision-evidence-portability-v0.1.md index 44f5896..9010c1d 100644 --- a/working-specifications/decision-evidence-portability-v0.1.md +++ b/working-specifications/decision-evidence-portability-v0.1.md @@ -1,110 +1,22 @@ -# Decision Evidence Portability Specification +# Decision Evidence Portability Specification — Deprecated Title -**Subtitle:** A cross-regime assurance specification for evidence generated by AI-assisted decisions and agentic workflows -**Status:** Working specification -**Version:** 0.1.1 -**Date:** 2026-07-15 -**Author:** Mark Julius Banasihan -**ORCID:** 0009-0001-8121-2878 +**Status:** Deprecated working title +**Last version under this title:** 0.1.1 +**Superseded by:** [`Decision Evidence Applicability Specification v0.2.0`](decision-evidence-applicability-v0.2.md) +**Deprecation date:** 2026-07-17 -## Bounded thesis +The title **Decision Evidence Portability Specification (DEPS)** is deprecated. -A recurring set of decision-evidence objects can support AI assurance across multiple regulatory and standards regimes. Each regime assigns its own legal meaning, thresholds, obligations, and consequences to those objects. +The word *portability* implied more than the working specification could support. Decision evidence may remain structurally reusable while changing legal meaning, assurance weight, sufficiency threshold, regulated-actor relationship, procedural effect, or enforcement consequence across governance regimes. -This specification tests which decision evidence can travel across regimes, which evidence requires a jurisdiction-specific overlay, and where apparent equivalence fails. It does not treat governance regimes as interchangeable. +The canonical working specification is now **Decision Evidence Applicability Specification (DEAS)**. -## Problem +DEAS evaluates: -AI governance crosswalks usually compare terms, controls, or framework categories. Those mappings can conceal differences in legal force, regulated actor, assurance burden, and enforcement consequence. +1. whether a defined evidence artifact is applicable to a specified governance requirement; +2. what bounded assurance claim it may support; +3. what additional local evidence is required; +4. where the artifact is insufficient; +5. where apparent equivalence fails. -The missing layer is a machine-readable account of the decision evidence itself: who held authority, what evidence was available, how AI influenced the decision, where intervention remained possible, what reasons were recorded, and who owned appeal, repair, and system change. - -## Scope boundary - -This specification does not define an agent runtime governor, compile policy into harness configuration, or prescribe a universal set of agent controls. - -It specifies the evidence that configured controls, enforcement events, and human interventions must generate for cross-regime assurance and decision reconstruction. Runtime controls are treated as evidence-producing mechanisms. - -The central traceability path is: - -**external obligation → configured control → enforcement event → decision evidence → assurance claim** - -## Relationship to existing work - -This working specification extends Governed Decision Intelligence. - -- Governed Decision Intelligence structures the governed decision record. -- Human Influence Telemetry tests whether formal human authority retained causal force. -- Decision Evidence Portability maps reusable decision evidence to external governance obligations while preserving jurisdiction-specific differences. - -## Adjacent work - -Credo AI's Agent Governor and Agent Governance Configuration work address harness-level agent governance. That work translates policy and risk posture into runtime configurations, applies controls before agent actions execute, and produces enforcement telemetry. - -Decision Evidence Portability begins at the resulting control event and decision record. Its question is whether the institution can prove authorization, intervention, reasoning, monitoring, appeal, repair, and system-change authority across multiple governance regimes. - -The two approaches can interoperate. A runtime governor may produce control and enforcement events that become inputs to a Decision Evidence Portability record. - -## Initial ontology - -The v0.1 ontology contains ten top-level objects: - -1. actor -2. institutional_role -3. decision_authority -4. ai_influence -5. evidence_access -6. intervention_capability -7. recorded_reasoning -8. approval_conditions -9. monitoring_and_incident_record -10. appeal_repair_and_system_change_authority - -Each object must resolve into observable artifacts, validation rules, failure signs, and provenance. - -## Mapping record - -Each cross-regime mapping must record: - -- source authority -- source type -- legal or normative force -- regulated actor -- governed object -- required evidence -- assurance test -- jurisdiction-specific condition -- point of non-equivalence -- mapping confidence - -## Initial scope - -Version 0.1 will test one agentic workflow in which an AI system retrieves evidence, recommends a decision, prepares an action, and requires human authorization before execution. - -The initial comparison set is: - -- NIST AI Risk Management Framework -- ISO/IEC 42001 and selected supporting standards -- European Union Artificial Intelligence Act -- selected Chinese agent-governance requirements as a comparative annex - -## Acceptance conditions - -The working specification will advance beyond v0.1 only when: - -1. every mapping links to an authoritative source; -2. each regime contains at least one explicit point of non-equivalence; -3. one core decision record validates under at least two jurisdictional overlays without changing historical facts; -4. each overlay can require additional local evidence; -5. missing authority, intervention, or logging evidence produces a defined failure; -6. external obligations map to identifiable controls, enforcement events, and resulting evidence artifacts in the agent workflow; -7. outputs state evidence sufficiency and never declare legal compliance; -8. the draft receives legal or standards review and technical assurance review. - -## Current status - -This document establishes authorship, scope, terminology, the initial research claim, and the boundary between runtime agent governance and cross-regime decision assurance. The ontology, schema extension, mappings, test workflow, and case studies remain in development. - -## Citation - -Banasihan, Mark Julius. “Decision Evidence Portability Specification.” Working specification, version 0.1.1, July 15, 2026. +This file remains as a migration pointer so earlier links and citations do not silently break. It is not the current specification and must not be used for new implementations or claims.