-
Notifications
You must be signed in to change notification settings - Fork 16
(MOT-4162) harness: minimal identity prompt + on-demand agent playbook skills #558
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
rohitg00
wants to merge
5
commits into
main
Choose a base branch
from
harness-minimal-prompt
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
e2f307a
feat(harness): split identity prompt into minimal core plus on-demand…
rohitg00 2f6dd57
test(harness): drop dead registry-id allowance from prompt invariants
rohitg00 ae9db57
(MOT-4162) fix(harness): repair SDK reference URLs in building playbook
rohitg00 5968bda
(MOT-4162) docs(harness): documented-call exemption + validator link …
rohitg00 afe5cf7
ci: rerun against current main
rohitg00 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
Large diffs are not rendered by default.
Oops, something went wrong.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,77 @@ | ||
| --- | ||
| name: harness/building | ||
| description: >- | ||
| When nothing registered fits: search and install workers from the public registry, edit code files with coder::*, and read the SDK reference before writing any worker code. | ||
| --- | ||
|
|
||
| # Building new things | ||
|
|
||
| First check what already exists with `engine::functions::list` and | ||
| `engine::triggers::list`. Do not carry patterns from other ecosystems (standalone servers, | ||
| package managers, ad-hoc processes) — iii has its own way, and foreign patterns do not run | ||
| here. | ||
|
|
||
| If no registered function fits, search the public registry: | ||
|
|
||
| Step 1. Call `directory::registry::workers::list { search: "<capability>" }` to find a | ||
| worker. | ||
| Step 2. Call `directory::registry::workers::info { name: "<name>" }` to see its functions, | ||
| config, and dependencies before installing. Both registry calls are documented here, so you | ||
| do not need to fetch their contracts first. | ||
| Step 3. Installing runs new code, so say what you are about to install and why. Then install | ||
| it with `worker::add { source: { kind: "registry", name: "<name>" } }`. | ||
| Step 4. Check it worked: confirm the new function ids appear with | ||
| `engine::functions::list { prefix: "<worker>::" }`. Then fetch each contract with | ||
| `engine::functions::info` before calling. The registry detail is a preview, not the contract. | ||
|
|
||
| If no `directory::*` function is registered: look in `worker::list` for a stopped | ||
| directory worker and start it. If it is not installed, install it with | ||
| `worker::add { source: { kind: "registry", name: "iii-directory" } }`. If the registry is | ||
| still unreachable, tell the user and continue with what is registered. | ||
|
|
||
| <example> | ||
| user: Email me the weekly report. | ||
| assistant: [calls engine::functions::list { search: "email" } — nothing registered fits] | ||
| [calls directory::registry::workers::list { search: "email" } and finds "email"] | ||
| [calls directory::registry::workers::info { name: "email" } to judge fit before installing] | ||
| I am installing the "email" worker from the public registry so I can send the report. | ||
| [calls engine::functions::info { function_id: "worker::add" } for the install contract] | ||
| [calls worker::add { source: { kind: "registry", name: "email" } }] | ||
| [calls engine::functions::list { prefix: "email::" } — the new function ids appear] | ||
| [calls engine::functions::info { function_id: "email::send" } to get the contract] | ||
| [calls agent_trigger with function: "email::send", payload: { ...per the contract }] | ||
| </example> | ||
|
|
||
| To create, edit, move, or delete code files, use the `coder::*` functions — they are | ||
| served by the shell worker (no separate install). Confirm they are available with | ||
| `engine::functions::list { prefix: "coder::" }`. Its functions include `coder::read-file`, | ||
| `coder::search`, `coder::list-folder`, `coder::tree`, `coder::create-file`, | ||
| `coder::update-file`, `coder::move`, and `coder::delete-file` — the prefix check shows | ||
| the full inventory. Use `coder::move` for renames and moves, never delete-then-recreate. Plain | ||
| file browsing outside code work (like `shell::fs::ls`) is still fine. Fetch each contract | ||
| first, as always. | ||
|
|
||
| To author a worker: import ONLY `registerWorker` from the SDK. Its return value has the | ||
| methods `registerFunction`, `registerTrigger`, and `trigger` — call them as | ||
| `iii.registerFunction(...)`. They are NOT top-level exports. Destructuring them throws | ||
| `TypeError: registerFunction is not a function`. Give every function a `description`, | ||
| `request_format`, and `response_format` — that becomes the contract that | ||
| `engine::functions::info` shows to callers. Before writing code, inspect the runtime with | ||
| `engine::workers::info { name }`. | ||
|
|
||
| Before you write the FIRST line of worker code — a new worker, or new registrations on an | ||
| existing one — read the SDK reference for the language you will use. Do not write SDK code | ||
| from memory: names and config keys from memory are often wrong, and a trigger registered with | ||
| wrong keys never fires. Fetch the reference as Markdown. | ||
| Pick the URL for the implementation language: | ||
| - https://iii.dev/docs/reference/sdk-node — Node/TypeScript | ||
| - https://iii.dev/docs/reference/sdk-python — Python | ||
| - https://iii.dev/docs/reference/sdk-rust — Rust | ||
| - https://iii.dev/docs/reference/sdk-browser — browser | ||
| - https://iii.dev/docs/reference/engine-protocol — the raw WebSocket protocol, for any other | ||
| language | ||
| Add `.md` to a docs URL to get the raw markdown source. If a fetch fails, use the index at | ||
| https://iii.dev/docs/llms.txt — it lists every doc page. If the docs stay unreachable, | ||
| say so and proceed with extra care: verify every registration with a real call. Do not fetch | ||
| docs for an ordinary call — `engine::functions::info` is the reference for calling | ||
| functions. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,146 @@ | ||
| --- | ||
| name: harness/finishing | ||
| description: >- | ||
| How a fan-out run ends: the turn_complete state key, mechanical (fp::pipe) vs judgment (sub-agent) validators, the deadline cron, wave dispatch, the armed-wake invariant. Mandatory before ending the first turn of any fan-out run. | ||
| --- | ||
|
|
||
| ## Finishing: the turn-complete key | ||
|
|
||
| A fan-out request is NOT answered in one turn. You wire, spawn, and stop; the answer is | ||
| composed later, when the work is actually done. The stop signal is a state key you watch: | ||
|
|
||
| Step 1. Pick a run scope unique to this run — a slug plus random characters, e.g. | ||
| `run-headlines-b4k9`. Every child writes its own key inside this scope. | ||
|
|
||
| Step 2. Register a one-shot notify trigger (NO `function_id`) on | ||
| `{ trigger_type: "state", config: { scope: "<run scope>", key: "turn_complete" } }`. Its | ||
| firing is your wake-up call to finish. NEVER bind this (or the deadline in Step 4) to | ||
| `harness::react` or a `call` — a call's result is discarded in the void; only a plain | ||
| notify injects into YOUR session and wakes you. | ||
|
|
||
| Step 3. Register the validator that decides when the run is done — always with | ||
| `continue_on_error: true` on every edge so failed children still reach it (a validator is | ||
| an outcome checker, not a success handler). For a couple of children, one edge on | ||
| `harness::turn-completed` with `config { parent_session_id: "<this session>" }` is enough; | ||
| for larger fan-ins use a join with `expect` covering every child (a single subscription's | ||
| fires are rate-limited to ~10 per minute — a join accumulates instead of firing per event). | ||
| `parent_session_id` only matches children that CARRY the link: ones you `harness::spawn` | ||
| from this session, or reaction-spawned children whose `metadata.parent_session_id` names | ||
| this session. Children dispatched by reactions or later waves without that link never | ||
| match the filter — for those, key the validator on the run-scope state keys the children | ||
| write (a `state` trigger on the scope) or a join with explicit per-child `expect`, or | ||
| their completions never reach the validator and `turn_complete` is never written. | ||
| Pick the validator KIND by the rule: | ||
|
|
||
| - MECHANICAL rule (keys present, a count reaching N, a value matching) → a `call` reaction | ||
| running `fp::pipe`: zero tokens, milliseconds, no sessions. The `fp::when` guard stops | ||
| the pipe when the condition fails, so `turn_complete` is only written on pass: | ||
|
|
||
| engine::register_trigger { | ||
| trigger_type: "harness::turn-completed", | ||
| function_id: "harness::react", | ||
| config: { parent_session_id: "<this session>" }, | ||
| once: false, | ||
| metadata: { continue_on_error: true, call: { function_id: "fp::pipe", payload: { | ||
| through: [ | ||
| { function: "database::query", payload: { db: "primary", | ||
| sql: "SELECT COUNT(*) AS n FROM <run table>" } }, | ||
| { function: "fp::get", payload: { path: "/rows/0/n" } }, | ||
| { function: "fp::when", payload: { op: ">=", to: <N> } }, | ||
| { function: "state::set", payload: { scope: "<run scope>", | ||
| key: "turn_complete", value: { done: true } }, into: "/value/count" } | ||
| ] } } } | ||
| } | ||
|
|
||
| The same shape counts state keys (start from `state::list`) or checks any readable | ||
| store. A standing (`once: false`) call edge re-checks on every child completion and | ||
| passes exactly once the threshold holds; unregister it when you finish. | ||
|
|
||
| A pipe GATES and moves ONE value — it does NOT compute a new value from several. `fp::*` | ||
| has no arithmetic (no sum/average/reduce), so a pipe that `state::get`s two keys just | ||
| threads the LAST one, never their total. To AGGREGATE (sum children into a subtotal, | ||
| combine results), gate on all inputs being PRESENT — a join over their keys, or a count | ||
| reaching N — and let that gate WAKE YOU with a plain notify (no `function_id`) so the sum | ||
| runs in YOUR OWN next turn, where you already hold your permissions. Do NOT put the | ||
| compute in a react/join DOWNSTREAM TASK unless you also grant it write access: | ||
| a trigger-fired downstream is PARENTLESS — it starts from the read-only baseline and its | ||
| `state::set` is denied, so the subtotal silently never writes and the whole tree above it | ||
| stalls. Either wake yourself and compute (simplest), or pass | ||
| `metadata: { options: { functions: { allow: ["state::get","state::set"] } } }` on the | ||
| downstream. And never gate a MULTI-input condition with a one-shot watcher on ONE key: the | ||
| fire can beat the sibling writes, short-circuit, and strand the check forever — use a join | ||
| (it accumulates every predecessor) instead. | ||
|
|
||
| A pipe READS | ||
| (`state::get`, `state::list`, read-only `database::query`, `fp::*` transforms) and | ||
| writes ONLY `state::set`: steps run with WORKER authority, so `harness::spawn`, every | ||
| DATABASE WRITE (`database::execute`/transactions/prepared statements), and all | ||
| turn-control, session, trigger-control, credential, router, and shell/coder steps are | ||
| REFUSED in a pipe. The gate writes `turn_complete` to STATE; anything else — a db | ||
| marker row, a respawn — happens on the woken turn or a react edge, with agent | ||
| authority. For a mechanical retry, write failures to their OWN key (e.g. | ||
| `<key>-failed`) and register a react edge on it whose REACTION is the respawn | ||
| (`metadata.task` + `model` + `session_id` + `options`) — that spawn runs with agent | ||
| authority under your policy. The threaded value lands at each | ||
| step's `into` (default `/value`) and REPLACES any literal there — on a write step aim | ||
| it at a scratch subfield (`into: "/value/_seen"`) or the literal `value` you meant to | ||
| write is silently overwritten. | ||
| - JUDGMENT rule (quality, coherence, "is this actually an answer") → a sub-agent | ||
| validator. Pin it into its OWN session — `metadata.session_id: | ||
| "validator-<run>-<rand>"`, with `metadata.parent_session_id: "<this session>"` for | ||
| console nesting; NEVER leave its `session_id` out, or the reaction delivers into THIS | ||
| chat. Grant it `metadata.options: { functions: { allow: ["state::get", "state::set"] } }`. | ||
| Its task is self-contained, e.g.: "For each of <scope>/<key1>, <scope>/<key2>: | ||
| `state::get` it, expecting `{ ok, value | error }`. All present with ok true → | ||
| `state::set` <scope>/turn_complete = { done: true, keys: [...], summary: '<one line>' }. | ||
| Anything missing or ok false → `state::set` <scope>/turn_complete = { done: false, | ||
| missing: [...], errors: [...] }. ALWAYS write turn_complete." (`state::list` returns | ||
| values without keys — verify named keys with `state::get`.) | ||
|
|
||
| Step 4. Register a one-shot `cron` notify at this run's deadline — and pass | ||
| `once: true` EXPLICITLY: cron is the one trigger type that defaults to RECURRING, so a | ||
| deadline registered without it fires every period forever (and is not tracked as your | ||
| armed wake). If it fires before `turn_complete` is set, a child never terminated: | ||
| compose what state holds, report the gap, unregister everything, stop. A run must never | ||
| depend on every child behaving. | ||
|
|
||
| Step 5. Spawn the children (up to your per-turn spawn cap), then end the turn: "running — | ||
| I'll report when the results land." If there are MORE items than one turn can spawn, this | ||
| is a WAVE: record the full work-list and a cursor in the run scope, spawn the first wave, | ||
| then arm ONE RELIABLE dispatch-wake to bring you back for the next wave — a single | ||
| recurring short `cron` (e.g. every minute) registered ONCE for the whole run, or a | ||
| self-notify you RE-ARM as the first act of each woken turn. Pick ONE mechanism, ONCE: | ||
| never arm a fresh cron per wave — stacked crons all keep firing for the rest of the run, | ||
| burning a wake-turn each, every minute. Do NOT rely on child-`turn-completed` events to | ||
| pace dispatch: they stop firing once the current wave's children finish (and are | ||
| rate-capped ~10/min), so a run that waits on them deadlocks with items still | ||
| un-dispatched. On each dispatch-wake, spawn the next wave from the cursor; once the | ||
| cursor is exhausted, drop the dispatch-wake (unregister the cron; a self-notify you | ||
| simply stop re-arming) — but ONLY if the `turn_complete` gate notify or the deadline | ||
| still stands to wake you. INVARIANT for every turn you end: while children are in flight | ||
| or the gate has not passed, at least ONE armed wake must survive the turn — gate notify, | ||
| dispatch-wake, or deadline. End a turn with zero armed wakes and the run strands at | ||
| whatever the children last wrote, done work with no one left to collect it. A fan-out that | ||
| stops after one wave silently covers only the first slice and the gate never reaches its | ||
| count. And do NOT begin a large downstream deliverable (a UI, a compiled report) until the | ||
| producer loop is finished and the gate is satisfied: interleaving a secondary build with an | ||
| unfinished fan-out is how a run wanders off and stalls half-done. Finish the loop, THEN | ||
| build the extras. | ||
|
|
||
| When the `[notification]` for `turn_complete` arrives, read the scope and finish: | ||
| `done: true` → compose the final answer FROM STATE (what was actually written, never your | ||
| memory of what you asked for), unregister the deadline and any standing watchers with | ||
| `engine::unregister_trigger`, and stop. `done: false` → report the failure with what | ||
| `missing`/`errors` name — or re-arm a fresh one-shot notify on `turn_complete` FIRST, then | ||
| respawn only the named work. | ||
|
|
||
| # End-of-run checklist | ||
|
|
||
| If this is a fan-out run, check the finish path: validator pinned to its OWN session? | ||
| One-shot state notify armed on the `turn_complete` key itself — the deadline cron is the | ||
| FALLBACK, never the loop's clock (cron-only wakes stall every round until the next | ||
| boundary)? Will `turn_complete` be written on EVERY path — success, failure, and timeout? | ||
|
|
||
| If you end with reactions armed, check each one: can its producer actually produce the watched | ||
| key or event — is the write inside the producer's allowed functions, and will the filtered | ||
| session exist? A reaction armed on something nothing can produce waits forever, silently. | ||
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.