Repository navigation
Instrument the facility goal (and fix the matcher that made it read zero) - #28
Merged
Merged
Conversation
The open question is no longer accuracy — it is whether anything depends on this memory. That was never measured, and the ad-hoc measurement was wrong: MCP tools appear in transcripts as mcp__engram__engram_recall, so grepping engram_recall reported zero reads and confirmed the premise it was meant to test. The real history is 171 calls across 19 sessions. eval/facility.py reports the L0-L4 ladder from real signal only: coverage and ingest lag by source, a structural tool-call matcher with an mcp__ prefix strip (reverting it drops L2 to 0, and a test pins that), a control count so a broken matcher cannot masquerade as a finding, supersession rate, and recall reach labelled as the ceiling it is. L3 usefulness prints NOT MEASURED rather than inventing a proxy, and L4 says it needs a human. Counts and rates only — no memory content, no paths. First reading: L1 fails on lag, not coverage — 77.4% closed but p50 3.2 days against a 60s goal, and the report states the watcher's 15-45m floor cannot reach it while a session-end hook can. It also detects that the write target and the read target are two different memories. engram-watch --install-hook writes that SessionEnd hook, with dry-run, backup and an uninstall that leaves every other hook byte-identical.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
eval/facility.pyreports the L0-L4 ladder from real signal, plusengram-watch --install-hookfor the SessionEnd hook that unblocks L1.The matcher correction matters on its own: MCP tools are named
mcp__engram__engram_recallin transcripts, so every ad-hocengram_recallgrep reported zero reads. Real history is 171 calls across 19 sessions. A test pins the prefix strip — reverting it drops L2 to 0.First reading: L1 fails on lag (p50 3.2 days vs a 60s goal), not coverage (77.4%). The report also flags that the write target and the read target are currently two different memories.