Skip to content

fix(memory): respect providers and quota cooldowns in reflection - #671

Open
HaningZS wants to merge 1 commit into
HarnessMD:mainfrom
HaningZS:fix/provider-aware-reflection-596
Open

HaningZS wants to merge 1 commit into
HarnessMD:mainfrom
HaningZS:fix/provider-aware-reflection-596

Conversation

@HaningZS

@HaningZS HaningZS commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What & why

Closes #596. Reflection now resolves the worker's provider/model from registry + roster (with explicit reflectProvider / reflectModel / reflectCommand overrides), rather than always running the global Claude command with Haiku/Claude flags.

  • Claude retains the existing hidden interactive/subscription runner and cheap default; an explicit reflect model overrides it. All tools are denied for this transformation. Quota banners without a transcript are preserved through an opt-in diagnostic path.
  • OpenCode uses its own run --format json adapter, stdin prompt, deny-all-tools primary agent, --pure, disabled sharing and explicit title. Only text events enter the summary; error events take precedence. An explicit title prevents the extra automatic title-model request discovered by the real CLI experiment.
  • Selected OpenCode local endpoint/model registration and only the selected cloud backend's broker key are forwarded. No new dependencies or automatic CLI installation.
  • Limit responses are rate-limited with provider/model/detail/retryAt, not generic parse failures. Calls sharing the same home/provider/binary/model respect a cooldown, including manual reflect; unknown resets back off exponentially. Original memory, pinned facts, newest sections, lossless backups and verification/atomic-swap gates are preserved.
  • Unsupported providers return an explicit unsupported-provider without silently launching Claude. An explicit supported reflector override is available for those floors.

Before

Same nine MemoryReflector regression scenarios with upstream f44b986a: 2 pass / 7 fail. Actual node:test output rendered as an image; CLI calls in this unit suite are mocked, memory files and production reflector are real.

Global Claude and quota failures

After

9 pass / 0 fail for the same scenarios. Additional runner/main-wiring/hidden-PTY/parser coverage is listed below.

Provider-aware reflector regression tests

Validation

macOS / Node 24.13.0:

  • 33 focused reflection tests pass: worker/override/legacy selection, quota reset and exponential/shared backoff, malformed config, memory preservation, framed parsing, CLI argument/env transport, stderr errors, timeout/output bound, unsupported providers, paths with spaces, Windows JS/native npm shim decoding, main registry/roster/local-endpoint/key wiring, and quota banners without a Claude transcript.
  • Full suite with independent catalog-contract PR test(models): validate hot and offline catalogs independently #669: 992 tests / 991 passed / 1 platform skip / 0 failed. A temporary test loader substitutes only test(models): validate hot and offline catalogs independently #669's catalog test, not production code or submitted files.
  • Integration checkout combining all three contributions (test(models): validate hot and offline catalogs independently #669, fix(worktree): persist isolation origins for agent restore #670 and this PR): 1004 tests / 1003 passed / 1 platform skip / 0 failed, plus typecheck/build passed. Production source combines the three actual commits; the full suite runs without a test-file substitution. There were no merge conflicts. The only subsequent amend preserves the architecture document's original final blank line (no source/test changes).
  • Independent full-suite run retains only the pre-existing catalog equality failure (CI never runs the test suite; model-catalog-remote is failing on main #652/test(models): validate hot and offline catalogs independently #669). This PR does not bundle that unrelated fix.
  • npm run typecheck, npm run build, git diff --check: passed.
  • Real official OpenCode 1.18.34 CLI, reproducible via OPENCODE_BIN=/path/to/opencode node test/reflect-opencode.manual.cjs: one loopback request with the selected fixture model, stdin → NDJSON → framed summary, zero advertised tools. A second deterministic quota-banner scenario runs through the real CLI and production MemoryReflector: original memory remains byte-identical, reason/deadline are recorded, and the next reflect call issues zero requests.

The CLI experiment uses an isolated deterministic OpenAI-compatible fixture, not real inference quality or a real exhausted account. No paid model/backend or account credentials are used. Server SSH timed out on two read-only attempts, so no Linux/A800 runtime or GPU claim is made.

Scope and self-review

  • No Settings UI rewrite: explicit overrides live in config.json and are documented in docs/memory-reflection.md, linked from the contributor architecture map.
  • Validated adapters are Claude and OpenCode only. Other CLIs fail visibly and preserve memory; their safe noninteractive adapters are not guessed.
  • OpenCode must support --pure; older CLIs fail safely with diagnostic output. Current official help and CLI behavior were checked. CLI reference, permissions, config precedence, Claude tool denial.
  • Cooldowns are in-process; app restart may make one new attempt. Unknown timezone/stale reset banners are not guessed. There is no new fleet dashboard or model-quality claim.
  • Cross-platform process/path helpers and Windows shim tests avoid shell interpolation; prompts travel on stdin. Real Windows GUI/subscription and Linux CLI smoke tests remain unverified. Timeout/output-overflow process-tree cleanup is tested at the process boundary.
  • Living-doc pre/post checks: 0 errors / 0 warnings; only the existing missing-profile/entrypoint info findings. No unrelated memory-system migration or doc rewrite.

Why:
- All-OpenCode floors still launched a hidden Claude session with Claude flags/model (HarnessMD#596).

What:
- Resolve per-agent or explicitly configured reflection provider, command and model.
- Add a bounded no-tools OpenCode adapter and retain the Claude subscription runner.
- Classify quota banners and share reset/backoff deadlines without rewriting memory.
- Document overrides, adapter limits and reproducible official-CLI validation.

Risk:
- Safe adapters are Claude and OpenCode; other providers require an explicit supported override.
- Cooldowns are in-process; old OpenCode without --pure fails safely.

Tests:
- 33 focused reflection tests pass; real OpenCode 1.18.34 loopback fixture validates no tools and quota cooldown.
- Full suite with HarnessMD#669: 991 passed, 1 skipped, 0 failed; independent suite retains known HarnessMD#652 error.
- npm run typecheck, npm run build and git diff --check passed on macOS / Node 24.13.0.

Live Docs:
- docs/memory-reflection.md, linked from docs/ARCHITECTURE.md.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant