Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
85c5e45
feat: extract single-tool CodeAct and Todo delegation prerequisites
furgalep Sep 14, 2026
34d2ae4
feat(bench): backport current benchmark agents and runner
furgalep Sep 14, 2026
345c645
test: pin ACP subprocesses to the checkout under test
furgalep Sep 14, 2026
8f3186a
fix: address bounded context and prefill trace review findings
furgalep Sep 14, 2026
6da0bf2
docs: explain shared state and delegated context boundaries
furgalep Sep 14, 2026
726ebbe
refactor(bench): separate RLM agent and trim unused backport scope
furgalep Sep 14, 2026
30831fb
fix(bench): drain background summaries before releasing resources
furgalep Sep 14, 2026
6208438
fix(bench): align behavior metrics with real trajectory exports
furgalep Sep 14, 2026
b4ebfd6
fix(bench): address metrics, capability, and cleanup review findings
furgalep Sep 16, 2026
c4f1d23
Address benchmark review; promote CodeActV2 and remove CodeActLite
furgalep Sep 16, 2026
f44e38d
Make CurrentCall lifecycle fields ordinary mutable attributes
furgalep Sep 16, 2026
70c18b7
Address remaining benchmark review summary and viewer follow-ups
furgalep Sep 16, 2026
adb0504
fix: close bench review gaps and harden viewer requests
furgalep Sep 16, 2026
18fd5bd
fix: retain authenticated HTTP viewer compatibility with warning
furgalep Sep 16, 2026
b3f033f
Clarify agent context and shell contracts; fix prompt binding
furgalep Sep 16, 2026
f01d08d
Generalize benchmark verification and demonstrate todo activation
furgalep Sep 16, 2026
3438752
Omit automatic cell-state inventory from benchmark prompts
furgalep Sep 16, 2026
ca87df8
Restore context usage status for benchmark agents
furgalep Sep 16, 2026
66c1fd1
Shorten context status and retain compaction guidance
furgalep Sep 16, 2026
b337133
Accept integer collapse tags with model-visible guidance
furgalep Sep 16, 2026
c19998e
Accept valid integer collapse tags silently
furgalep Sep 16, 2026
96cc76f
Clarify context block editing guidance
furgalep Sep 16, 2026
8706f60
Fix benchmark review regressions and use standard delegation arguments
furgalep Sep 16, 2026
8dcaf13
fix: align context reserves with model requests and support streaming…
furgalep Sep 16, 2026
94792d6
Reject reply reserves that exhaust the known context window
furgalep Sep 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
70 changes: 70 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,76 @@ to follow semantic versioning.

## [Unreleased]

- `ShellTools.run_stream` and the coding activity wrapper now accept
`command, *, stdin=None, timeout=30.0`, matching `run`. Streaming uses the
same stdin handling; pass an existing positional timeout as `timeout=...`.
- Context status, percentages, automatic summarization and overflow recovery now
reserve the selected UnifiedLLM client's effective reply cap, including
reasoning levels and per-call overrides. Unknown caps use a labelled planning
reserve. Automatic summaries trigger at 80% of the usable input window;
explicit summary thresholds remain fixed across model switches. Responses
requests translate reply-cap aliases to `max_output_tokens`, and cap overrides
replace inherited aliases rather than sending conflicting limits.
Context management rejects a configured reply cap at or above the known
context window instead of repeatedly summarizing against a one-token budget.
- Restore legacy Todo notes and statuses through the stored-session deserializer,
and retain completed worker results when a delegated Todo disappears.
Cleanup handles child-task re-entry and continues after a callback is cancelled,
without cancelling unrelated callers.
- Keep CodeAct call correlation IDs separate from task display tags, report live
input types after reassignment, and correct V2 tool/delegation hints. Benchmark
working-directory context is untraced; failed trajectory exports no longer
reuse a previous task's metrics. No-ID trace attribution requires matching code
before selecting a later LLM turn.
- `self.events.collapse()` accepts integer endpoints that identify existing events,
including mixed string/integer ranges over prior summaries, without warnings.
Invalid numeric endpoints leave history unchanged.
- Rename the benchmark `TaskResult.command_to_verify` field to `how_to_verify`
("How to Verify"): concrete verification steps and expected results, not
necessarily a shell command. Result JSON and runner answers use the new field.
- Reject ambiguous `ShellTools.replace(match, old, new)` calls before file access,
with guidance for full-region versus path-based substring replacement.
- Add `CodeActV2`, the single-`python_cell` strategy with in-cell `return_result`.
The benchmark agents use it; `CodeActStrategy` remains the default.
Its cacheable Python-cell context includes the execution namespace's typed
stub without a second execution-context block. Names and runtime helpers use
one Python-style block; internal delegation errors are not advertised there.
Benchmark agents omit the automatic `python_cell_state` inventory block.
- Trace explorer viewer requests now send configured viewer authentication and
honor proxy environment settings, including `NO_PROXY` for direct access.
Authenticated HTTP requests warn that bearer tokens are unencrypted; existing
HTTP viewer/exporter setups remain supported. Use HTTPS or a trusted local
connection/tunnel. The warning includes neither the token nor the URL.
- `CurrentCall` is a mutable invocation record; strategies bind its task tag and
live execution namespace with ordinary public-field assignment during setup.
- Breaking: remove `CodeActLiteStrategy` and its experimental exports. Use
`CodeActStrategy` for the existing two-tool contract or `CodeActV2` for the
single-tool contract. The evaluation CLI option is now `codeact_v2`.
- Preserve inline completion values in replay and archived events in benchmark
trajectories; make Todo updates/restores atomic and delegation merge failures
recoverable. Behavior reports use schema version 2; regenerate older reports
from their trajectories before comparing results.
- Benchmark agents release resources through `aclose()` as well as `close()`.
Supplied delegation context uses ordinary method-argument formatting, without
a custom renderer or redaction policy. Todo metadata
and comment read-back methods are now included in model-facing documentation.
Cancellation during shutdown is propagated only after background cleanup drains.
- `CodeActStrategy` remains the default strategy, but its model-facing behavior
changes: revised delegation guidance, validated inline completion values in
PythonOutput (None on validation failure), no replay of synthetic inline-return
tool pairs, and explicit error/retry feedback for non-object tool arguments.
- Todo snapshot upgrades preserve legacy tasks, but downgrading to the previous
implementation silently loses descriptions, active-task selection and comment
IDs. Back up sessions before downgrading. `TodoVars` is now an alias for
`PersistentVars`; helper-name keys must use explicit `get`/`set` access, and
private/helper attribute writes are rejected. `InteractiveAgent.v` retains its
separate `AgentVars` implementation.
- Delegation merge conflicts raise `DelegationMergeError` carrying the completed
result and worker state. Benchmark agents no longer pre-seed a planning Todo;
they expose tools through `python_cell_tools`, retain concise `context_usage`
status and compaction guidance, and recreate the shell for each evaluation's
working directory.

- Responses clients now honor the cached renderer's stable-prefix boundary by default,
without a cache setting in the model registry. Requests without a usable boundary
retain provider-default caching; `cache_breakpoint=None` opts out of NOOA markers.
Expand Down
2 changes: 1 addition & 1 deletion packages/nooa-acp/src/nooa_acp/event_bridge.py
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,7 @@ def _on_agent_message(self, event: EventBase) -> None:
def _on_tool_call(self, event: EventBase) -> None:
if (
not isinstance(event, ToolCallEvent)
or event.name != "execute_python"
or event.name not in {"execute_python", "python_cell"}
or event.metadata.get("prefill") is True
# codeact manufactures an execute_python call to carry a prose-only
# reply. Nothing ran, so showing it as a Python card would present
Expand Down
37 changes: 37 additions & 0 deletions packages/nooa-acp/tests/conftest.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# SPDX-FileCopyrightText: Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Keep protocol subprocesses on the checkout and configuration under test."""

import os
from pathlib import Path

import pytest


@pytest.fixture(autouse=True)
def _protocol_subprocess_environment(monkeypatch):
"""Preserve test sources/config through ACP's sanitized subprocess environment."""
import acp.transports

original = acp.transports.default_environment
root = Path(__file__).resolve().parents[3]
sources = [root / "src"] + [
root / "packages" / package / "src"
for package in ("nooa-cli", "nooa-acp", "nooa-memory", "nooa-bench")
]

def environment():
values = original()
values["PYTHONPATH"] = os.pathsep.join(str(path) for path in sources)
for name in (
"NEMO_OO_USER_DIR",
"NEMO_OO_PROJECT_DIR",
"NEMO_OO_SETTINGS",
"NOOA_SESSIONS_DIR",
"NOOA_ACP_MCP_TRACE",
):
if name in os.environ:
values[name] = os.environ[name]
return values

monkeypatch.setattr(acp.transports, "default_environment", environment)
5 changes: 3 additions & 2 deletions packages/nooa-acp/tests/test_event_bridge.py
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,8 @@ def _content_text(content: ContentToolCallContent) -> str:
return block.text


async def test_bridge_preserves_message_tool_and_usage_order(tmp_path):
@pytest.mark.parametrize("tool_name", ["execute_python", "python_cell"])
async def test_bridge_preserves_message_tool_and_usage_order(tmp_path, tool_name):
agent = CodingAgent(llm=FakeLLMClient(), cwd=tmp_path)
client = _RecordingClient()
bridge = ACPEventBridge(agent, client, "session-1") # type: ignore[arg-type]
Expand All @@ -69,7 +70,7 @@ async def test_bridge_preserves_message_tool_and_usage_order(tmp_path):
agent.event_manager.add(
ToolCallEvent(
tool_call_id="call-1",
name="execute_python",
name=tool_name,
arguments={"code": "print('hello')"},
)
)
Expand Down
25 changes: 25 additions & 0 deletions packages/nooa-acp/tests/test_protocol.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,31 @@
_HANG_TIMEOUT = 30


def test_protocol_subprocess_imports_this_checkout(tmp_path):
import json
import subprocess

from acp.transports import default_environment

root = Path(__file__).resolve().parents[3]
result = subprocess.run(
[
sys.executable,
"-c",
"import json, nooa, nooa_acp; print(json.dumps([nooa.__file__, nooa_acp.__file__]))",
],
cwd=tmp_path,
env=default_environment(),
capture_output=True,
text=True,
check=True,
)
assert [Path(path).resolve() for path in json.loads(result.stdout)] == [
root / "src/nooa/__init__.py",
root / "packages/nooa-acp/src/nooa_acp/__init__.py",
]


class _RecordingClient:
def __init__(self) -> None:
self.updates: list[tuple[str, object]] = []
Expand Down
66 changes: 63 additions & 3 deletions packages/nooa-bench/README.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# nooa-bench

Benchmark agent (`BenchAgent`) and Harbor runner for
[NOOA](https://github.com/NVIDIA-NeMo/labs-OO-Agents). Reproduces the SWE-bench
and Terminal-Bench results from the NOOA tech report.
Coding benchmark agents and Harbor runner for
[NOOA](https://github.com/NVIDIA-NeMo/labs-OO-Agents), supporting SWE-bench
and Terminal-Bench tasks.

```bash
uv add nooa-bench
Expand All @@ -12,4 +12,64 @@ nemo-harbor --help
See the [main repository](https://github.com/NVIDIA-NeMo/labs-OO-Agents) for
documentation.

Two agent variants are available through `nemo-harbor --agent-type`:

- `bench` — `BenchAgent` in `nooa_bench.bench_agent`: compact CodeAct baseline
with automatic summarization and optional delegation.
- `rlm` — `RLMBenchAgent` in `nooa_bench.rlm_bench_agent`: the same capabilities
with instructions emphasizing delegation for bounded work.

Both use `CodeActV2` with the single `python_cell` tool and return a structured `TaskResult`.
Its `how_to_verify` field describes concrete checks and expected results; commands
are optional. The `evidence` field records results the agent actually observed.
Both delegate through an awaited call returning a `TaskResult`; neither exposes
the interactive coding agent's background `spawn()` / job-handle API.
The strategy allows ten retries, uses a 1,800-second cell timeout, and has no
fixed iteration cap; configure the enclosing benchmark's time/token budget.
Awaited delegation runs inside that same parent cell deadline. A timeout cancels
the worker and merges no partial Todo state; the parent receives a cell timeout
error and may try again. Cell timeouts do not consume the strategy's retry counter,
so the enclosing harness budget is the overall limit on repeated delegations.
Workers use the same agent type, model client and working directory, with their
own execution context and shell. Delegation defaults to a maximum depth of four.
Passing a Todo gives the worker an independent task copy; successful worker
updates are merged after cleanup. Conflicts or worker-only dependencies raise
`DelegationMergeError`, retaining the completed `result` and full `worker_state`
for explicit reconciliation without rerunning the worker. Failed execution or
cleanup does not merge partial state. Task-local state stays on Todos, and automatic
summarization handles context maintenance. Method-writing tools are available in
both variants.

Context usage is the last provider-reported input count divided by the model
window minus UnifiedLLM's effective reply cap (including reasoning-level and
per-call settings). For example, 32,000 input tokens with a 128,000-token window
and 64,000-token reply cap is 50%. Automatic summarization triggers at 80% of
that usable input window. Explicit summarization thresholds stay fixed; an
unknown reply cap uses a labelled planning reserve, not a claimed model limit.

The runner writes `result.json`, `trajectory.json` and aggregate `behavior.json`
under `/logs/agent`, and the verification instructions to `/app/answer.txt`. Behavior
metrics count both Python tool names and exclude framework prefill. Set
`NOOA_INTERFACE_CHANGE_ID` to label a comparison; the default is `baseline`.
Set `NOOA_TASK_ID` to identify the task when logs share the `/logs/agent` path.
Exports include archived events after summarization, with the actual event IDs.
The original task inputs also remain in bounded, non-summarizable prompt context.
Metrics cover the controller's history; delegated workers keep separate histories
and their cells are not included. Recovery and retry metrics are omitted until
framework events carry explicit attempt linkage. Schema version 2 removes
unsupported error-code guesses from stdout. Code metrics count syntactic call
sites, not actual runtime loop iterations; fan-out recognizes direct and starred
`asyncio.gather` arguments, including comprehensions and same-cell list aliases.
Simple same-cell aliases of Todo, shell and repo objects are recognized; this is
not general cross-cell dataflow analysis. Reports with another schema, content
policy or unknown metrics are rejected; regenerate them from `trajectory.json`.
Supplied delegation context is an ordinary worker-method argument, displayed by
NOOA's standard parameter formatting. There is no delegation-specific renderer
or redaction policy; pass only the data the worker needs.
Failure to generate the behavior report does not fail an otherwise completed
task. Agents close their shells; the runner closes the shared model client.

These are the current agent prompts and strategy. Reproducing a historical tech
report run requires its original code revision and configuration.

Apache-2.0 licensed.
2 changes: 1 addition & 1 deletion packages/nooa-bench/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
name = "nooa-bench"
# Version is derived from git tags at build time by uv-dynamic-versioning.
dynamic = ["version"]
description = "Benchmark agent (BenchAgent) and Harbor runner for the NOOA framework — reproduces the tech report's SWE-bench and Terminal-Bench results"
description = "Coding benchmark agents, Harbor runner and trajectory analysis for the NOOA framework"
license = {text = "Apache-2.0"}
readme = "README.md"
# Matches nooa and nooa-cli (both >=3.12,<3.14), which this depends on.
Expand Down
9 changes: 4 additions & 5 deletions packages/nooa-bench/src/nooa_bench/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,18 +2,17 @@
# SPDX-License-Identifier: Apache-2.0
"""Benchmark agent and Harbor runner for the NOOA framework.

Minimal reproducibility package for the NOOA tech report: the benchmark-agnostic
``BenchAgent`` (SWE-bench Verified, Terminal-Bench 2.0), the ``nemo-harbor``
container runner, and the trace analyzer used to extract the per-task token
statistics reported in the paper.
Benchmark-agnostic coding agents for SWE-bench and Terminal-Bench, the
``nemo-harbor`` container runner, and trajectory/token analysis utilities.
"""

from __future__ import annotations

# Maps --agent-type CLI values to dotted import paths: "module:ClassName"
AGENT_CLASSES: dict[str, str] = {
# Unified SWE-bench + Terminal-Bench agent (the tech report's BenchAgent)
# Unified SWE-bench + Terminal-Bench baseline.
"bench": "nooa_bench.bench_agent:BenchAgent",
"rlm": "nooa_bench.rlm_bench_agent:RLMBenchAgent",
}

__all__ = ["AGENT_CLASSES"]
Loading
Loading