Skip to content

fix: bound manifest jinja collection traversal - #1561

Merged
mldangelo-oai merged 7 commits into
mainfrom
mdangelo/codex/fix-manifest-jinja-traversal-budget
Jun 9, 2026
Merged

mldangelo-oai merged 7 commits into
mainfrom
mdangelo/codex/fix-manifest-jinja-traversal-budget

Conversation

@mldangelo-oai

@mldangelo-oai mldangelo-oai commented Jun 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • replace recursive embedded-Jinja collection with explicit depth and item budgets
  • fail closed on incomplete collection while preserving templates recovered before and after over-depth branches
  • expand shared containers once per semantic mode and chat-template context, while revisiting them at a shallower depth when needed
  • terminate recursive aliases without duplicate-template amplification
  • prioritize direct template fields before untrusted metadata under the global item budget
  • skip collection when the embedded Jinja scanner is excluded
  • redact and bound attacker-controlled traversal and delegated template-location evidence
  • preserve every template when redacted locations collide
  • fall back to safe defaults for invalid and infinite traversal-budget values
  • make the earlier model-name, URL, and weak-hash manifest walks cycle-safe so recursive YAML aliases cannot prevent Jinja analysis
  • normalize non-string YAML keys in those walks and bound path construction before concatenation

False-positive / false-negative audit

  • 100,000 randomized acyclic manifests matched the prior collector's normalized template semantics
  • 1,000 references to one shared mapping remain bounded and do not become falsely inconclusive
  • recursive aliases terminate through the public YAML scan path and preserve sibling malicious templates
  • numeric YAML keys no longer abort structured analysis before embedded-Jinja detection
  • over-depth branches are pruned locally, so later malicious siblings are still analyzed
  • aliases first encountered deeply are revisited at shallower paths
  • direct malicious template fields are analyzed before wide metadata can exhaust the item budget
  • JSON, YAML, TOML, and INI delegation probes detect nested malicious templates
  • credential-bearing manifest keys do not leak through Jinja findings or oversized-template reports
  • benign template locations and collision-distinct templates are preserved
  • a 64 KiB repeated path segment at depth 64 dropped from about 140 MB traced peak allocation to about 1.3 MB; evidence paths remain capped at 240 characters

Validation

  • PROMPTFOO_DISABLE_TELEMETRY=1 UV_CACHE_DIR=/tmp/modelaudit-uv-cache uv run --frozen pytest -q tests/scanners/test_manifest_scanner.py -> 97 passed, 1 skipped (optional jinja2.sandbox unavailable)
  • scoped Ruff check and format checks passed
  • scoped mypy passed
  • git diff --check origin/main passed
  • current main integrated at 43d2fb80731878757b307e3fc018f7097b072fae
  • exact remote tree verified: 7565c505d23bc0e794e30a51098a2f4975637282
  • published head: 6758ae22782d5c73449260648c95b00fa3a431f6

The full local suite was intentionally left to CI per the PR-audit workflow.

@github-actions

github-actions Bot commented Jun 8, 2026 •

Copy link
Copy Markdown
Contributor

Workflow run and artifacts

Performance Benchmarks

Compared 12 shared benchmarks with a regression threshold of 15%.
Status: 0 regressions, 0 improved, 12 stable, 0 new, 0 missing.
Aggregate shared-benchmark median: 1.815s -> 1.823s (+0.4%).

Workload Benchmark Target Size Files Baseline Current Change Status
warm-cache-rescan tests/benchmarks/test_scan_benchmarks.py::test_scan_warm_cached_repository_rescan release-candidate 547.3 KiB 32 98.38ms 89.55ms -9.0% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_raw] nested_raw 78 B 1 491.4us 453.1us -7.8% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_hex] nested_hex 130 B 1 506.3us 492.0us -2.8% stable
direct-malicious-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_direct_malicious_upload malicious_reduce 52 B 1 440.5us 429.4us -2.5% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_base64] nested_base64 98 B 1 482.7us 471.6us -2.3% stable
mixed-model-repository tests/benchmarks/test_scan_benchmarks.py::test_scan_release_candidate_repository release-candidate 547.3 KiB 32 594.23ms 604.25ms +1.7% stable
single-checkpoint-preflight tests/benchmarks/test_scan_benchmarks.py::test_scan_single_checkpoint_before_load single_checkpoint.pkl 183.0 KiB 1 113.48ms 114.80ms +1.2% stable
duplicate-heavy-registry tests/benchmarks/test_scan_benchmarks.py::test_scan_duplicate_registry_snapshot registry-snapshot 915.2 KiB 13 599.94ms 604.65ms +0.8% stable
chunked-upload-stream tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_chunked_upload_stream chunked_stream 278.2 KiB 1 113.03ms 113.14ms +0.1% stable
clean-training-checkpoint tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_clean_training_checkpoint safe_large 278.2 KiB 1 110.44ms 110.46ms +0.0% stable
padded-multi-stream-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_padded_multi_stream_upload multi_stream_padded 4.1 KiB 1 545.2us 545.1us -0.0% stable
suspicious-pickle-intake tests/benchmarks/test_scan_benchmarks.py::test_scan_suspicious_pickle_intake suspicious-intake 183.8 KiB 4 183.52ms 183.55ms +0.0% stable

@mldangelo-oai
mldangelo-oai marked this pull request as ready for review June 8, 2026 19:36

Copy link
Copy Markdown
Contributor Author

Critical review and focused QA complete.

Fixed a false positive where manifests could fail the new embedded-Jinja traversal budget even when jinja2_template was explicitly excluded by scanner selection. The policy gate now runs before collection, with regression coverage for a benign nested manifest under a tiny depth budget.

Validation:

  • PROMPTFOO_DISABLE_TELEMETRY=1 uv run pytest tests/scanners/test_manifest_scanner.py -q (81 passed)
  • uv run ruff check modelaudit/scanners/manifest_scanner.py tests/scanners/test_manifest_scanner.py
  • uv run ruff format --check modelaudit/scanners/manifest_scanner.py tests/scanners/test_manifest_scanner.py
  • uv run mypy modelaudit/scanners/manifest_scanner.py tests/scanners/test_manifest_scanner.py

Adversarial checks covered exact item/depth boundaries, malicious templates collected before exhaustion, disabled scanner selection, and cyclic input to the collector. A pre-existing recursive-YAML alias can still fail closed earlier in other manifest walkers with RecursionError; it is not a silent bypass or introduced by this PR.

Published fix: 17a0c584e1317a00b66275a1facfccf8048e40b1.

Replace recursive embedded-template collection with bounded traversal, preserve recovered findings on incomplete scans, redact budget evidence, and avoid treating plain nested template metadata as executable Jinja.
@mldangelo-oai
mldangelo-oai force-pushed the mdangelo/codex/fix-manifest-jinja-traversal-budget branch from 17a0c58 to cfe4a53 Compare June 8, 2026 20:07

@mldangelo-oai mldangelo-oai left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found two traversal correctness issues that should be fixed before merging:

  1. Repeated YAML aliases are expanded once per reference. A list containing 1,000 references to the same small mapping consumes the item budget and records hundreds of duplicate templates, creating a false-inconclusive result.
  2. A recursive alias encountered before a sibling chat_template repeatedly descends until the depth limit. The current depth-limit break then stops the entire traversal, so the sibling malicious template is never analyzed. This is a false-negative path, even though the scan is marked inconclusive.

I prepared local commit de309af7 that tracks expanded container identities per traversal mode and skips only the over-depth branch instead of aborting all remaining siblings. It also adds regressions for shared aliases, recursive aliases, and a malicious template ordered after an over-depth branch.

Focused validation:

  • tests/scanners/test_manifest_scanner.py: 85 passed
  • traversal subset: 9 passed
  • Ruff check/format: clean
  • scoped mypy: clean
  • direct probes: shared alias case finishes in 1,004 visits with one template; recursive alias finishes in 3 visits and preserves the sibling template

I could not push the commit because the available shell GitHub token is invalid, and these files are too large for the connector's whole-file update endpoint. The finding is therefore still present on remote head cfe4a53f.

Copy link
Copy Markdown
Contributor Author

Addressed both traversal blockers from the review in eabc0e06c8c588e9995012173cb5f506b2f04838 (published head 608eb38e7fad8021ecb4470b657da22b3c9a8c12).

  • container identities are expanded once per traversal mode, so repeated/recursive YAML aliases cannot amplify templates or consume the budget through repeated subtree walks
  • depth overflow now marks analysis incomplete and prunes only that branch; remaining siblings continue, including malicious templates ordered afterward
  • mode-aware identity tracking still revisits a shared container when it moves from ordinary metadata to a template field

Focused QA: complete manifest scanner module 85 passed, 1 skipped; blocker subset 4 passed; Ruff, format, mypy, and diff checks clean.

Copy link
Copy Markdown
Contributor Author

Addressed the existing traversal review and completed an additional false-positive/false-negative pass on head 84d9ff678c6eff11acfefe3a96415005ff3e4a21.

Fixes now published:

  • shared YAML aliases are expanded once per semantic context instead of once per reference
  • recursive and over-depth branches no longer abort later siblings
  • aliases first seen deeply are revisited when a shallower path can recover descendants
  • generic/template and chat-template alias contexts are both retained, closing a benign-macro false positive
  • direct template fields are prioritized before wide metadata under the global item cap
  • infinite numeric budget values safely fall back to defaults

Scoped QA: 91 manifest-scanner tests, 8 focused adversarial regressions, Ruff check/format, mypy, direct probes, and exact remote tree/blob verification. No inline review threads remain. Full-suite validation is left to CI.

Copy link
Copy Markdown
Contributor Author

Integrated current main and republished the exact reviewed tree at f96b111dcb3c6c281a2cd894160a8e9c4a6abeac (tree 72fa4b96c9f0809739f387dda8a231f5777bc443). The complete associated manifest scanner module still passes: 91 tests. Scoped Ruff, format, mypy, and diff checks are clean; no inline review threads remain. This refreshes the stale merge base and hands the full-suite gate back to CI.

@mldangelo-oai
mldangelo-oai marked this pull request as draft June 9, 2026 10:24
Merge current main and prevent attacker-controlled manifest keys from leaking through delegated Jinja findings or oversized-template reports while preserving chat-template context and collisions.
@mldangelo-oai
mldangelo-oai marked this pull request as ready for review June 9, 2026 14:29
@mldangelo-oai
mldangelo-oai merged commit 23b6ca2 into main Jun 9, 2026
20 of 24 checks passed
@mldangelo-oai
mldangelo-oai deleted the mdangelo/codex/fix-manifest-jinja-traversal-budget branch June 9, 2026 15:06
@github-actions github-actions Bot mentioned this pull request Jun 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant