Skip to content

fix(jinja): require executable SSTI context - #1647

Merged
mldangelo-oai merged 3 commits into
mainfrom
mdangelo/codex/hf-fp-t25-jinja-ssti-executable-context-20260610
Jun 11, 2026
Merged

mldangelo-oai merged 3 commits into
mainfrom
mdangelo/codex/hf-fp-t25-jinja-ssti-executable-context-20260610

Conversation

@mldangelo-oai

Copy link
Copy Markdown
Contributor

Summary

  • Scope static Jinja SSTI regex matching to executable {{ ... }} and {% ... %} spans instead of full extracted template strings.
  • Ignore literal template data, comments, and raw blocks for SSTI regexes while preserving active expressions, statements, obfuscated filters, and malformed active spans.
  • Add regressions for the Cohere-style three-entry chat template prose false positive plus active and malformed malicious controls.

Root Cause

The Jinja scanner extracted chat_template values correctly, but then applied SSTI regexes across the entire template string. Literal system-preamble prose in CohereLabs/North-Mini-Code-1.0 includes phrases like user's requests.; the requests\. critical pattern matched that literal prose as if it were an executable request call.

Security Tradeoff

This narrows static SSTI indicators to structurally active Jinja spans. Benign prose outside delimiters, Jinja comments, and raw blocks no longer produce security findings. Active payloads such as {{ requests.get(...) }}, {% set x = requests.post(...) %}, dunder traversal, attr(...) obfuscation, command execution, and unterminated active expressions remain detected. When the Jinja lexer is unavailable, a delimiter fallback still scans executable spans conservatively, including unterminated active tags.

Validation

  • PROMPTFOO_DISABLE_TELEMETRY=1 uv run pytest tests/scanners/test_jinja2_template_scanner.py -q -> 419 passed, 1 skipped.
  • PROMPTFOO_DISABLE_TELEMETRY=1 uv run pytest tests/scanners/test_gguf_scanner.py tests/test_simple_jinja2.py -q -> 58 passed, 3 skipped.
  • uv run ruff format --check modelaudit/ packages/modelaudit-picklescan/src packages/modelaudit-picklescan/tests tests/ -> 419 files already formatted.
  • uv run ruff check modelaudit/ packages/modelaudit-picklescan/src packages/modelaudit-picklescan/tests tests/ -> all checks passed.
  • uv run mypy modelaudit/ packages/modelaudit-picklescan/src packages/modelaudit-picklescan/tests tests/ -> success, 474 source files.
  • PROMPTFOO_DISABLE_TELEMETRY=1 uv run pytest -n auto -m "not slow and not integration" --maxfail=1 -> 18546 passed, 793 skipped, 40 warnings.
  • git diff --check -> passed.

Pinned Hugging Face QA

Before fix on exact pinned revision:

PROMPTFOO_DISABLE_TELEMETRY=1 HF_HUB_DISABLE_PROGRESS_BARS=1 uv run modelaudit scan https://huggingface.co/CohereLabs/North-Mini-Code-1.0/resolve/effaeda477c041c107d5a3d8c599cb5d6c5878ef/tokenizer_config.json --scanners jinja2_template --no-cache --max-size 5MB --timeout 180 --format json --output /tmp/modelaudit-t25-before-allci.json --quiet

Outcome: exit 1; three CRITICAL requests\. components in chat_template[0].template, chat_template[1].template, and chat_template[2].template.

After fix on the same pinned revision:

PROMPTFOO_DISABLE_TELEMETRY=1 HF_HUB_DISABLE_PROGRESS_BARS=1 uv run modelaudit scan https://huggingface.co/CohereLabs/North-Mini-Code-1.0/resolve/effaeda477c041c107d5a3d8c599cb5d6c5878ef/tokenizer_config.json --scanners jinja2_template --no-cache --max-size 5MB --timeout 180 --format json --output /tmp/modelaudit-t25-after-final.json --quiet

Outcome: exit 0 clean; JSON summary success=true, files_scanned=1, issues=0, failed_checks=0.

Scope static Jinja SSTI regexes to executable template spans so literal chat-template prose does not trigger critical request-call findings. Preserve active expressions, statements, obfuscated traversal, and malformed active payload detection with focused regressions.
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Workflow run and artifacts

Performance Benchmarks

Compared 12 shared benchmarks with a regression threshold of 15%.
Status: 0 regressions, 0 improved, 12 stable, 0 new, 0 missing.
Aggregate shared-benchmark median: 1.437s -> 1.421s (-1.1%).

Workload Benchmark Target Size Files Baseline Current Change Status
single-checkpoint-preflight tests/benchmarks/test_scan_benchmarks.py::test_scan_single_checkpoint_before_load single_checkpoint.pkl 183.0 KiB 1 74.48ms 70.78ms -5.0% stable
duplicate-heavy-registry tests/benchmarks/test_scan_benchmarks.py::test_scan_duplicate_registry_snapshot registry-snapshot 915.2 KiB 13 406.34ms 389.13ms -4.2% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_hex] nested_hex 130 B 1 627.5us 604.5us -3.7% stable
direct-malicious-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_direct_malicious_upload malicious_reduce 52 B 1 527.3us 511.0us -3.1% stable
padded-multi-stream-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_padded_multi_stream_upload multi_stream_padded 4.1 KiB 1 649.0us 632.7us -2.5% stable
warm-cache-rescan tests/benchmarks/test_scan_benchmarks.py::test_scan_warm_cached_repository_rescan release-candidate 547.3 KiB 32 103.55ms 105.83ms +2.2% stable
suspicious-pickle-intake tests/benchmarks/test_scan_benchmarks.py::test_scan_suspicious_pickle_intake suspicious-intake 183.8 KiB 4 144.49ms 141.44ms -2.1% stable
chunked-upload-stream tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_chunked_upload_stream chunked_stream 278.2 KiB 1 110.55ms 111.98ms +1.3% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_raw] nested_raw 78 B 1 578.4us 583.6us +0.9% stable
clean-training-checkpoint tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_clean_training_checkpoint safe_large 278.2 KiB 1 107.89ms 108.77ms +0.8% stable
mixed-model-repository tests/benchmarks/test_scan_benchmarks.py::test_scan_release_candidate_repository release-candidate 547.3 KiB 32 486.99ms 490.46ms +0.7% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_base64] nested_base64 98 B 1 580.3us 576.9us -0.6% stable

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🎉

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@mldangelo-oai
mldangelo-oai requested a review from mldangelo June 11, 2026 01:41
@mldangelo-oai
mldangelo-oai enabled auto-merge (squash) June 11, 2026 01:41
ianw-oai
ianw-oai previously approved these changes Jun 11, 2026

@ianw-oai ianw-oai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved as a focused executable-span Jinja SSTI false-positive fix with preserved active-context coverage and green checks.

@ianw-oai
ianw-oai dismissed their stale review June 11, 2026 16:19

Dismissed by approval sweep because GitHub recomputed this current head as dirty after approval; needs conflict resolution before the simple-approval path.

Avoid letting malformed ignored comment/raw regions consume later active Jinja spans in the no-Jinja fallback parser. Add synthetic regressions for later active requests payloads after malformed ignored regions.
@mldangelo-oai
mldangelo-oai merged commit 4cd9a18 into main Jun 11, 2026
29 checks passed
@mldangelo-oai
mldangelo-oai deleted the mdangelo/codex/hf-fp-t25-jinja-ssti-executable-context-20260610 branch June 11, 2026 17:29
@github-actions github-actions Bot mentioned this pull request Jun 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants