Skip to content

fix(llamafile): stream executable runtime coverage - #1622

Merged
mldangelo-oai merged 12 commits into
mainfrom
mdangelo/codex/fix-llamafile-runtime-coverage
Jun 11, 2026
Merged

mldangelo-oai merged 12 commits into
mainfrom
mdangelo/codex/fix-llamafile-runtime-coverage

Conversation

@mldangelo-oai

Copy link
Copy Markdown
Contributor

Summary

  • stream executable runtime bytes instead of relying only on head, tail, and marker previews
  • preserve overlap across chunks so command and network indicators spanning boundaries remain detectable
  • mark the scan inconclusive whenever runtime bytes are omitted by the configured bound or a stream read is incomplete
  • retain bounded embedded-payload and Torch7 handling

This closes the clean-result coverage gap for indicators hidden between preview windows.

Tests

  • pytest tests/scanners/test_llamafile_scanner.py -q (62 passed)
  • ruff check modelaudit/scanners/llamafile_scanner.py tests/scanners/test_llamafile_scanner.py
  • ruff format --check modelaudit/scanners/llamafile_scanner.py tests/scanners/test_llamafile_scanner.py
  • mypy modelaudit/scanners/llamafile_scanner.py tests/scanners/test_llamafile_scanner.py
  • git diff --check

@github-actions

github-actions Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Workflow run and artifacts

Performance Benchmarks

Compared 12 shared benchmarks with a regression threshold of 15%.
Status: 0 regressions, 0 improved, 12 stable, 0 new, 0 missing.
Aggregate shared-benchmark median: 1.434s -> 1.433s (-0.1%).

Workload Benchmark Target Size Files Baseline Current Change Status
warm-cache-rescan tests/benchmarks/test_scan_benchmarks.py::test_scan_warm_cached_repository_rescan release-candidate 547.3 KiB 32 109.86ms 105.87ms -3.6% stable
mixed-model-repository tests/benchmarks/test_scan_benchmarks.py::test_scan_release_candidate_repository release-candidate 547.3 KiB 32 480.52ms 490.76ms +2.1% stable
direct-malicious-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_direct_malicious_upload malicious_reduce 52 B 1 523.6us 513.3us -2.0% stable
suspicious-pickle-intake tests/benchmarks/test_scan_benchmarks.py::test_scan_suspicious_pickle_intake suspicious-intake 183.8 KiB 4 146.89ms 144.10ms -1.9% stable
chunked-upload-stream tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_chunked_upload_stream chunked_stream 278.2 KiB 1 114.56ms 112.53ms -1.8% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_raw] nested_raw 78 B 1 584.7us 575.0us -1.7% stable
clean-training-checkpoint tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_clean_training_checkpoint safe_large 278.2 KiB 1 110.33ms 109.16ms -1.1% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_base64] nested_base64 98 B 1 565.8us 569.9us +0.7% stable
padded-multi-stream-upload tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_padded_multi_stream_upload multi_stream_padded 4.1 KiB 1 637.5us 640.0us +0.4% stable
single-checkpoint-preflight tests/benchmarks/test_scan_benchmarks.py::test_scan_single_checkpoint_before_load single_checkpoint.pkl 183.0 KiB 1 72.54ms 72.28ms -0.4% stable
duplicate-heavy-registry tests/benchmarks/test_scan_benchmarks.py::test_scan_duplicate_registry_snapshot registry-snapshot 915.2 KiB 13 396.77ms 395.71ms -0.3% stable
nested-payload-review tests/benchmarks/test_picklescan_benchmarks.py::test_picklescan_nested_payload_review[nested_hex] nested_hex 130 B 1 612.4us 611.1us -0.2% stable

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 37e8678a2b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread modelaudit/scanners/llamafile_scanner.py Outdated
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

Second-pass sidecar review on exact head 37e8678a2b13facb0de8f4c618f7c9d50ccf73e7 found several additional substantive issues beyond the current open thread. Summary:

  1. High: unbounded evidence retention defeats bounded streaming.

    • modelaudit/scanners/llamafile_scanner.py:437 and :453
    • The scanner retains every distinct hit and then sorts all of them even though it only emits five. The sidecar reproduced 69,905 unique strings from a 1 MiB probe; at the 512 MiB default limit this can grow into multi-GB memory use.
  2. High: the first parseable in-band GGUF header truncates later runtime coverage.

    • modelaudit/scanners/llamafile_scanner.py:305 and :1035
    • A valid empty GGUF decoy earlier in the file can suppress malicious runtime evidence later in the same file because candidate search stops at the first parseable header. A malformed first marker can also mask a later valid malicious GGUF.
  3. High: embedded GGUF incompleteness fails open.

    • modelaudit/scanners/llamafile_scanner.py:565 and :571
    • The child GGUF path can return an inconclusive/resource-limited result, but the outer scan still reports success with exit 0 and no outcome reason. The sidecar also confirmed the inverse false-positive shape from the current unresolved thread: recognized-but-inconclusive GGUF bytes can still be scanned as runtime.
  4. Medium: payload bytes are mislabeled as executable runtime.

    • modelaudit/scanners/llamafile_scanner.py:455
    • Tail preview analysis is not clipped to runtime_end, so inert GGUF metadata / payload text can surface as critical runtime findings.
  5. Medium: fixed overlap changes whole-record semantics.

    • modelaudit/scanners/llamafile_scanner.py:470 and :473
    • Retaining only 512 bytes of a partial printable record can both miss long malicious records and fabricate false positives from truncated safe records.

The sidecar validated these against the live head with 62 targeted Llamafile tests plus Ruff/format/mypy green. Its merge-safety call was explicit: do not merge yet.

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Exact head: d54e4092834ece78c8a640b75cee21aa14c61430.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Can't wait for the next one!

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Exact head: 08b2cdf0fdc224e73dc88fd56360912bb1ade865 after merging latest main (37289644f219a57e4c043dcc083c3eea434aaca3).

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 08b2cdf0fd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread modelaudit/scanners/llamafile_scanner.py Outdated
Prioritize the selected payload boundary while sharing the bounded carve budget across retained GGUF candidates.
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Please review exact head e9901fa8f89dd815c67ebe406ce39f961c735103.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e9901fa8f8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread modelaudit/scanners/llamafile_scanner.py
Comment thread modelaudit/scanners/llamafile_scanner.py Outdated
Comment thread modelaudit/scanners/llamafile_scanner.py
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Please review exact head 9907c8c9560619e5c1bed184ba248c699ccfd4db. All previously actionable threads are resolved. Local validation: 536 Llamafile tests pass; repo-wide Ruff format/check and mypy pass; malicious-positive and benign-negative runtime regressions are included.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9907c8c956

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread modelaudit/scanners/llamafile_scanner.py
@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

@mldangelo-oai

Copy link
Copy Markdown
Contributor Author

@codex review

Please review exact head f7c0af9af82ad6e43ffc50f7ed80ea0a3b101d12. The remaining heredoc fallback thread is resolved on this head; local validation includes the malicious-positive and benign heredoc negative regressions, 537 Llamafile tests, Ruff format/check, mypy, and the broad practical floor.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Chef's kiss.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@mldangelo-oai
mldangelo-oai merged commit 8d6c486 into main Jun 11, 2026
29 checks passed
@mldangelo-oai
mldangelo-oai deleted the mdangelo/codex/fix-llamafile-runtime-coverage branch June 11, 2026 00:57
@mldangelo-oai
mldangelo-oai requested a review from mldangelo June 11, 2026 00:58
@github-actions github-actions Bot mentioned this pull request Jun 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant