Skip to content

Commit 64a858d

Browse files
authored
fix(ci): restore Codex Security scan execution (#3124)
* refactor(ci): resolve Codex Security range in Python Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(ci): allow unprivileged userns for Codex sandbox Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com>
1 parent e64b035 commit 64a858d

7 files changed

Lines changed: 509 additions & 445 deletions

File tree

‎.github/workflows/codex-security.yml‎

Lines changed: 47 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -49,6 +49,11 @@ on:
4949
description: Optional previous stable tag override
5050
required: false
5151
type: string
52+
upload_sarif:
53+
description: Upload manual-run results to Code Scanning
54+
required: false
55+
default: false
56+
type: boolean
5257
allow_full_bootstrap:
5358
description: Allow a full scan when no previous stable tag exists
5459
required: false
@@ -65,13 +70,16 @@ concurrency:
6570
cancel-in-progress: true
6671

6772
env:
73+
CODEX_SECURITY_MAX_CONCURRENT_THREADS: "8"
6874
CODEX_SECURITY_REASONING_EFFORT: medium
6975
NVIDIA_INFERENCE_BASE_URL: https://inference-api.nvidia.com/v1
7076
NVIDIA_INFERENCE_MODEL: openai/openai/gpt-5.6-sol
7177

7278
jobs:
7379
analyze:
7480
name: Codex Security (${{ inputs.candidate_ref || github.ref_name }})
81+
# The agent executes no shell commands on the repository self-hosted runner,
82+
# so its preflight never scopes the diff and it seals no draft.
7583
runs-on: ubuntu-latest
7684
timeout-minutes: 120
7785
outputs:
@@ -93,6 +101,17 @@ jobs:
93101
with:
94102
python-version: "3.14"
95103

104+
# Codex confines model-run commands with bubblewrap, which needs
105+
# unprivileged user namespaces. Ubuntu 24.04 restricts those through
106+
# AppArmor, so bubblewrap cannot set up the sandbox network namespace and
107+
# the scan agent executes nothing at all.
108+
- name: Allow unprivileged user namespaces
109+
run: |
110+
set -euo pipefail
111+
if [ -e /proc/sys/kernel/apparmor_restrict_unprivileged_userns ]; then
112+
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
113+
fi
114+
96115
- name: Install Codex Security
97116
run: |
98117
set -euo pipefail
@@ -111,6 +130,25 @@ jobs:
111130
test -x "$CODEX_SECURITY_BIN"
112131
"$CODEX_SECURITY_BIN" --version
113132
133+
# The range resolver has to come from the workflow's own revision. A
134+
# scanned candidate predates it, and running the resolver from the
135+
# revision under scan would let that revision pick its own scan range.
136+
- name: Check out the workflow revision
137+
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
138+
with:
139+
path: workflow-revision
140+
sparse-checkout: tasks/scripts
141+
persist-credentials: false
142+
143+
- name: Stage the range resolver
144+
run: |
145+
set -euo pipefail
146+
install -d -m 700 "$RUNNER_TEMP/range-resolver"
147+
cp workflow-revision/tasks/scripts/release.py \
148+
workflow-revision/tasks/scripts/codex_security_range.py \
149+
"$RUNNER_TEMP/range-resolver/"
150+
rm -rf workflow-revision
151+
114152
- name: Check out the pre-release
115153
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
116154
with:
@@ -139,17 +177,17 @@ jobs:
139177
args+=(--allow-full-bootstrap)
140178
fi
141179
142-
node tasks/scripts/codex-security-release-range.mjs "${args[@]}"
180+
python3 "$RUNNER_TEMP/range-resolver/codex_security_range.py" "${args[@]}"
143181
144182
- name: Scan changes since the previous stable
145183
env:
146184
BASE_SHA: ${{ steps.range.outputs.base_sha }}
147185
CODEX_SECURITY_BIN: ${{ runner.temp }}/codex-security/node_modules/.bin/codex-security
148-
CODEX_SECURITY_STATE_DIR: ${{ runner.temp }}/codex-security-state
186+
CODEX_SECURITY_STATE_DIR: ${{ runner.temp }}/codex-security-state-${{ github.run_id }}-${{ github.run_attempt }}
149187
HEAD_SHA: ${{ steps.range.outputs.candidate_sha }}
150188
NVIDIA_INFERENCE_API_KEY: ${{ secrets.CODEX_SECURITY_API_KEY }}
151189
OPENAI_API_KEY: ${{ secrets.CODEX_SECURITY_API_KEY }}
152-
SCAN_DIR: ${{ runner.temp }}/codex-security-results
190+
SCAN_DIR: ${{ runner.temp }}/codex-security-results-${{ github.run_id }}-${{ github.run_attempt }}
153191
SCAN_SCOPE: ${{ steps.range.outputs.scan_scope }}
154192
run: |
155193
set -euo pipefail
@@ -174,14 +212,16 @@ jobs:
174212
--codex 'model_providers.nvidia.env_key="NVIDIA_INFERENCE_API_KEY"' \
175213
--codex 'model_providers.nvidia.wire_api="responses"' \
176214
--codex 'model_providers.nvidia.supports_websockets=false' \
215+
--codex "features.multi_agent_v2.max_concurrent_threads_per_session=$CODEX_SECURITY_MAX_CONCURRENT_THREADS" \
216+
--codex 'approval_policy="never"' \
177217
--output-dir "$SCAN_DIR" \
178218
--headless > /dev/null
179219
180220
- name: Export SARIF
181221
env:
182222
CODEX_SECURITY_BIN: ${{ runner.temp }}/codex-security/node_modules/.bin/codex-security
183223
SARIF_FILE: ${{ runner.temp }}/codex-security.sarif
184-
SCAN_DIR: ${{ runner.temp }}/codex-security-results
224+
SCAN_DIR: ${{ runner.temp }}/codex-security-results-${{ github.run_id }}-${{ github.run_attempt }}
185225
run: |
186226
set -euo pipefail
187227
"$CODEX_SECURITY_BIN" export "$SCAN_DIR" \
@@ -205,6 +245,7 @@ jobs:
205245
echo
206246
echo "- Train: \`$TRAIN\`"
207247
echo "- Candidate: \`$CANDIDATE_TAG\`"
248+
echo "- Maximum concurrent agent threads: $CODEX_SECURITY_MAX_CONCURRENT_THREADS"
208249
if [ "$SCAN_SCOPE" = "diff" ]; then
209250
echo "- Previous stable: \`$BASE_TAG\`"
210251
echo "- Commits in cumulative diff: $COMMIT_COUNT"
@@ -218,6 +259,7 @@ jobs:
218259
} >> "$GITHUB_STEP_SUMMARY"
219260
220261
- name: Upload SARIF to Code Scanning
262+
if: ${{ github.event_name != 'workflow_dispatch' || inputs.upload_sarif }}
221263
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
222264
with:
223265
sarif_file: ${{ runner.temp }}/codex-security.sarif
@@ -227,7 +269,7 @@ jobs:
227269

228270
result:
229271
name: OpenShell / Codex Security (informational)
230-
if: always()
272+
if: ${{ always() }}
231273
needs: analyze
232274
runs-on: ubuntu-latest
233275
permissions: {}

‎architecture/build.md‎

Lines changed: 58 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -288,7 +288,7 @@ the release tag.
288288

289289
## CI and E2E
290290

291-
Required checks run on GitHub Actions. Workflows that use NVIDIA self-hosted runners trigger from copy-pr-bot mirror branches, so trusted PRs are mirrored into `pull-request/<N>` branches before those workflows run. `main` also uses GitHub merge queue so the final queued integration commit is validated before it merges.
291+
Required checks run on GitHub Actions. Pull-request workflows that use NVIDIA self-hosted runners trigger from copy-pr-bot mirror branches, so trusted PRs are mirrored into `pull-request/<N>` branches before those workflows run. `main` also uses GitHub merge queue so the final queued integration commit is validated before it merges.
292292

293293
The high-level CI model:
294294

@@ -307,16 +307,19 @@ synthetic activity from contributing to product usage metrics.
307307
Static security checks are deliberately outside the mirror-branch path. They run
308308
directly on GitHub-hosted runners and none of them consume NVIDIA self-hosted
309309
capacity. The change-oriented ones receive no secrets, so they also cover fork
310-
pull requests; Codex Security release qualification is the exception because it
311-
needs a scoped API key. That key routes Codex Security's model calls to
312-
NVIDIA-hosted inference; the job itself still runs on a GitHub-hosted runner and
313-
uses no NVIDIA self-hosted runner. Scanner jobs request `security-events: write`
314-
and upload SARIF to Code Scanning directly on every event they run on, including
315-
fork and Dependabot pull requests, which Code Scanning permits for
310+
pull requests. Codex Security release qualification is the exception: it needs a
311+
scoped API key, which routes its model calls to NVIDIA-hosted inference while
312+
the job itself stays GitHub-hosted. That placement is load-bearing rather than
313+
incidental: on the repository self-hosted runner the scan agent executes no
314+
shell commands at all, so its preflight never scopes the diff and it seals no
315+
draft. Scanner jobs request `security-events: write` and upload SARIF to Code
316+
Scanning directly on every event they run on, including fork and Dependabot
317+
pull requests, which Code Scanning permits for
316318
`pull_request` runs despite their read-only `GITHUB_TOKEN`. No privileged
317-
intermediate workflow relays those uploads. Report retention differs by scanner:
318-
Actionlint, Zizmor, and CodeQL keep their reports as workflow artifacts, and
319-
Codex Security keeps no raw report.
319+
intermediate workflow relays those uploads. Manually dispatched Codex Security
320+
runs are the one opt-in exception, described below. Report retention differs by
321+
scanner: Actionlint, Zizmor, and CodeQL keep their reports as workflow artifacts,
322+
and Codex Security keeps no raw report.
320323
Triggers differ by workflow: `.github/workflows/workflow-security.yml` runs on
321324
`pull_request`, `merge_group`, `main`, and a weekly schedule;
322325
`.github/workflows/dependency-review.yml` runs on `pull_request` and
@@ -361,11 +364,23 @@ a pull request or merge group.
361364
calls go to NVIDIA-hosted inference at `https://inference-api.nvidia.com/v1`,
362365
declared as a custom Codex provider named `nvidia` that uses the Responses
363366
wire API with WebSockets disabled. The scan runs `openai/openai/gpt-5.6-sol`
364-
at `medium` reasoning effort. The `CODEX_SECURITY_API_KEY` secret holds the
367+
at `medium` reasoning effort, with the multi-agent runtime capped at eight
368+
concurrent threads through
369+
`features.multi_agent_v2.max_concurrent_threads_per_session`. The
370+
`CODEX_SECURITY_API_KEY` secret holds the
365371
NVIDIA key and is exposed to the scan step alone, as `OPENAI_API_KEY` so the
366372
CLI selects API-key auth and as `NVIDIA_INFERENCE_API_KEY`, the provider
367-
`env_key` read by the Codex child process.
368-
`tasks/scripts/codex-security-release-range.mjs` resolves the scan range: the
373+
`env_key` read by the Codex child process. `CODEX_SECURITY_STATE_DIR` and
374+
`SCAN_DIR` are suffixed with `github.run_id` and `github.run_attempt` and
375+
created mode `700`, so no scanner state or result set from a previous run or
376+
retry attempt is reused even on a runner with a reusable temp directory.
377+
`tasks/scripts/codex_security_range.py` resolves the scan range, reusing the
378+
tag parsers in `tasks/scripts/release.py` so both stay on one definition of a
379+
release tag while requiring the `v` prefix that a release workflow needs. The
380+
job stages both files out of the workspace from the workflow's own revision
381+
and runs the resolver by absolute path, because a scanned candidate predates
382+
them and a revision under scan must not choose its own scan range. The range
383+
itself is resolved against the checked-out candidate: the
369384
candidate must be a `vX.Y.Z-pre.N` tag that is an ancestor of `origin/main`,
370385
and the base is the newest stable `vX.Y.Z` tag merged into the candidate that
371386
is strictly older than the release train `vX.Y.Z` the candidate targets. A
@@ -374,12 +389,38 @@ a pull request or merge group.
374389
stable-to-candidate diff, so later candidates re-cover earlier ones. SARIF is
375390
uploaded against `refs/heads/main` at the candidate commit under the
376391
train-scoped category `codex-security/vX.Y.Z`, which makes each candidate's
377-
analysis replace the previous one for that train. Codex Security 0.1.24 cannot
378-
apply `--max-cost` to a slash-qualified model identifier, so the run has no
379-
CLI-enforced cost ceiling. Spend is bounded instead by the 120-minute job
380-
timeout, a single repository-wide concurrency group that serializes
392+
analysis replace the previous one for that train. Automatic pre-release tag
393+
pushes and `workflow_call` runs always upload. `workflow_dispatch` runs still
394+
perform the scan and the SARIF export, but skip the Code Scanning upload
395+
unless the caller sets the `upload_sarif` input, so manual diagnostics do not
396+
overwrite a train's published analysis by default. Codex Security 0.1.24
397+
cannot apply `--max-cost` to a slash-qualified model identifier, so the run
398+
has no CLI-enforced cost ceiling. Spend is bounded instead by the 120-minute
399+
job timeout, a single repository-wide concurrency group that serializes
381400
qualification so starting a newer candidate cancels an in-flight one, and
382401
NVIDIA account-side controls. No raw report is retained.
402+
- The job clears `kernel.apparmor_restrict_unprivileged_userns` before
403+
installing the scanner. Codex confines model-run commands with bubblewrap,
404+
which needs unprivileged user namespaces; Ubuntu 24.04 restricts those through
405+
AppArmor, so bubblewrap fails to configure the sandbox network namespace
406+
(`bwrap: loopback: Failed RTM_NEWADDR`) and the agent executes no commands at
407+
all. The failure is silent: the agent retries its shell tool, gives up, and
408+
seals no draft, while the scanner only reports a missing or incomplete draft.
409+
Lifting a kernel restriction on the runner is what allows the sandbox that
410+
confines the agent to start, and the runner is ephemeral and GitHub-hosted.
411+
- The scan sets `approval_policy="never"`. Codex Security keeps
412+
`approvals_reviewer="auto_review"` unconditionally, and that reviewer runs on
413+
its own model rather than the configured one. Because the workflow declares a
414+
single provider that serves only `openai/openai/gpt-5.6-sol`, any approval
415+
request reaches a model the endpoint does not serve, so the agent never gets a
416+
shell command approved and seals no draft. The scan stays confined by its
417+
`workspace-write` sandbox with network access disabled and by the scanner's
418+
own permission profile, which grants read access to the filesystem root and
419+
write access only to the workspace roots.
420+
- A scan that cannot execute commands reports only a missing or incomplete
421+
draft, so diagnosing one means reading the scanner's session rollouts under
422+
`CODEX_SECURITY_STATE_DIR`, where every shell command the agent ran is
423+
recorded. No command at all is the signal that the sandbox failed to start.
383424

384425
Findings never fail these checks; scanner and build failures do. A scanner that
385426
cannot run, a CodeQL analyzer that does not complete, an unexpected Dependency

0 commit comments

Comments
 (0)