You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: architecture/build.md
+58-17Lines changed: 58 additions & 17 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -288,7 +288,7 @@ the release tag.
288
288
289
289
## CI and E2E
290
290
291
-
Required checks run on GitHub Actions. Workflows that use NVIDIA self-hosted runners trigger from copy-pr-bot mirror branches, so trusted PRs are mirrored into `pull-request/<N>` branches before those workflows run. `main` also uses GitHub merge queue so the final queued integration commit is validated before it merges.
291
+
Required checks run on GitHub Actions. Pull-request workflows that use NVIDIA self-hosted runners trigger from copy-pr-bot mirror branches, so trusted PRs are mirrored into `pull-request/<N>` branches before those workflows run. `main` also uses GitHub merge queue so the final queued integration commit is validated before it merges.
292
292
293
293
The high-level CI model:
294
294
@@ -307,16 +307,19 @@ synthetic activity from contributing to product usage metrics.
307
307
Static security checks are deliberately outside the mirror-branch path. They run
308
308
directly on GitHub-hosted runners and none of them consume NVIDIA self-hosted
309
309
capacity. The change-oriented ones receive no secrets, so they also cover fork
310
-
pull requests; Codex Security release qualification is the exception because it
311
-
needs a scoped API key. That key routes Codex Security's model calls to
312
-
NVIDIA-hosted inference; the job itself still runs on a GitHub-hosted runner and
313
-
uses no NVIDIA self-hosted runner. Scanner jobs request `security-events: write`
314
-
and upload SARIF to Code Scanning directly on every event they run on, including
315
-
fork and Dependabot pull requests, which Code Scanning permits for
310
+
pull requests. Codex Security release qualification is the exception: it needs a
311
+
scoped API key, which routes its model calls to NVIDIA-hosted inference while
312
+
the job itself stays GitHub-hosted. That placement is load-bearing rather than
313
+
incidental: on the repository self-hosted runner the scan agent executes no
314
+
shell commands at all, so its preflight never scopes the diff and it seals no
315
+
draft. Scanner jobs request `security-events: write` and upload SARIF to Code
316
+
Scanning directly on every event they run on, including fork and Dependabot
317
+
pull requests, which Code Scanning permits for
316
318
`pull_request` runs despite their read-only `GITHUB_TOKEN`. No privileged
317
-
intermediate workflow relays those uploads. Report retention differs by scanner:
318
-
Actionlint, Zizmor, and CodeQL keep their reports as workflow artifacts, and
319
-
Codex Security keeps no raw report.
319
+
intermediate workflow relays those uploads. Manually dispatched Codex Security
320
+
runs are the one opt-in exception, described below. Report retention differs by
321
+
scanner: Actionlint, Zizmor, and CodeQL keep their reports as workflow artifacts,
322
+
and Codex Security keeps no raw report.
320
323
Triggers differ by workflow: `.github/workflows/workflow-security.yml` runs on
321
324
`pull_request`, `merge_group`, `main`, and a weekly schedule;
322
325
`.github/workflows/dependency-review.yml` runs on `pull_request` and
@@ -361,11 +364,23 @@ a pull request or merge group.
361
364
calls go to NVIDIA-hosted inference at `https://inference-api.nvidia.com/v1`,
362
365
declared as a custom Codex provider named `nvidia` that uses the Responses
363
366
wire API with WebSockets disabled. The scan runs `openai/openai/gpt-5.6-sol`
364
-
at `medium` reasoning effort. The `CODEX_SECURITY_API_KEY` secret holds the
367
+
at `medium` reasoning effort, with the multi-agent runtime capped at eight
368
+
concurrent threads through
369
+
`features.multi_agent_v2.max_concurrent_threads_per_session`. The
370
+
`CODEX_SECURITY_API_KEY` secret holds the
365
371
NVIDIA key and is exposed to the scan step alone, as `OPENAI_API_KEY` so the
366
372
CLI selects API-key auth and as `NVIDIA_INFERENCE_API_KEY`, the provider
367
-
`env_key` read by the Codex child process.
368
-
`tasks/scripts/codex-security-release-range.mjs` resolves the scan range: the
373
+
`env_key` read by the Codex child process. `CODEX_SECURITY_STATE_DIR` and
374
+
`SCAN_DIR` are suffixed with `github.run_id` and `github.run_attempt` and
375
+
created mode `700`, so no scanner state or result set from a previous run or
376
+
retry attempt is reused even on a runner with a reusable temp directory.
377
+
`tasks/scripts/codex_security_range.py` resolves the scan range, reusing the
378
+
tag parsers in `tasks/scripts/release.py` so both stay on one definition of a
379
+
release tag while requiring the `v` prefix that a release workflow needs. The
380
+
job stages both files out of the workspace from the workflow's own revision
381
+
and runs the resolver by absolute path, because a scanned candidate predates
382
+
them and a revision under scan must not choose its own scan range. The range
383
+
itself is resolved against the checked-out candidate: the
369
384
candidate must be a `vX.Y.Z-pre.N` tag that is an ancestor of `origin/main`,
370
385
and the base is the newest stable `vX.Y.Z` tag merged into the candidate that
371
386
is strictly older than the release train `vX.Y.Z` the candidate targets. A
@@ -374,12 +389,38 @@ a pull request or merge group.
374
389
stable-to-candidate diff, so later candidates re-cover earlier ones. SARIF is
375
390
uploaded against `refs/heads/main` at the candidate commit under the
376
391
train-scoped category `codex-security/vX.Y.Z`, which makes each candidate's
377
-
analysis replace the previous one for that train. Codex Security 0.1.24 cannot
378
-
apply `--max-cost` to a slash-qualified model identifier, so the run has no
379
-
CLI-enforced cost ceiling. Spend is bounded instead by the 120-minute job
380
-
timeout, a single repository-wide concurrency group that serializes
392
+
analysis replace the previous one for that train. Automatic pre-release tag
393
+
pushes and `workflow_call` runs always upload. `workflow_dispatch` runs still
394
+
perform the scan and the SARIF export, but skip the Code Scanning upload
395
+
unless the caller sets the `upload_sarif` input, so manual diagnostics do not
396
+
overwrite a train's published analysis by default. Codex Security 0.1.24
397
+
cannot apply `--max-cost` to a slash-qualified model identifier, so the run
398
+
has no CLI-enforced cost ceiling. Spend is bounded instead by the 120-minute
399
+
job timeout, a single repository-wide concurrency group that serializes
381
400
qualification so starting a newer candidate cancels an in-flight one, and
382
401
NVIDIA account-side controls. No raw report is retained.
402
+
- The job clears `kernel.apparmor_restrict_unprivileged_userns` before
403
+
installing the scanner. Codex confines model-run commands with bubblewrap,
404
+
which needs unprivileged user namespaces; Ubuntu 24.04 restricts those through
405
+
AppArmor, so bubblewrap fails to configure the sandbox network namespace
406
+
(`bwrap: loopback: Failed RTM_NEWADDR`) and the agent executes no commands at
407
+
all. The failure is silent: the agent retries its shell tool, gives up, and
408
+
seals no draft, while the scanner only reports a missing or incomplete draft.
409
+
Lifting a kernel restriction on the runner is what allows the sandbox that
410
+
confines the agent to start, and the runner is ephemeral and GitHub-hosted.
411
+
- The scan sets `approval_policy="never"`. Codex Security keeps
412
+
`approvals_reviewer="auto_review"` unconditionally, and that reviewer runs on
413
+
its own model rather than the configured one. Because the workflow declares a
414
+
single provider that serves only `openai/openai/gpt-5.6-sol`, any approval
415
+
request reaches a model the endpoint does not serve, so the agent never gets a
416
+
shell command approved and seals no draft. The scan stays confined by its
417
+
`workspace-write` sandbox with network access disabled and by the scanner's
418
+
own permission profile, which grants read access to the filesystem root and
419
+
write access only to the workspace roots.
420
+
- A scan that cannot execute commands reports only a missing or incomplete
421
+
draft, so diagnosing one means reading the scanner's session rollouts under
422
+
`CODEX_SECURITY_STATE_DIR`, where every shell command the agent ran is
423
+
recorded. No command at all is the signal that the sandbox failed to start.
383
424
384
425
Findings never fail these checks; scanner and build failures do. A scanner that
385
426
cannot run, a CodeQL analyzer that does not complete, an unexpected Dependency
0 commit comments