Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
49 commits
Select commit Hold shift + click to select a range
8d473c9
chore(bench): remove the llm-d lane, SWE-bench Lite and native power …
adibarra Oct 2, 2026
dec7571
chore(bench): drop the vLLM chat-template-kwargs and TRT-LLM perf-met…
cquil11 Oct 2, 2026
35813cc
chore(bench): remove benchmark_lib.sh helpers nothing calls and llm-d…
cquil11 Oct 2, 2026
9aa0f62
chore(agentx): stop generating post-run plots
cquil11 Oct 2, 2026
f2c7776
refactor(bench): replace benchmark_lib.sh with the infx.bench package
adibarra Oct 2, 2026
ebfbb8b
chore(agentx): drop the post-run plot step from infx.bench.agentic
cquil11 Oct 2, 2026
43d0890
refactor: retire client-side GPU sampling
edwingao28 Oct 2, 2026
8010ec5
docs: link power cleanup changelog to PR 3684
edwingao28 Oct 2, 2026
a6bfe0f
fix: stabilize Kimi archive test collection IDs
edwingao28 Oct 2, 2026
cb6901a
fix: wire native power collection for single-node jobs
edwingao28 Oct 4, 2026
0a7fd2b
refactor: name single-node result finalization explicitly
edwingao28 Oct 4, 2026
6de7062
refactor(bench): address review on infx.bench
adibarra Oct 5, 2026
b9907c3
fix: reconcile native telemetry with benchmark client refactor
edwingao28 Oct 5, 2026
eb46b79
docs: simplify native power telemetry setup guidance
edwingao28 Oct 5, 2026
ce046ca
chore: reconcile native telemetry with main
edwingao28 Oct 6, 2026
2a4f2e3
chore: sync native telemetry branch with main
edwingao28 Oct 6, 2026
a529889
fix: retain llm-d after base synchronization
edwingao28 Oct 6, 2026
d2c0fec
chore: merge native telemetry lane from #3684
edwingao28 Oct 7, 2026
6ef006f
feat: accept SRT temperature sample artifacts
edwingao28 Oct 2, 2026
4a76ebf
docs: link temperature artifacts changelog to PR 3685
edwingao28 Oct 2, 2026
1d93404
chore: add prometheus-client to the test group for srt-slurm power im…
edwingao28 Oct 7, 2026
b7d9e20
chore: drop the legacy single-node gpu_metrics.csv power consumer
edwingao28 Oct 7, 2026
56e4903
feat: run the AMD power exporter through srt-slurm #572
edwingao28 Oct 7, 2026
8af9204
fix: accept any recorded power metric in srt-slurm manifests
edwingao28 Oct 7, 2026
e17362b
fix: cap v3 sample temperatures at the producer's 200 C bound
edwingao28 Oct 7, 2026
f6c693c
docs: describe the srt-slurm #572 AMD exporter schema and single-node…
edwingao28 Oct 7, 2026
5002dcb
test: replay 641a07f2 AMD and DCGM power packages through the consumers
edwingao28 Oct 7, 2026
1006773
Merge branch 'feat/powerx-amd-upstream-572-launcher' into feat/powerx…
edwingao28 Oct 7, 2026
da439eb
docs: append the srt-slurm #572 AMD exporter changelog entry
edwingao28 Oct 7, 2026
0fe9a32
Merge branch 'feat/powerx-amd-upstream-572-contract' into feat/powerx…
edwingao28 Oct 7, 2026
b81502e
Merge branch 'feat/powerx-amd-upstream-572-docs' into feat/powerx-amd…
edwingao28 Oct 7, 2026
616a066
docs: drop the retired single-node power consumer command
edwingao28 Oct 7, 2026
e48397f
docs: link the srt-slurm #572 AMD exporter changelog entry to PR 3781
edwingao28 Oct 7, 2026
ed75366
fix: honor require-power on single-node jobs and stop the AgentX laun…
edwingao28 Oct 7, 2026
a6a571b
fix: keep clock-sync-refused power packages unpublishable in the cons…
edwingao28 Oct 7, 2026
4847f2d
fix: publish the single-node AgentX power sidecar beside the result
edwingao28 Oct 7, 2026
4172b89
chore: merge main into feat/powerx-amd-upstream-572
edwingao28 Oct 7, 2026
1086a39
Merge branch 'main' into feat/powerx-amd-upstream-572
edwingao28 Oct 7, 2026
e6d28aa
chore: restore main's docs after the branch rebase replayed old sync …
edwingao28 Oct 7, 2026
c9fbe32
docs: keep one perf-changelog entry for PR 3781
edwingao28 Oct 7, 2026
bd0ce47
fix: keep single-node power staging best-effort unless power is required
edwingao28 Oct 7, 2026
c5794f6
chore: sync AMD telemetry PR with main
edwingao28 Oct 8, 2026
110b8de
chore: pin the rebuilt AMD exporter image
edwingao28 Oct 8, 2026
c591f0c
fix: give the single-node exporter 300 s to come up
edwingao28 Oct 8, 2026
9594478
chore: cite only the measured mi355x pull time for the exporter timeout
edwingao28 Oct 8, 2026
050bf0b
chore: restore the multinode workflow template to main
edwingao28 Oct 7, 2026
25bc5c2
chore: sync AMD power telemetry PR with main
edwingao28 Oct 9, 2026
3a4ed70
chore: sync AMD telemetry PR with current main
edwingao28 Oct 9, 2026
b7825b6
chore: sync AMD telemetry PR with current main
edwingao28 Oct 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .github/AGENT_OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,15 +68,15 @@ Eval selection, the `--no-evals` / `--evals-only` / `--all-evals` flags, and cha

## Power telemetry

Multinode srt-slurm results may include `power_valid`, `avg_power_w`, `avg_total_gpu_power_w`, `total_gpu_energy_j`, and joules per query/input/output/total token. Invalid telemetry records `power_valid: 0` without energy metrics and fails only with `REQUIRE_POWER=1`. Single-node results carry no power fields until srt-slurm telemetry covers those lanes.
Multinode srt-slurm results, and single-node fixed-sequence and AgentX results with a retained native telemetry package, may include `power_valid`, `avg_power_w`, `avg_total_gpu_power_w`, `total_gpu_energy_j`, and joules per query/input/output/total token. Invalid telemetry records `power_valid: 0` without energy metrics and fails only with `REQUIRE_POWER=1`.

Multinode disaggregated results add `prefill_gpu_energy_j`, `decode_gpu_energy_j`, `prefill_avg_power_w`, `decode_avg_power_w`, `prefill_joules_per_input_token`, and `decode_joules_per_output_token`. Role energy covers the full formal benchmark window, not kernel-level phases, and the role watts are that energy divided by the same window and by the role's GPU count.

Every power result, valid or invalid, carries `power_metric_schema_version`. Version 2 defines each unprefixed `joules_per_*` field as whole-deployment GPU-board energy over the named denominator; role-scoped energy uses the explicit `prefill_*` / `decode_*` keys. Rows without the field predate the whole-deployment switch and their unprefixed joules are not comparable across topologies.

For srt-slurm recipes, `telemetry.enabled: true` with `telemetry.dcgm_exporter` enables official energy collection. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for every srt-slurm job, including TileRT. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated.
Multinode srt-slurm recipes opt into official energy collection with `telemetry.enabled: true`; the launcher enables it for single-node jobs. Both use the cluster's `default_gpu_exporter` from `inferencex-e2e/configs/runners.yaml` unless the recipe sets `telemetry.dcgm_exporter`: DCGM on NVIDIA, the `kind: custom` AMD device-metrics-exporter on AMD. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for every srt-slurm job, including TileRT. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated.

Power audit artifacts are named `power_audit_<result>` and contain the multinode `power_validation_<result>_*.json` sidecars. They are uploaded even when validation fails.
Power audit artifacts are named `power_audit_<result>` and contain the `power_validation*.json` sidecars plus, for native telemetry jobs, `LOGS/power/` (`samples.csv` v3 with `temperature_c`, `manifest.json`, `windows/`). They are uploaded even when validation fails.

## Result artifacts and metrics

Expand Down
35 changes: 35 additions & 0 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,11 @@ on:
required: false
type: string
default: ""
require-power:
description: "Fail result processing when GPU power telemetry is invalid"
type: boolean
required: false
default: false
env:
PYTHONPATH: ${{ github.workspace }}/inferencex-e2e
INFERENCEX_E2E_ROOT: ${{ github.workspace }}
Expand All @@ -86,6 +91,7 @@ env:
SALLOC_TIME_LIMIT: '480'
KEEP_LOGS: '0'
IS_MULTINODE: 'false'
REQUIRE_POWER: ${{ inputs.require-power && '1' || '0' }}
RANDOM_RANGE_RATIO: 0.8
HF_TOKEN: ${{ secrets.INFERENCEX_OFFICIAL_RO_HF_TOKEN }}
HF_HUB_CACHE: '/mnt/hf_hub_cache/'
Expand Down Expand Up @@ -454,6 +460,35 @@ jobs:
${{ env.INFERENCEX_E2E_ROOT }}/srt-setup.log
if-no-files-found: ignore

- name: Upload power audit bundle
if: ${{ always() && !inputs.eval-only }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: power_audit_${{ env.RESULT_FILENAME }}
path: |
${{ env.INFERENCEX_E2E_ROOT }}/${{ env.RESULT_FILENAME }}.json
${{ env.INFERENCEX_E2E_ROOT }}/agg_${{ env.RESULT_FILENAME }}.json
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics.csv
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics*_context.json
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics_energy_start.csv
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics_energy_end.csv
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics_identity.json
${{ env.INFERENCEX_E2E_ROOT }}/gpu_metrics_identity.csv
${{ env.INFERENCEX_E2E_ROOT }}/power_validation_${{ env.RESULT_FILENAME }}.json
${{ env.INFERENCEX_E2E_ROOT }}/results/gpu_metrics*.csv
${{ env.INFERENCEX_E2E_ROOT }}/results/gpu_metrics*_context.json
${{ env.INFERENCEX_E2E_ROOT }}/results/gpu_metrics_identity.json
${{ env.INFERENCEX_E2E_ROOT }}/results/agentic_power_window.json
${{ env.INFERENCEX_E2E_ROOT }}/results/agentic_power_timezone_offset.txt
${{ env.INFERENCEX_E2E_ROOT }}/results/power_validation.json
${{ env.INFERENCEX_E2E_ROOT }}/LOGS/power/**
${{ env.INFERENCEX_E2E_ROOT }}/LOGS/agentic/**/agentic_power_concurrency_*.json
${{ env.INFERENCEX_E2E_ROOT }}/LOGS/agentic/**/power_validation.json
${{ env.INFERENCEX_E2E_ROOT }}/LOGS/${{ env.RESULT_FILENAME }}.json
${{ env.INFERENCEX_E2E_ROOT }}/power-producer-sha.txt
${{ env.INFERENCEX_E2E_ROOT }}/exporter-image.sha256
if-no-files-found: ignore

- name: Upload eval results (if any)
id: upload-eval
if: ${{ always() && (env.RUN_EVAL == 'true' || inputs.eval-only) }}
Expand Down
2 changes: 2 additions & 0 deletions .github/workflows/e2e-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -395,6 +395,7 @@ jobs:
secrets:
INFERENCEX_OFFICIAL_RO_HF_TOKEN: ${{ secrets.INFERENCEX_OFFICIAL_RO_HF_TOKEN }}
with:
require-power: ${{ inputs.require-power }}
config: ${{ toJSON(matrix.config) }}
klaud-run: ${{ inputs.klaud-run }}
runner: ${{ matrix.config.runner }}
Expand Down Expand Up @@ -505,6 +506,7 @@ jobs:
secrets:
INFERENCEX_OFFICIAL_RO_HF_TOKEN: ${{ secrets.INFERENCEX_OFFICIAL_RO_HF_TOKEN }}
with:
require-power: ${{ inputs.require-power }}
config: ${{ toJSON(matrix.config) }}
klaud-run: ${{ inputs.klaud-run }}
runner: ${{ matrix.config.runner }}
Expand Down
24 changes: 24 additions & 0 deletions inferencex-e2e/configs/CONFIGS.md
Original file line number Diff line number Diff line change
Expand Up @@ -229,3 +229,27 @@ schema; unknown keys fail.
resolve to the main image and to the staged nginx, and `outputs`,
`shared-run-root` and `uv-cache-root` are the directories the srt-slurm launcher
itself uses.
- `slurm.srt-slurm.extra.default_gpu_exporter` is the cluster's native power exporter.
srtctl inherits it as the recipe's `telemetry.dcgm_exporter` for multi-node recipes that
enable telemetry and for every single-node throughput and AgentX job; eval-only jobs
collect no power, a cluster without the block stops preparation before the benchmark
starts, and an invalid measurement fails the job under `REQUIRE_POWER=1` or otherwise
records an invalid verdict without energy metrics. NVIDIA clusters run `dcgm-exporter`
on port `9401`. AMD clusters use srt-slurm's
[`kind: custom` schema](https://github.com/NVIDIA/srt-slurm/blob/641a07f2d465847fe51d8d8db275366651d9ebef/docs/power-telemetry.md#gpu-exporter-labels-and-metrics):
`gpu_labels` names the index and identity labels (`gpu_id`, `serial_number`) and
`gpu_metrics` the power, utilization and temperature metrics (`gpu_power_usage` with its
recorded `scope`, `gpu_gfx_activity`, `gpu_junction_temperature`); unknown keys such as
the former `power_profile` are rejected. The image is the public
`ghcr.io#semianalysisai/amd-device-metrics-exporter@sha256:8a3fe70b8a848ca10a7fd90862d1d9e41e9c7b6314669a2e8340c646e6dae15c`,
AMD nightly `build-dme-10.2.0a20261001` with AMD's 255 W power-reading fix plus our
cache-TTL patch, started with `env AMD_GPU_GET_CACHE_TTL=0s /home/amd/tools/entrypoint.sh`
on port `19500` and resolved by digest through the cluster's `squash` settings like any
other image. It reads
[`runners/srt-slurm/exporters/amd-power.json`](../runners/srt-slurm/exporters/amd-power.json),
which the driver mounts at `/etc/metrics/config.json`. Until NVIDIA/srt-slurm#573 merges,
[`573-participating-gpus.patch`](../runners/srt-slurm/patches/README.md) keeps
worker-node sample rows to the GPUs the job uses; without it a TP4 job on an eight-GPU
node records `unexpected_device`. Single-node matrix rows still reject `require-power`; a
manual `e2e-tests.yml` dispatch passes its `require-power` input to the single-node
throughput and AgentX jobs.
49 changes: 48 additions & 1 deletion inferencex-e2e/configs/runners.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -804,6 +804,22 @@ clusters:
/dev/dri: /dev/dri
extra:
visible_devices_env: ROCR_VISIBLE_DEVICES
default_gpu_exporter:
container_image: ghcr.io#semianalysisai/amd-device-metrics-exporter@sha256:8a3fe70b8a848ca10a7fd90862d1d9e41e9c7b6314669a2e8340c646e6dae15c
port: 19500
command: env AMD_GPU_GET_CACHE_TTL=0s /home/amd/tools/entrypoint.sh
kind: custom
gpu_labels:
index: gpu_id
identity: serial_number
gpu_metrics:
power:
metric: gpu_power_usage
scope: gpu_device_power_as_reported_by_amd_device_metrics_exporter
gpu_util:
metric: gpu_gfx_activity
temperature:
metric: gpu_junction_temperature
# Barite shares its login and NFS with mi300x-amd, but not its GPU partition.
mi325x-amd:
gpus-per-node: 8
Expand Down Expand Up @@ -834,6 +850,22 @@ clusters:
/dev/dri: /dev/dri
extra:
visible_devices_env: ROCR_VISIBLE_DEVICES
default_gpu_exporter:
container_image: ghcr.io#semianalysisai/amd-device-metrics-exporter@sha256:8a3fe70b8a848ca10a7fd90862d1d9e41e9c7b6314669a2e8340c646e6dae15c
port: 19500
command: env AMD_GPU_GET_CACHE_TTL=0s /home/amd/tools/entrypoint.sh
kind: custom
gpu_labels:
index: gpu_id
identity: serial_number
gpu_metrics:
power:
metric: gpu_power_usage
scope: gpu_device_power_as_reported_by_amd_device_metrics_exporter
gpu_util:
metric: gpu_gfx_activity
temperature:
metric: gpu_junction_temperature
default_bash_preamble: |-
export XDG_CACHE_HOME="/tmp/xdg-cache-$SLURM_JOB_ID"
export TRITON_CACHE_DIR="/tmp/triton-cache-$SLURM_JOB_ID"
Expand Down Expand Up @@ -885,5 +917,20 @@ clusters:
/it-share/hf_home: /it-share/hf_home
extra:
visible_devices_env: ROCR_VISIBLE_DEVICES
default_gpu_exporter: null
default_gpu_exporter:
container_image: ghcr.io#semianalysisai/amd-device-metrics-exporter@sha256:8a3fe70b8a848ca10a7fd90862d1d9e41e9c7b6314669a2e8340c646e6dae15c
port: 19500
command: env AMD_GPU_GET_CACHE_TTL=0s /home/amd/tools/entrypoint.sh
kind: custom
gpu_labels:
index: gpu_id
identity: serial_number
gpu_metrics:
power:
metric: gpu_power_usage
scope: gpu_device_power_as_reported_by_amd_device_metrics_exporter
gpu_util:
metric: gpu_gfx_activity
temperature:
metric: gpu_junction_temperature
nginx_raise_ulimit: false
4 changes: 3 additions & 1 deletion inferencex-e2e/docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -265,7 +265,9 @@ The current processing paths share these helpers:
- [`Parallelism`](../infx/results/topology.py) shares GPU-count calculation, parallelism result fields, and normalization when there are no separate decode GPUs. Fixed-sequence results retain explicit allocation counts; AgentX derives counts from its workers. Each caller retains its environment defaults, validation order, errors, and throughput denominators.
- [`with_power_metrics`](../infx/results/power/__init__.py) returns a copy with the supplied metric family replaced, removes stale validity reasons, and validates and rounds new metrics. Callers supply metric keys and schema version, then own artifact writes and validation sidecars. This allows another metric family to reuse the transformation without changing its implementation.

The power telemetry engine also lives in [`infx.results.power`](../infx/results/power): `multinode.run` validates srt-slurm artifact packages. Its benchmark-window parsing, per-device integration, aggregate replacement, and audit serialization live in `common.py`. Fixed-sequence and AgentX adapters import the engine directly; new result formats can supply their benchmark window and token counts to it. Single-node results carry no power until srt-slurm telemetry covers those lanes.
The power telemetry engine also lives in [`infx.results.power`](../infx/results/power): `multinode.run` validates srt-slurm artifact packages. Its benchmark-window parsing, per-device integration, aggregate replacement, and audit serialization live in `common.py`. Fixed-sequence and AgentX adapters import the engine directly; new result formats can supply their benchmark window and token counts to it. Native srt-slurm telemetry also covers single-node throughput and AgentX lanes.

SRT fixed-sequence and AgentX clients use native srt-slurm power sampling; they no longer launch a local NVIDIA/AMD SMI sampler. Single-node fixed-sequence execution requires `SRT_MEASUREMENT_WINDOW_DIR` before running throughput, then writes the completed window from the benchmark result. AgentX marks its native window regardless of node count; a missing contract records invalid power under the existing best-effort/`REQUIRE_POWER` policy. The launcher must enable native telemetry, retain its power package and producer identity, and finalize AgentX power after collection. Fixed-sequence processing selects that package through `POWER_ARTIFACT_DIR`, including single-node jobs.

The `infx` package runs with no installation step or new runtime dependency. Run the engine with `python -m infx.results.power.multinode` from `inferencex-e2e/`.

Expand Down
2 changes: 1 addition & 1 deletion inferencex-e2e/docs/eval-agentx-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -257,7 +257,7 @@ For each concurrency retain:
- server/frontend logs and every metrics endpoint represented.
- run URL/ID, attempt, head SHA, recipe/config identity, image, topology, fast flag, and any override.

The runner writes the command before replay and validates raw results after aggregation ([execution path](../infx/bench/agentic/run.py#L122-L197)). Aggregation preserves dataset provenance and hardware/model/topology fields ([aggregate construction](../infx/results/agentic/__init__.py)). Raw workflow uploads intentionally omit very large `inputs.json` and `profile_export_raw.jsonl`. If those are required for an investigation, preserve them from the live allocation before cleanup ([single-node artifact contract](../../.github/workflows/benchmark-tmpl.yml#L382-L391), [multi-node contract](../../.github/workflows/benchmark-multinode-tmpl.yml#L466-L475)).
The runner writes the command before replay and validates raw results after aggregation ([execution path](../infx/bench/agentic/run.py#L135-L218)). Aggregation preserves dataset provenance and hardware/model/topology fields ([aggregate construction](../infx/results/agentic/__init__.py)). Raw workflow uploads intentionally omit very large `inputs.json` and `profile_export_raw.jsonl`. If those are required for an investigation, preserve them from the live allocation before cleanup ([single-node artifact contract](../../.github/workflows/benchmark-tmpl.yml#L382-L391), [multi-node contract](../../.github/workflows/benchmark-multinode-tmpl.yml#L466-L475)).

## 9. Debug long AgentX runs from live evidence

Expand Down
Loading
Loading