Repository navigation
[PowerX] tidy the #3781 power consumers and trim their driver test / [PowerX] 整理 #3781 功耗消费端并精简其 driver 测试 - #3798
edwingao28 wants to merge 3 commits into
Conversation
There was a problem hiding this comment.
LGTM — this is a pure refactor with no behavior change, and I verified the two logic extractions preserve original semantics.
- Verified
_power_topologyin power_adapter.py returns outputs identical to the inlined original logic for both the single-node (expected_num_gpusset) and multinode (disagg/aggregate) branches, including the invalid-topology paths that route into the existingreasonsfailure flow. - Verified
SAMPLES_HEADERSdict consolidation in multinode.py is a mechanical equivalent of the prior three-way if/elif and inline-dict-literal lookups. - Confirmed the dropped
mi325x-amd/mi300x-amdparametrize cases in test_srt_driver.py are justified: both used the samecustomexporter kind/port as the retainedmi355x-amdscase (so added no new exporter-kind coverage to that test), andtest_srt_config.pystill parametrizes over all three AMD clusters for config-rendering coverage. - The remaining two test files are mechanical rename-only updates (
run_multinode_agentic_power→run_native_agentic_power) with no change to test bodies.
Extended reasoning...
The diff is a readability-only refactor across two source files (power_adapter.py, multinode.py) and three test files, touching power-telemetry validation logic with no auth/crypto/data-exposure surface. I manually traced both extracted helpers (_power_topology and SAMPLES_HEADERS) against the original inline logic line-by-line and confirmed behavioral equivalence, and checked that the trimmed test_srt_driver.py parametrize cases are still covered by test_srt_config.py as the author claims. No findings were reported by the bug hunt, the change is small and mechanical, and there are no outstanding reviewer objections in the timeline.
7899c87 to
fe41eaf
Compare
|
Generally most InferenceX submissions don't need to add changes to the infx package. On the off chance they do, generally the diff is small. Please prompt your agent to clean your PR to keep your diff to only necessary & minimal changes too. We have an increased code quality bar for changes to the infx package. 中文大多数 InferenceX 提交通常不需要修改 infx 包。即使确实需要,改动通常也很小。请让你的 agent 清理 PR,使改动只包含必要且最小的内容。infx 包的改动有更高的代码质量要求。 |
中文:samples CSV 的版本识别与行校验改为共用同一张表头表。
The entry point serves single-node points too, and the GPU-topology decision now lives in one function instead of two interleaved branches. 中文:入口函数同时服务单节点测试点,故按原生功耗包命名;GPU 拓扑判断集中到一个函数,不再分散在两段交错分支中。
中文:单节点原生功耗矩阵测试缩减为每种 exporter 一个集群;三个 AMD 集群渲染一致由 test_srt_config 覆盖。
fe41eaf to
c84aec3
Compare
Summary
Stacked on #3781 with base
feat/powerx-amd-upstream-572; merge into that branch, notmain.Three readability changes, no behavior change. Samples CSV version detection and row length derive from one
SAMPLES_HEADERStable, so a future v4 is one entry. The AgentX adapter is named for native packages,run_native_agentic_power, with its GPU-topology decision in_power_topologyinstead of two interleaved branches. The single-node driver test runs one cluster per exporter kind;test_srt_configalready covers the three AMD records.Testing: full CI pytest command and ruff pass locally.
AI model disclosure
claude-fable-5-1(Claude Code)PR checklist and change type
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesfull-sweep-fail-fast(recommended),full-sweep-enabled, ornon-canary-full-sweep-enabled. Optional modifiersall-evals,evals-only, andagentx-fastrequire a primary label; the last two block reuse while applied.OWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.中文
基于 #3781,base 为
feat/powerx-amd-upstream-572;请合入该分支,而非main。三项可读性整理,不改变行为。samples CSV 的版本识别与行长度统一来自
SAMPLES_HEADERS一张表,日后新增 v4 只需加一项。AgentX adapter 按原生功耗包命名为run_native_agentic_power,GPU 拓扑判断集中在_power_topology,不再分散在两段交错分支。单节点 driver 测试每种 exporter 只跑一个集群;三个 AMD 集群已由test_srt_config覆盖。测试: 本地完整 CI pytest 命令与 ruff 均通过。
AI 模型:
claude-fable-5-1(Claude Code),负责实现与验证。