Skip to content

docs(e2e-deploy): fix the guidellm load test for the v0.7+ CLI - #485

Merged
jtechapps merged 1 commit into
llm-d:mainfrom
shimib:fix/e2e-guidellm-cli
Oct 6, 2026
Merged

jtechapps merged 1 commit into
llm-d:mainfrom
shimib:fix/e2e-guidellm-cli

Conversation

@shimib

@shimib shimib commented Oct 6, 2026

Copy link
Copy Markdown
Member

What does this PR do?

The guidellm Job used guidellm benchmark run with flat flags, the v0.6 CLI. guidellm v0.7.0 (June 2026) replaced it with guidellm run and kind=... options, so with the latest tag the job exits immediately with No such command 'benchmark' and Option B's saturation test runs with no load. latest currently resolves to v0.7.4.

Pins the image to v0.8.0 and expresses the same benchmark in the new syntax (openai_http backend on /v1/completions, constant 50 req/s, 3600 s duration constraint, synthetic 256/512 tokens). The tokenizer now resolves from the backend model, so --processor and USER go away. The guide waits for the job pod to be Ready before sampling and says why the manifest must not go back to latest.

How was this tested?

  • Manual testing performed

Ran the Job against guidellm mock-server from the same image: backend validated, constant rate held at about 41 req/s against the 50 target, 2,243 requests served, JSON report written.

Release note

NONE

The guidellm Job used `guidellm benchmark run` with flat flags, the CLI of
guidellm v0.6. guidellm v0.7.0 (2026-06-29) replaced it with `guidellm run`
and registry-backed `kind=...` options, so with the rolling `latest` tag the
job exits at once with "No such command 'benchmark'" and the saturation
test runs with no load.

Pin the image to v0.8.0 and express the same benchmark in the new syntax:
openai_http backend against the gateway with /v1/completions, constant
50 req/s, a 3600 s duration constraint, and synthetic 256/512-token data.
The tokenizer now resolves from the backend model, so `--processor` and
the USER variable go away. The guide waits for the job pod to be Ready
before sampling, since the first run pulls the image and the tokenizer,
and notes why the manifest must not go back to `latest`.

Verified against `guidellm mock-server` from the same image: backend
validated, constant rate held, 2243 requests served, JSON report written.

Signed-off-by: Shimi Bandiel <shimib@google.com>
@jtechapps
jtechapps merged commit 2ac021f into llm-d:main Oct 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants