Repository navigation
docs(e2e-deploy): fix the guidellm load test for the v0.7+ CLI - #485
Merged
Merged
Conversation
The guidellm Job used `guidellm benchmark run` with flat flags, the CLI of guidellm v0.6. guidellm v0.7.0 (2026-06-29) replaced it with `guidellm run` and registry-backed `kind=...` options, so with the rolling `latest` tag the job exits at once with "No such command 'benchmark'" and the saturation test runs with no load. Pin the image to v0.8.0 and express the same benchmark in the new syntax: openai_http backend against the gateway with /v1/completions, constant 50 req/s, a 3600 s duration constraint, and synthetic 256/512-token data. The tokenizer now resolves from the backend model, so `--processor` and the USER variable go away. The guide waits for the job pod to be Ready before sampling, since the first run pulls the image and the tokenizer, and notes why the manifest must not go back to `latest`. Verified against `guidellm mock-server` from the same image: backend validated, constant rate held, 2243 requests served, JSON report written. Signed-off-by: Shimi Bandiel <shimib@google.com>
shimib
requested review from
RishabhSaini,
ahg-g,
evacchi and
jtechapps
as code owners
October 6, 2026 18:20
jtechapps
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
The guidellm Job used
guidellm benchmark runwith flat flags, the v0.6 CLI. guidellm v0.7.0 (June 2026) replaced it withguidellm runandkind=...options, so with thelatesttag the job exits immediately withNo such command 'benchmark'and Option B's saturation test runs with no load.latestcurrently resolves to v0.7.4.Pins the image to v0.8.0 and expresses the same benchmark in the new syntax (openai_http backend on
/v1/completions, constant 50 req/s, 3600 s duration constraint, synthetic 256/512 tokens). The tokenizer now resolves from the backend model, so--processorandUSERgo away. The guide waits for the job pod to be Ready before sampling and says why the manifest must not go back tolatest.How was this tested?
Ran the Job against
guidellm mock-serverfrom the same image: backend validated, constant rate held at about 41 req/s against the 50 target, 2,243 requests served, JSON report written.Release note