Skip to content

Run an end-to-end gfx1151 serving qualification sweep #21

Description

@randomvariable

Objective

Run a reproducible end-to-end serving qualification sweep for the fork's gfx1151 path after runtime blockers and image provenance are resolved.

Related fork issues

Test matrix

  • Context-length scalability through the selected model's supported range.
  • Deterministic quality probes plus a named evaluation suite where applicable.
  • Cold/warm TTFT, prefill throughput, decode throughput, and concurrency scaling.
  • Peak unified memory, graph-reserved memory, KV-cache capacity, and host-memory headroom.
  • Backend-selection evidence proving intended gfx1151 AITER/Triton/HIP paths run without CPU fallback.

Record exact model revision, quantization, ROCm/Torch/vLLM revisions, launch settings, prompt/output lengths, repetitions, and summary statistic.

Acceptance criteria

  • All blocking correctness and crash issues are linked to minimized reproducers.
  • Results are captured in a machine-readable artifact and concise comparison table.
  • No hidden CPU fallback or unsupported FP8/FP4/MLA path is presented as acceleration.
  • Any serving migration proposal names its reference backend and requires API/quality compatibility plus equal-or-better latency/throughput on the agreed workload.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/ideaOptimization idea candidate for evaluation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions