benchmark-vllm-gpu
vLLM GPU serving benchmark: vllm serve with the largest valid tensor-parallel
size for each model (not always all visible GPUs — see
Tensor parallelism), then
GuideLLM sweep (default 3 steps) per workload.
| Arch | Base |
|---|---|
| amd64 + arm64 | vllm/vllm-openai:v{VLLM_VERSION} (multi-arch Hub image) |
On version bumps, confirm the pinned tag lists both linux/amd64 and linux/arm64 on Docker Hub.
Published as ghcr.io/sparecores/benchmark-vllm-gpu:main. Pins: VLLM_VERSION, GUIDELLM_VERSION. Harness: benchmark.py.
Models (default ladder)
SmolLM2-135M, Qwen2.5-0.5B, Gemma-2-2B, Llama-3.1-8B, Phi-4, Llama-3.3-70B bnb-4bit (large VRAM).
Gemma and Llama-3.1-8B are gated on Hugging Face — set HF_TOKEN and accept each model license.
Workloads
- chat: 256 prompt / 128 output tokens
- rag: 1024 / 256
- long: 4096 / 512 (GPU only)
Output
JSONL lines (benchmark=vllm_serving) with TTFT/TPOT/ITL/E2EL percentiles and throughputs per GuideLLM strategy. See vllm-common/README.