Skip to main content

benchmark-vllm-gpu

vLLM GPU serving benchmark: vllm serve with the largest valid tensor-parallel size for each model (not always all visible GPUs — see Tensor parallelism), then GuideLLM sweep (default 3 steps) per workload.

ArchBase
amd64 + arm64vllm/vllm-openai:v{VLLM_VERSION} (multi-arch Hub image)

On version bumps, confirm the pinned tag lists both linux/amd64 and linux/arm64 on Docker Hub.

Published as ghcr.io/sparecores/benchmark-vllm-gpu:main. Pins: VLLM_VERSION, GUIDELLM_VERSION. Harness: benchmark.py.

Models (default ladder)​

SmolLM2-135M, Qwen2.5-0.5B, Gemma-2-2B, Llama-3.1-8B, Phi-4, Llama-3.3-70B bnb-4bit (large VRAM). Gemma and Llama-3.1-8B are gated on Hugging Face — set HF_TOKEN and accept each model license.

Workloads​

  • chat: 256 prompt / 128 output tokens
  • rag: 1024 / 256
  • long: 4096 / 512 (GPU only)

Output​

JSONL lines (benchmark=vllm_serving) with TTFT/TPOT/ITL/E2EL percentiles and throughputs per GuideLLM strategy. See vllm-common/README.