No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-16 02:28:02 +00:00
defaults Add vllm_extra_args passthrough to the serve command 2026-09-16 02:28:02 +00:00
files Point vLLM tok/sec bench at projects/llm-bench 2026-08-17 15:50:30 +00:00
meta Initialize vLLM deployment role 2026-08-14 17:43:12 +00:00
tasks Persist caches and verify vLLM readiness 2026-08-14 18:27:16 +00:00
.gitignore Initialize vLLM deployment role 2026-08-14 17:43:12 +00:00
LICENSE Initialize vLLM deployment role 2026-08-14 17:43:12 +00:00
README.md Point vLLM tok/sec bench at projects/llm-bench 2026-08-17 15:50:30 +00:00

ansible-roles-vllm

Deploys vLLM OpenAI-compatible inference server as a rootless Podman container.

Task Configuration

- name: Setup vLLM
  hosts: somehost
  become: true
  roles:
    - role: vllm
      vllm_model: meta-llama/Llama-3.2-1B
      vllm_hf_token: "{{ vault_vllm_hf_token }}"

Deployment and Removal

systemctl --user stop container-vllm.service

Deploy

./run.sh deploy actual --tags vllm --limit somehost

Remove

./run.sh deploy actual --tags vllm --extra-vars "deployment_state=absent" --limit somehost

Benchmark

files/vllm-benchmark.sh is a wrapper around projects/llm-bench/openai-chat-bench.sh (set LLM_BENCH if the workspace copy is not found). It measures client-observed time to first token and an approximate sequential decode rate.

./files/vllm-benchmark.sh \
  --url http://somehost:8000/v1 \
  --model organization/model \
  --reasoning off \
  --runs 3 \
  --ignore-eos

The reported decode rate is (output tokens - 1) / (last chunk - first chunk). SSE chunks are not token-aligned, so this is a client-observed approximation, not engine token timing. Runs are sequential; they measure single-request generation speed, not concurrent serving throughput. Each request gets a unique cache salt unless --reuse-prefix-cache is set. Use --ignore-eos for comparable synthetic fixed-output runs; omit it when measuring normal end-to-end behavior.