- Shell 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| defaults | ||
| files | ||
| meta | ||
| tasks | ||
| .gitignore | ||
| LICENSE | ||
| README.md | ||
ansible-roles-vllm
Deploys vLLM OpenAI-compatible inference server as a rootless Podman container.
- https://docs.vllm.ai/en/stable/deployment/docker.html
- https://hub.docker.com/r/vllm/vllm-openai/tags
Task Configuration
- name: Setup vLLM
hosts: somehost
become: true
roles:
- role: vllm
vllm_model: meta-llama/Llama-3.2-1B
vllm_hf_token: "{{ vault_vllm_hf_token }}"
Deployment and Removal
systemctl --user stop container-vllm.service
Deploy
./run.sh deploy actual --tags vllm --limit somehost
Remove
./run.sh deploy actual --tags vllm --extra-vars "deployment_state=absent" --limit somehost
Benchmark
files/vllm-benchmark.sh is a wrapper around
projects/llm-bench/openai-chat-bench.sh (set LLM_BENCH if the workspace
copy is not found). It measures client-observed time to first token and an
approximate sequential decode rate.
./files/vllm-benchmark.sh \
--url http://somehost:8000/v1 \
--model organization/model \
--reasoning off \
--runs 3 \
--ignore-eos
The reported decode rate is (output tokens - 1) / (last chunk - first chunk).
SSE chunks are not token-aligned, so this is a client-observed approximation,
not engine token timing. Runs are sequential; they measure single-request
generation speed, not concurrent serving throughput. Each request gets a unique
cache salt unless --reuse-prefix-cache is set. Use --ignore-eos for
comparable synthetic fixed-output runs; omit it when measuring normal end-to-end
behavior.