- Shell 72.3%
- Jinja 27.7%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
The previous files/sglang-benchmark.sh was a copy of the vLLM client. The wrapper execs the shared script. Agent: grok |
||
| defaults | ||
| files | ||
| meta | ||
| tasks | ||
| templates | ||
| .gitignore | ||
| LICENSE | ||
| README.md | ||
ansible-roles-sglang
Deploys SGLang as a rootless Podman OpenAI-compatible inference server.
Task Configuration
- name: Setup SGLang
hosts: somehost
become: true
roles:
- role: sglang
sglang_model: unsloth/Qwen3.8-27B-NVFP4
sglang_hf_token: "{{ vault_sglang_hf_token }}"
Podman and NVIDIA CDI must be configured on the managed host. A recent NVIDIA Container Toolkit creates the CDI specification automatically; otherwise, generate it with:
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
nvidia-ctk cdi list
sglang_nvidia_start_service installs a small host oneshot unit by default so
NVIDIA device nodes exist before rootless services start after boot.
Deployment and Removal
systemctl --user stop container-sglang.service
Deploy
./run.sh deploy actual --tags sglang --limit somehost
Remove
./run.sh deploy actual --tags sglang --extra-vars "deployment_state=absent" --limit somehost
Benchmark
files/sglang-benchmark.sh is a wrapper around
projects/llm-bench/openai-chat-bench.sh (set LLM_BENCH if the workspace
copy is not found). It measures client-observed time to first token and an
approximate sequential decode rate.
./files/sglang-benchmark.sh \
--url http://somehost:8000/v1 \
--model organization/model \
--reasoning both \
--runs 3 \
--ignore-eos
The reported decode rate is (output tokens - 1) / (last chunk - first chunk).
SSE chunks are not token-aligned, so this is a client-observed approximation,
not engine token timing. Runs are sequential; they measure single-request
generation speed, not concurrent serving throughput. Each request gets a unique
cache salt unless --reuse-prefix-cache is set. Use --ignore-eos for
comparable synthetic fixed-output runs; omit it when measuring normal end-to-end
behavior.