• Shell 72.3%
  • Jinja 27.7%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
agent 71dac0f285 Point SGLang tok/sec bench at projects/llm-bench
The previous files/sglang-benchmark.sh was a copy of the vLLM
client. The wrapper execs the shared script.

Agent: grok
2026-08-17 15:50:30 +00:00
defaults Pass the SGLang launch command without a custom entrypoint 2026-08-16 18:33:02 +00:00
files Point SGLang tok/sec bench at projects/llm-bench 2026-08-17 15:50:30 +00:00
meta Add SGLang deployment role 2026-08-03 19:57:44 +00:00
tasks Pass the SGLang launch command without a custom entrypoint 2026-08-16 18:33:02 +00:00
templates Add SGLang deployment role 2026-08-03 19:57:44 +00:00
.gitignore Add SGLang deployment role 2026-08-03 19:57:44 +00:00
LICENSE Add SGLang deployment role 2026-08-03 19:57:44 +00:00
README.md Point SGLang tok/sec bench at projects/llm-bench 2026-08-17 15:50:30 +00:00

ansible-roles-sglang

Deploys SGLang as a rootless Podman OpenAI-compatible inference server.

Task Configuration

- name: Setup SGLang
  hosts: somehost
  become: true
  roles:
    - role: sglang
      sglang_model: unsloth/Qwen3.8-27B-NVFP4
      sglang_hf_token: "{{ vault_sglang_hf_token }}"

Podman and NVIDIA CDI must be configured on the managed host. A recent NVIDIA Container Toolkit creates the CDI specification automatically; otherwise, generate it with:

sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
nvidia-ctk cdi list

sglang_nvidia_start_service installs a small host oneshot unit by default so NVIDIA device nodes exist before rootless services start after boot.

Deployment and Removal

systemctl --user stop container-sglang.service

Deploy

./run.sh deploy actual --tags sglang --limit somehost

Remove

./run.sh deploy actual --tags sglang --extra-vars "deployment_state=absent" --limit somehost

Benchmark

files/sglang-benchmark.sh is a wrapper around projects/llm-bench/openai-chat-bench.sh (set LLM_BENCH if the workspace copy is not found). It measures client-observed time to first token and an approximate sequential decode rate.

./files/sglang-benchmark.sh \
  --url http://somehost:8000/v1 \
  --model organization/model \
  --reasoning both \
  --runs 3 \
  --ignore-eos

The reported decode rate is (output tokens - 1) / (last chunk - first chunk). SSE chunks are not token-aligned, so this is a client-observed approximation, not engine token timing. Runs are sequential; they measure single-request generation speed, not concurrent serving throughput. Each request gets a unique cache salt unless --reuse-prefix-cache is set. Use --ignore-eos for comparable synthetic fixed-output runs; omit it when measuring normal end-to-end behavior.