No description
- Shell 72.1%
- Jinja 27.9%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
Wait on /health like vLLM and SGLang. Expose cache-ram, fit, seccomp, and memlock for the Qwen GGUF overlay. Benchmark is the shared projects/llm-bench client. Agent: grok |
||
| defaults | ||
| files | ||
| meta | ||
| tasks | ||
| templates | ||
| .gitignore | ||
| LICENSE | ||
| README.md | ||
ansible-roles-llamacpp
This role deploys a rootless podman based llama.cpp server with OpenAI-compatible API support.
Task Configuration
- name: Setup llama.cpp
hosts: somehost
become: true
roles:
- role: llamacpp
llamacpp_user: llm
llamacpp_model: my-model.Q4_K_M.gguf
llamacpp_listen: 0.0.0.0:8080
llamacpp_image_tag: server-cuda
llamacpp_gpus: all
llamacpp_cuda_visible_devices: "0"
Host paths under the service user (e.g. ~/.config/llamacpp/):
| Directory | Purpose |
|---|---|
models/ |
Local GGUF files (-m) |
cache/ |
Hugging Face downloads (-hf, via LLAMA_CACHE) |
prompts/ |
Optional --system-prompt-file content |
Place files under files/<hostname>/llamacpp/ (or files/llamacpp/) and redeploy to sync them.
Deployment and Removal
systemctl --user stop container-llamacpp.service
Deploy
./run.sh deploy actual --tags llamacpp --limit somehost
Remove
./run.sh deploy actual --tags llamacpp --extra-vars "deployment_state=absent" --limit somehost
Benchmark
files/llamacpp-benchmark.sh is a wrapper around
projects/llm-bench/openai-chat-bench.sh. llama.cpp on Amboy is :11435.
Use --strict-openai if the server rejects cache_salt / ignore_eos.
./files/llamacpp-benchmark.sh \
--url http://somehost:11435/v1 \
--reasoning off \
--runs 3 \
--strict-openai