No description
  • Shell 72.1%
  • Jinja 27.9%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
agent a83e5601a8 Add llama.cpp health wait and tok/sec wrapper
Wait on /health like vLLM and SGLang. Expose cache-ram, fit,
seccomp, and memlock for the Qwen GGUF overlay. Benchmark is
the shared projects/llm-bench client.

Agent: grok
2026-08-17 15:50:30 +00:00
defaults Add llama.cpp health wait and tok/sec wrapper 2026-08-17 15:50:30 +00:00
files Add llama.cpp health wait and tok/sec wrapper 2026-08-17 15:50:30 +00:00
meta Add rootless podman deployment role for llama.cpp server 2026-06-16 03:43:23 +00:00
tasks Add llama.cpp health wait and tok/sec wrapper 2026-08-17 15:50:30 +00:00
templates Add rootless podman deployment role for llama.cpp server 2026-06-16 03:43:23 +00:00
.gitignore Add rootless podman deployment role for llama.cpp server 2026-06-16 03:43:23 +00:00
LICENSE Add rootless podman deployment role for llama.cpp server 2026-06-16 03:43:23 +00:00
README.md Add llama.cpp health wait and tok/sec wrapper 2026-08-17 15:50:30 +00:00

ansible-roles-llamacpp

This role deploys a rootless podman based llama.cpp server with OpenAI-compatible API support.

Task Configuration

- name: Setup llama.cpp
  hosts: somehost
  become: true
  roles:
    - role: llamacpp
      llamacpp_user: llm
      llamacpp_model: my-model.Q4_K_M.gguf
      llamacpp_listen: 0.0.0.0:8080
      llamacpp_image_tag: server-cuda
      llamacpp_gpus: all
      llamacpp_cuda_visible_devices: "0"

Host paths under the service user (e.g. ~/.config/llamacpp/):

Directory Purpose
models/ Local GGUF files (-m)
cache/ Hugging Face downloads (-hf, via LLAMA_CACHE)
prompts/ Optional --system-prompt-file content

Place files under files/<hostname>/llamacpp/ (or files/llamacpp/) and redeploy to sync them.

Deployment and Removal

systemctl --user stop container-llamacpp.service

Deploy

./run.sh deploy actual --tags llamacpp --limit somehost

Remove

./run.sh deploy actual --tags llamacpp --extra-vars "deployment_state=absent" --limit somehost

Benchmark

files/llamacpp-benchmark.sh is a wrapper around projects/llm-bench/openai-chat-bench.sh. llama.cpp on Amboy is :11435. Use --strict-openai if the server rejects cache_salt / ignore_eos.

./files/llamacpp-benchmark.sh \
  --url http://somehost:11435/v1 \
  --reasoning off \
  --runs 3 \
  --strict-openai