No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-26 18:26:27 +00:00
defaults agent fixes 2026-09-26 18:26:27 +00:00
meta update template, add heartbeat 2026-04-05 18:37:05 +00:00
tasks support inventory backlog keys 2026-05-19 20:48:32 +00:00
templates agent fixes 2026-09-26 18:26:27 +00:00
.gitignore initial commit 2025-10-06 22:03:16 +00:00
AGENTS.md update contributor workflow guidance 2026-05-19 20:11:50 +00:00
LICENSE initial commit 2025-10-06 22:03:16 +00:00
README.md agent fixes 2026-09-26 18:26:27 +00:00

ansible-roles-backlog

An Ansible role for installing per-user ZFS backup scripts, SSH configuration, ZFS permissions, and cron entries.

Task Configuration

backlog_key files are expected under files/backlog/ in the inventory directory, the playbook directory, or this role. Inventory-local keys are preferred, for example deploy/beta/files/backlog/<key-name>. Each backlog_targets item configures one local user and dataset.

- name: Setup backlog
  hosts: somehost
  become: true
  roles:
    - role: backlog
      backlog_key: backlog-somekey-file
      backlog_targets:
        - user: someservice
          services:
            - name: pod-someservice
              type: systemd
              scope: user
          dataset: storage/DATA/home/someservice
          retention: 3
          time: "16 0 * * 0"
          destinations:
            - name: somehost-backlog
              host: somehost.sh
              data: batterywharf/DATA/backlog/someservice
              heartbeat: https://example.com/ping/somehost-backlog

Target Options

  • user: local user that owns the backup script and cron entry.
  • dataset: local ZFS dataset to snapshot and send.
  • services: services to stop before snapshot creation and restart afterward.
  • time: five-field cron expression. If omitted, backlog_cron_time is used.
  • destinations: remote backup targets. name is the SSH host alias generated in ~/.ssh/config; host is the real hostname; data is the remote dataset. Optional heartbeat sends a curl ping after that destination receives a full or incremental snapshot successfully.
  • template: optional script template name without .sh.j2. The role resolves this as templates/<template>.sh.j2.
  • retry_limit / retry_delay: per-destination cool-down retry within one run. A destination that is merely unreachable is retried after retry_delay seconds rather than failing the run outright.
  • max_runtime: soft ceiling for a single run. It bounds when a new retry starts and caps the cool-down sleep; a transfer already in flight can exceed it.
  • send_intermediates: true sends every intermediate snapshot (zfs send -I) so a destination keeps the same restore points as the source; false (default) replicates latest state only (-i), which is cheaper but leaves gaps on a destination that missed a run. Choose per the invariant you want.

Running the script

backlog.sh [run|send|prune] [snapshot-suffix]

run (the default, and what cron invokes) snapshots and sends. send retries the newest existing snapshot without creating another, which is how a failed send is retried the same day. Runs take an exclusive lock, so a long cool-down retry cannot overlap the next scheduled run.

Destination states

The script distinguishes three remote states, because only one of them may safely escalate to a full send:

state action
dataset absent full send (first seed)
dataset present, no matching snapshots refuse, report permanent failure
listing failed (transport, auth, remote error) send nothing, retry later

An interrupted zfs receive -s leaves a resume token. Before choosing a base the script resumes from it; it only runs zfs receive -A — which discards the partial receive — after zfs send -nvt reports the stream definitively unresumable. This requires zfs get receive_resume_token on the destination, which zfshell.sh permits in exactly that form.

Holds and retention

New snapshots are held under backlog_zfs_hold. Holds are released, and retention applied, only after every destination succeeds. A snapshot that is still some destination's incremental base is never pruned — including for a destination that was unreachable during the run, whose last confirmed snapshot is recorded in backlog_state_file.

Template Versions

templates/backlog.sh.j2 is the default template used by the role. The current template selector resolves target.template as templates/<template>.sh.j2, so the historical templates/backlog.sh-v*.j2 files are archival unless renamed or the task lookup is changed. Prefer updating the default template for new behavior and remove old versioned templates once no inventory references them.

Deployment and Removal

Deploy

ansible-playbook -i hosts site.yml --tags=backlog --limit=somehost