- Jinja 84.8%
- Shell 14.4%
- Python 0.8%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| defaults | ||
| files | ||
| filter_plugins | ||
| meta | ||
| tasks | ||
| templates | ||
| vars | ||
| .gitignore | ||
| LICENSE | ||
| README.md | ||
ansible-role-k8s
Ansible role for configuring k3s and rke2 kubernetes clusters
- https://docs.k3s.io/
- https://docs.rke2.io/
- https://kube-vip.io/
- https://github.com/sbstp/kubie
- https://kubernetes.io/docs/tasks/tools/
- https://helm.sh/
Requirements
There is an included helper script to install common tools scripts/get-kube-tools.sh
yqrequired on the local system for the kubectl formatting task which places an updated kubeconfig in the local ~/.kubekubectlrequired on the local system for basic cluster mangement and application of locally stored manifests or secretshelmrequired on the local system for helm deployments that use locally stored value files, otherwise this is handled on the bootstrap nodekubierecommened on the local system for context management after deployment
Setup
There is a helper script scripts/token-vault.sh which pre-generates a cluster token and places it in an encrypted vault file
Cluster Example
cluster hosts
[k8s_somecluster]
somecluster_control k8s_deploy_type=bootstrap
somecluster_agent_smith k8s_deploy_type=agent k8s_external_ip=x.x.x.x
somecluster_agent_jones k8s_deploy_type=agent k8s_external_ip=x.x.x.x
cluster tasks
- name: Setup k8s server node
hosts: somehost
become: true
roles:
- role: k8s
k8s_type: rke2
k8s_cluster_name: somecluster
k8s_cluster_url: somecluster.somewhere
k8s_cni_interface: enp1s0
k8s_selinux: true
- role: firewalld
firewalld_add:
- name: internal
interfaces:
- enp1s0
masquerade: true
forward: true
interfaces:
- enp1s0
services:
- dhcpv6-client
- ssh
- http
- https
ports:
- 6443/tcp # kubernetes API
- 9345/tcp # supervisor API
- 10250/tcp # kubelet metrics
- 2379/tcp # etcd client
- 2380/tcp # etcd peer
- 30000-32767/tcp # NodePort range
- 8472/udp # canal/flannel vxlan
- 9099/tcp # canal health checks
- name: trusted
sources:
- 10.42.0.0/16
- 10.43.0.0/16
- name: public
masquerade: true
forward: true
interfaces:
- enp7s0
services:
- http
- https
firewalld_remove:
- name: public
interfaces:
- enp1s0
services:
- dhcpv6-client
- ssh
Retrieve kube config from an existing cluster
This task will retrieve and format the kubectl config for an existing cluster, this runs automatically during cluster creation.
k8s_cluster_name sets the cluster context
k8s_cluster_url sets the server address
The working copy is ~/.kube/config-<cluster>.yaml so helm and kubectl can read it. When the playbook wrapper is present (run.sh at the playbook root, including two levels above inventory_dir), that file is also piped into ansible-vault encrypt at {{ inventory_dir }}/files/k8s/<cluster>/kubeconfig.yaml. Set k8s_kubeconfig_vault: false to skip the vaulted copy.
ansible-playbook -i prod/ site.yml --tags=k8s-get-kubeconf --limit=k8s_somecluster
Basic Cluster Interaction
kubie ctx <cluster-name>
kubectl get node -o wide
kubectl get pods,svc,ds --all-namespaces
Deployment and Removal
Deploy
ansible-playbook -i hosts site.yml --tags=firewalld,k8s --limit=k8s_somecluster
Adding a node, simply add the new host to the cluster group with its defined role and deploy
ansible-playbook -i hosts site.yml --tags=firewalld,k8s --limit=just_the_new_host
Remove firewall role
ansible-playbook -i hosts site.yml --tags=firewalld,k8s --extra-vars "firewall_action=remove" --limit=somehost
There is a task to completely destroy an existing cluster, this will ask for interactive user confirmation and should be used with caution.
ansible-playbook -i prod/ site.yml --tags=k8s --extra-vars 'k8s_action=destroy' --limit=some_innocent_cluster
Set k8s_allow_destruction: true in a cluster's inventory to skip the prompt for that cluster. To skip once without changing inventory, pass --extra-vars 'k8s_action=destroy k8s_destroy_confirm=true'.
Manual removal commands
/usr/local/bin/k3s-uninstall.sh
/usr/local/bin/k3s-agent-uninstall.sh
/usr/local/bin/rke2-uninstall.sh
/usr/local/bin/rke2-agent-uninstall.sh
Managing K3S Services
servers
systemctl status k3s.service
journalctl -u k3s.service -f
agents
systemctl status k3s-agent.service
journalctl -u k3s-agent -f
uninstall servers
/usr/local/bin/k3s-uninstall.sh
uninstall agents
/usr/local/bin/k3s-agent-uninstall.sh
Managing RKE2 Services
servers
systemctl status rke2-server.service
journalctl -u rke2-server -f
agents
systemctl status rke2-agent.service
journalctl -u rke2-agent -f
uninstall servers
/usr/bin/rke2-uninstall.sh
uninstall agents
/usr/local/bin/rke2-uninstall.sh
RKE2 stop vs kube-vip failover
rke2-server.service uses KillMode=process. systemctl stop rke2-server is
not a full node-down. That is useful as a half-stopped test, and the wrong
tool if the goal is VIP failover.
Half-stopped (systemctl stop rke2-server on the kube-vip lease holder):
- the main
rke2process dies, so supervisor:9345is gone - containerd-shims are reparented to PID 1, so kube-vip and kube-apiserver often keep running
- kube-vip keeps renewing
plndr-cp-lockand keeps the VIP on that node - VIP
:6443may still work; VIP:9345is RST - other control-plane nodes do not get the VIP; kube-vip does not probe
:9345
The VIP is a real /32 on the holder's NIC (ip addr), not a purely virtual
ARP advertisement like MetalLB or Cilium L2. Kish/uruk cannot remotely delete
that address. Even if they took the lease, the half-stopped holder would still
own .198 on the wire until kube-vip on that node drops it (clean exit) or
something on the holder removes it (rke2-killall.sh, ip addr del).
Use this on a disposable cluster when testing "node is only partly dead" (clients still reach the API, workers cannot reach the supervisor).
# on the current kube-vip holder
systemctl stop rke2-server
ss -lntp | grep -E '6443|9345'
ip -4 addr show | grep <vip>
kubectl -n kube-system get lease plndr-cp-lock
# recover
systemctl start rke2-server
Full stop / real failover: /usr/bin/rke2-killall.sh (stops the unit and
kills the shims, including kube-vip). After that the lease should move and
VIP :9345 should answer on the new holder. Then systemctl start rke2-server
to bring the node back. Do not use uninstall for a failover test.
override default cannal options
# /var/lib/rancher/rke2/server/manifests/rke2-canal-config.yaml
---
apiVersion: helm.cattle.io/v1
kind: HelmChartConfig
metadata:
name: rke2-canal
namespace: kube-system
spec:
valuesContent: |-
flannel:
iface: "eth1"
Enable flannels wireguard support under canal
kubectl rollout restart ds rke2-canal -n kube-system
# /var/lib/rancher/rke2/server/manifests/rke2-canal-config.yaml
---
apiVersion: helm.cattle.io/v1
kind: HelmChartConfig
metadata:
name: rke2-canal
namespace: kube-system
spec:
valuesContent: |-
flannel:
backend: "wireguard"