Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 11 additions & 2 deletions .github/workflows/litestream.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name: Headscale Litestream

# Runs tests/litestream/run.sh: the headscale_hosted provider with Litestream and a warm standby against an S3-compatible store, a real Tailscale client, promotion of the standby, and the headscale_in_cluster guard. Only on changes that can affect them.
# Runs tests/litestream/run.sh: the headscale_hosted provider with Litestream and a warm standby against an S3-compatible store, a real Tailscale client, promotion of the standby, and the headscale_in_cluster guard, once as is and once with the Headscale hosts pinning github_runner_cluster_docker_context while the current Docker context points at nothing. Only on changes that can affect them.
#
# Runs on GitHub-hosted ubuntu-latest, which has the Docker daemon the test needs. The work directory sits under runner.temp, which the job always has write access to.

Expand All @@ -9,6 +9,9 @@ on:
paths:
- roles/github_runner_cluster/tasks/mesh/headscale_*.yml
- roles/github_runner_cluster/tasks/headscale_promote.yml
- roles/github_runner_cluster/tasks/docker_context.yml
- plugins/filter/docker.py
- tests/lib/**
- roles/github_runner_cluster/templates/headscale/**
- roles/github_runner_cluster/files/headscale/**
- roles/github_runner_cluster/defaults/main.yml
Expand All @@ -26,9 +29,14 @@ permissions:

jobs:
litestream:
name: Litestream and standby promotion
name: Litestream and standby promotion${{ matrix.pin && ' (context pinned)' || '' }}
runs-on: ubuntu-latest
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
# Pinned: the Headscale hosts pin github_runner_cluster_docker_context while the current Docker context points at nothing (see tests/litestream/run.sh).
pin: [false, true]
env:
COMPOSE_VERSION: v5.5.1
steps:
Expand Down Expand Up @@ -67,6 +75,7 @@ jobs:
env:
GRTEST_WORK: ${{ runner.temp }}/grtest-ls
GRTEST_EXTRA_COLLECTIONS: ${{ runner.temp }}/collections
GRTEST_PIN_CONTEXT: ${{ matrix.pin && '1' || '0' }}
run: tests/litestream/run.sh

required-checks:
Expand Down
8 changes: 6 additions & 2 deletions .github/workflows/mesh-integration.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name: Mesh integration

# Brings up three k3s nodes in Docker and joins them through Headscale with the cluster role (tests/mesh/run.sh), once per scenario: headscale_hosted, headscale_in_cluster, and the headscale_existing policy check, plus the role's refusal of a Docker Compose release that recreates the containers it has just built. Only on changes that can affect how nodes join the mesh.
# Brings up three k3s nodes in Docker and joins them through Headscale with the cluster role (tests/mesh/run.sh), once per scenario: headscale_hosted, headscale_in_cluster, and the headscale_existing policy check, plus the role's refusal of a Docker Compose release that recreates the containers it has just built. headscale_in_cluster also runs with github_runner_cluster_docker_context pinned while the current Docker context points at nothing, which covers the cluster role's own docker calls and the recovery playbook's, and context_checks checks the role refuses a pin it cannot honour. Only on changes that can affect how nodes join the mesh.
#
# Runs on GitHub-hosted ubuntu-latest, whose Docker daemon runs the privileged containers the test needs (k3s, and Tailscale inside it). The work directory sits under runner.temp, which the job always has write access to.

Expand All @@ -13,6 +13,7 @@ on:
- k3s/**
- playbooks/recover_in_cluster_mesh.yml
- tests/mesh/**
- tests/lib/**
- .github/workflows/mesh-integration.yml
workflow_dispatch:

Expand All @@ -25,7 +26,7 @@ permissions:

jobs:
mesh:
name: Mesh (${{ matrix.scenario }}, Compose ${{ matrix.compose }})
name: Mesh (${{ matrix.scenario }}, Compose ${{ matrix.compose }}${{ matrix.pin && ', context pinned' || '' }})
runs-on: ubuntu-latest
timeout-minutes: 45
strategy:
Expand All @@ -38,6 +39,8 @@ jobs:
- {scenario: in_cluster, compose: v5.5.1}
- {scenario: hosted, compose: v2.39.0}
- {scenario: compose_faulty, compose: v2.38.2}
- {scenario: in_cluster, compose: v5.5.1, pin: true}
- {scenario: context_checks, compose: v5.5.1}
steps:
- uses: actions/checkout@v7

Expand Down Expand Up @@ -76,6 +79,7 @@ jobs:
env:
ANSIBLE_COLLECTIONS_PATH: ${{ runner.temp }}/collections
GRTEST_WORK: ${{ runner.temp }}/grtest-mesh
GRTEST_PIN_CONTEXT: ${{ matrix.pin && '1' || '0' }}
run: tests/mesh/run.sh "${{ matrix.scenario }}"

required-checks:
Expand Down
15 changes: 15 additions & 0 deletions playbooks/recover_in_cluster_mesh.yml
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@
msg: github_runner_cluster_mesh is {{ github_runner_cluster_mesh }}; this playbook only applies to headscale_in_cluster.
when: github_runner_cluster_mesh != 'headscale_in_cluster'

- name: Resolve the Docker context the recovery drives, the same one the cluster role uses on this host
# See the cluster role's docker_context.yml: the host's pinned github_runner_cluster_docker_context, or its current context. Every docker command below runs with it.
ansible.builtin.include_role:
name: exadev.github_runner.github_runner_cluster
tasks_from: docker_context

- name: Record where Headscale's files live
ansible.builtin.set_fact:
github_runner_cluster_recover_headscale_dir: "{{ github_runner_cluster_headscale_dir or github_runner_cluster_dir ~ '/headscale' }}"
Expand All @@ -36,11 +42,13 @@
- name: Restart the k3s container, which starts k3s off the mesh until Headscale answers
ansible.builtin.command:
argv: [docker, restart, "{{ github_runner_cluster_container_name }}"]
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: true

- name: Wait for the Headscale static pod to answer
ansible.builtin.command:
argv: "{{ github_runner_cluster_recover_probe }}"
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
failed_when: false
register: github_runner_cluster_recover_answers
Expand All @@ -54,6 +62,7 @@
- name: Find the k3s container's data volume, which holds Headscale's state
ansible.builtin.command:
argv: [docker, inspect, --format, "{{ '{{' }} range .Mounts {{ '}}' }}{{ '{{' }} if eq .Destination \"/var/lib/rancher/k3s\" {{ '}}' }}{{ '{{' }} .Name {{ '}}' }}{{ '{{' }} end {{ '}}' }}{{ '{{' }} end {{ '}}' }}", "{{ github_runner_cluster_container_name }}"]
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
register: github_runner_cluster_recover_volume

Expand All @@ -77,11 +86,13 @@
- "type=volume,src={{ github_runner_cluster_recover_volume.stdout | trim }},dst=/var/lib/headscale,volume-subpath=github-runner-headscale"
- "docker.io/headscale/headscale:v{{ github_runner_cluster_headscale_version }}"
- serve
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: true

- name: Wait for the temporary Headscale server to answer
ansible.builtin.command:
argv: "{{ github_runner_cluster_recover_probe }}"
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
failed_when: false
register: github_runner_cluster_recover_rescue_answers
Expand All @@ -97,6 +108,7 @@
- name: Wait for this node to rejoin the cluster on the mesh
ansible.builtin.command:
argv: [docker, exec, "{{ github_runner_cluster_container_name }}", kubectl, wait, --for=condition=Ready, "node/{{ github_runner_cluster_topology.hosts[inventory_hostname].node_name }}", --timeout=10s]
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
register: github_runner_cluster_recover_ready
until: github_runner_cluster_recover_ready.rc == 0
Expand All @@ -107,6 +119,7 @@
- name: Read the temporary Headscale server's log, for a recovery that did not bring the node back
ansible.builtin.command:
argv: [docker, logs, --tail, "40", "{{ github_runner_cluster_recover_rescue }}"]
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
failed_when: false
register: github_runner_cluster_recover_rescue_log
Expand All @@ -120,12 +133,14 @@
- name: Remove the temporary Headscale server, so the static pod can take its port back
ansible.builtin.command:
argv: [docker, rm, --force, "{{ github_runner_cluster_recover_rescue }}"]
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: true
failed_when: false

- name: Wait for the Headscale static pod to answer again
ansible.builtin.command:
argv: "{{ github_runner_cluster_recover_probe }}"
environment: "{{ github_runner_cluster_docker_cli_environment }}"
changed_when: false
register: github_runner_cluster_recover_final
until: github_runner_cluster_recover_final.rc == 0
Expand Down
32 changes: 32 additions & 0 deletions plugins/filter/docker.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
"""Filters the github_runner_cluster role uses to point the docker CLI at the same daemon as its community.docker module calls."""

from __future__ import annotations

from typing import Any

from ansible.errors import AnsibleFilterError


def docker_cli_environment(context: Any) -> dict[str, str]:
"""Return the environment that makes the docker CLI use a pinned Docker context.

DOCKER_CONTEXT selects the context for that one command, and for the Compose plugin it runs, without touching the host's current context. With no context pinned the environment is empty, so the CLI keeps following the host's current context or DOCKER_HOST.

:param context: the pinned context name, or an empty value for none. :returns: ``{"DOCKER_CONTEXT": context}``, or an empty mapping. :raises AnsibleFilterError: if the value is not a string.
"""
if context is None or context == "":
return {}
if not isinstance(context, str):
raise AnsibleFilterError(f"A Docker context name must be a string, not {type(context).__name__}: {context!r}.")
name = context.strip()
if name == "":
return {}
return {"DOCKER_CONTEXT": name}


class FilterModule:
"""Registers the Docker filters with Ansible."""

def filters(self) -> dict[str, Any]:
"""Return the filters this plugin provides."""
return {"docker_cli_environment": docker_cli_environment}
9 changes: 8 additions & 1 deletion roles/github_runner_cluster/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,12 @@ Which hosts are servers is decided automatically from the inventory group named

The role reads `docker compose version` first and fails, before changing anything on the host, if it is 2.37.1 or later but older than 2.39.0. Those releases create a container from an image they have just built without recording that image's ID on it ([docker/compose#13047](https://github.com/docker/compose/pull/13047)), so the next run sees a different image and recreates the container, which restarts every node on the run after the one that built the k3s image. Releases before 2.37.1 and from 2.39.0 on keep the containers. GitHub's ubuntu-latest runner image has shipped 2.38.2, so a CI job that runs the role there needs a different Compose installed first, as `.github/workflows/mesh-integration.yml` does.

## Docker contexts

By default the role drives whichever Docker daemon the host's current Docker context points at (`docker context show`), or the one `DOCKER_HOST` names when that is set. That is right for a host with one Docker engine, but the current context is shared, mutable state that other software changes. The case that matters: a Mac that runs its node in Colima and also has Docker Desktop installed. Launching Docker Desktop switches the current context to `desktop-linux` without asking, so the next run of the role would find no node in Docker Desktop and build a second, empty k3s node there, alongside the real one in Colima, instead of managing it.

Set `github_runner_cluster_docker_context` in the host's variables to pin the context by name (`colima` in that example) on any host with more than one engine. Every `community.docker` module call in the role then passes it as `cli_context`, and every `docker` command the role runs (Compose, `exec`, `inspect`, `volume`, image pulls, and the commands in `playbooks/recover_in_cluster_mesh.yml` and `playbooks/promote_headscale_standby.yml`) runs with `DOCKER_CONTEXT` set to it, so the modules and the CLI always reach the same daemon. The role checks the pin before it touches the host and fails if the context does not exist there, or if `DOCKER_HOST` is also set, since the two would name the daemon in conflicting ways. It never creates a context and never changes the host's current context; `docker context show` on the host reads the same before and after a run. The Headscale tasks that act on another host (the `headscale_hosted` server and standby, and the promotion playbook) use that host's own value. Run `docker context show` on each host to see what it currently points at before choosing a value.

## Mesh providers

`github_runner_cluster_mesh` picks the provider. Each one lives in `tasks/mesh/<provider>.yml` and publishes the same facts, so nothing else in the role depends on which is in use:
Expand Down Expand Up @@ -87,6 +93,7 @@ Each site is its own cluster: its own inventory group, bootstrap server and mesh
- `github_runner_cluster_headscale_*`: the Headscale providers' settings, documented in `defaults/main.yml`.
- `github_runner_cluster_node_name`: the k3s node name and mesh hostname, defaulting to the inventory name.
- Per-host overrides: `github_runner_cluster_node_role` (`server` or `agent`), `github_runner_cluster_bootstrap` (true picks the bootstrap host, false excludes a host), `github_runner_cluster_server_url` (join this server instead, for a cluster whose servers are outside the inventory), `github_runner_cluster_tls_sans`, `github_runner_cluster_node_address`. The run fails if the result has an even number of servers, no servers, or more than one bootstrap host.
- `github_runner_cluster_docker_context` (default empty): pins the Docker context this host's node runs in (see [Docker contexts](#docker-contexts)).
- `github_runner_cluster_compose_files`: extra Compose files for the host's k3s project, merged after `docker-compose.yml`, as paths relative to `github_runner_cluster_dir` or absolute. Several nodes on one Docker host each need their own `github_runner_cluster_dir` and `github_runner_cluster_compose_project`.
- `github_runner_cluster_wait_for_nodes`: wait for every node to be Ready afterwards, clearing stale node-password registrations if a node fails to rejoin (installs the kubernetes Python client through `github_runner_k8s_client`).

Expand Down Expand Up @@ -128,4 +135,4 @@ or read from 1Password by adding `exadev.github_runner.github_runner_secrets_one

`tests/litestream/run.sh` runs the `headscale_hosted` provider with Litestream and a warm standby against an S3-compatible store on this machine's Docker, registers a real Tailscale client, promotes the standby and checks that users, nodes and pre-auth keys survive, that the client reconnects and that the promoted server keeps replicating. It also checks that `headscale_in_cluster` refuses to run without Litestream. `.github/workflows/litestream.yml` runs it on pull requests that touch the Headscale providers.

`tests/mesh/run.sh` brings up three k3s nodes in Docker on one machine and joins them through Headscale with this role, for the `headscale_hosted` and `headscale_in_cluster` providers, and checks the `headscale_existing` policy check against a real server. Each cluster scenario runs the role a second time and fails if any node container restarted. The `compose_faulty` scenario checks the role refuses a Compose release that has that fault. `.github/workflows/mesh-integration.yml` runs them on pull requests that touch the role, the hosted scenario on 2.39.0, the first fixed release, as well as on a current one.
`tests/mesh/run.sh` brings up three k3s nodes in Docker on one machine and joins them through Headscale with this role, for the `headscale_hosted` and `headscale_in_cluster` providers, and checks the `headscale_existing` policy check against a real server. Each cluster scenario runs the role a second time and fails if any node container restarted. The `compose_faulty` scenario checks the role refuses a Compose release that has that fault. With `GRTEST_PIN_CONTEXT=1` the cluster scenarios pin `github_runner_cluster_docker_context` while the current Docker context the playbooks see points at a socket that does not exist, so any docker call that ignored the pin fails, and they check the current context is unchanged afterwards; `tests/litestream/run.sh` takes the same setting for the Headscale hosts it delegates to and the promotion playbook. The `context_checks` scenario checks the role refuses a pinned context that does not exist, and one set alongside `DOCKER_HOST`, before writing anything. `.github/workflows/mesh-integration.yml` runs them on pull requests that touch the role, the hosted scenario on 2.39.0, the first fixed release, as well as on a current one.
2 changes: 2 additions & 0 deletions roles/github_runner_cluster/defaults/main.yml
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,8 @@ github_runner_cluster_compose_file_list: >-
{{ ['docker-compose.yml']
+ (['docker-compose.mesh.yml'] if github_runner_cluster_mesh == 'headscale_in_cluster' and github_runner_cluster_effective_bootstrap | default(false) | bool else [])
+ github_runner_cluster_compose_files }}
# The Docker context this host's node runs in, pinned by name (for example "colima"). Empty, the default, follows the host's current Docker context, or DOCKER_HOST when that is set. When set, every community.docker module call in the role passes it as cli_context and every docker CLI command runs with DOCKER_CONTEXT set to it, so both always reach the same daemon, whatever the host's current context is at the time. The role fails before touching the host if the context does not exist there, or if DOCKER_HOST is set as well, and it never changes the host's current context. Set it on any host with more than one Docker engine installed: on a Mac running its node in Colima with Docker Desktop also installed, merely launching Docker Desktop makes desktop-linux the current context, and an unpinned run would then build a second, empty node in Docker Desktop instead of managing the real one. Per host, like every value here that names something on the host; the Headscale tasks read it from the inventory of the host they act on.
github_runner_cluster_docker_context: ""
# The node's k3s container, as docker-compose.yml names it. The headscale_in_cluster provider runs commands inside it; override it only alongside an extra Compose file that renames the container.
github_runner_cluster_container_name: github-runner-k3s

Expand Down
Loading
Loading