Skip to content

feat(cluster): pin the Docker context the role drives per host - #21

Merged
Mearman merged 3 commits into
mainfrom
feat/cluster-docker-context
Sep 26, 2026
Merged

Mearman merged 3 commits into
mainfrom
feat/cluster-docker-context

Conversation

@Mearman

@Mearman Mearman commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

The cluster role works out which Docker daemon to drive from the host's current Docker context (docker context show), and every docker CLI call it makes follows that context too. That breaks on a host with more than one engine. On a Mac that runs its node in Colima and also has Docker Desktop installed, just launching Docker Desktop switches the current context to desktop-linux, and the next run of the role then finds no node there and builds a second, empty k3s node in Docker Desktop instead of managing the real one.

This adds github_runner_cluster_docker_context, a per-host variable. Empty (the default) keeps today's behaviour. When it is set:

  • every community.docker module call in the role passes it as cli_context;
  • every docker CLI call runs with DOCKER_CONTEXT set to it: docker compose version, the in-cluster Headscale image pull and docker exec probes and key reads, the Headscale API key creation, the hosted server's volume inspect/rm and compose exec, the promotion's compose stop/rm/run and exec, and every command in recover_in_cluster_mesh.yml;
  • the role fails before touching the host if the context does not exist there (listing the ones that do) or if DOCKER_HOST is also set;
  • nothing changes the host's current context.

The resolution lives in tasks/docker_context.yml, which the role and both playbooks include. Tasks delegated to a Headscale server or standby resolve against that host and read its own value from the inventory. The promotion's stop of the old server reads the value from the inventory without probing, since that host may be down.

Tests: tests/lib/docker_context.sh gives the playbook runs a throwaway DOCKER_CONFIG whose current context points at a socket that doesn't exist, plus a context named for the pin that points at the real daemon, so any call that ignored the pin fails. The in_cluster mesh scenario (role plus recovery playbook) and the Litestream test (delegated Headscale hosts plus promotion) now also run in that mode, and afterwards check the current context hasn't changed. A new context_checks scenario checks the two refusals happen before anything is written. There's also a unit test for the filter that builds the environment.

I ran the branch's collection against a real fleet in --check --diff mode, both with the variable unset and with it set to each host's current context, and got the same results as the released collection.

The role followed the host's current Docker context, which other software
changes: on a Mac running its node in Colima with Docker Desktop also
installed, launching Docker Desktop switches the current context to
desktop-linux, and the next run would build a second, empty node there.

github_runner_cluster_docker_context (default empty, keeping the current
behaviour) names the context explicitly. docker_context.yml resolves it
per host, failing before the host is touched if the context does not
exist or DOCKER_HOST is also set, and publishes the cli_context for the
community.docker modules and a DOCKER_CONTEXT environment for every
docker CLI call, including the recovery and promotion playbooks. Tasks
delegated to a Headscale host take that host's own value. The host's
current context is never changed.
tests/lib/docker_context.sh gives the playbook runs a throwaway Docker
CLI configuration whose current context points at a socket that does not
exist, plus a named context for the real daemon, so any docker call that
ignored github_runner_cluster_docker_context fails. The in_cluster mesh
scenario (role and recovery playbook) and the Litestream test (delegated
Headscale hosts and promotion) run that way, check the current context is
unchanged afterwards, and a context_checks scenario checks the role
refuses a missing context and one set alongside DOCKER_HOST before
writing anything.
The pinned-context check after the playbook made run_role end on an if
whose condition was false, so it returned success even when the role
failed, and compose_faulty saw its expected refusal as a pass.
@Mearman
Mearman marked this pull request as ready for review September 26, 2026 09:56
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-26T09:59:50.789197Z 9c3aca3 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@Mearman
Mearman merged commit eab77e0 into main Sep 26, 2026
22 checks passed
@Mearman
Mearman deleted the feat/cluster-docker-context branch September 26, 2026 09:57
@github-actions

Copy link
Copy Markdown

🎉 This PR is included in version 1.4.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant