fix(cluster): refuse Docker Compose 2.37.1 to 2.38.x - #19
Merged
Merged
Conversation
Mearman
force-pushed
the
fix/compose-rerun-recreates
branch
from
September 26, 2026 06:50
7c945eb to
5126229
Compare
Compose 2.37.1 to 2.38.x build a service's image through Bake and key the build result by service name rather than image name (docker/compose#13047, fixed in 2.39.0), so the container created in the same `up` gets an empty com.docker.compose.image label. The next `up` sees that label differ from the image ID and recreates the container, restarting every k3s node on the run after the one that built the k3s image. The config hash and the image ID are unchanged between runs; the label is the only difference. The role now reads `docker compose version` before touching the host and fails with an explanation on a release in that range. Releases before it do not have the fault, so they stay accepted.
A compose_faulty scenario runs the cluster role under Compose 2.38.2 and asserts it refuses before writing to the host. The hosted scenario also runs on 2.39.0, the first fixed release, so the rerun check covers it.
Mearman
force-pushed
the
fix/compose-rerun-recreates
branch
from
September 26, 2026 06:51
5126229 to
556786c
Compare
Mearman
marked this pull request as ready for review
September 26, 2026 06:59
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
🎉 This PR is included in version 1.3.1 🎉 The release is available on:
Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #5
The recreation comes from Compose itself, not from anything the role does. On ubuntu-latest (Compose v2.38.2, Engine 28.0.4, overlay2 storage, not the containerd image store) I captured each node container's labels before and after the rerun in the hosted scenario. The config hash and the image ID were identical across the two runs; the only difference was the
com.docker.compose.imagelabel, which was empty on the containers the first run created.docker compose up --dry-runbefore the rerun showedRecreatefor all three nodes, and the module's result for the rerun showed the same. After that recreate the label is filled in, and a thirdupleaves the containers alone.The same thing happens with plain
docker compose up -dtwice on any service that hasbuild:(a one-lineFROM busyboxDockerfile is enough), with no Ansible involved. A service that only pulls its image is unaffected, because Compose fills the label in from the pulled image. Running the role's own Compose file twice under a range of Compose releases on ubuntu-latest:upkeeps the containeruprecreates it, third keeps itupkeeps itCOMPOSE_BAKE=false, or withdocker compose buildrun beforeup: keeps itSo it is the Bake build path. Before 2.39.0 Compose keyed Bake's build result by service name instead of image name, so the code that stamps the image ID on the new container found nothing. That was fixed in docker/compose#13047, released in 2.39.0. I did not reproduce the Docker Desktop case from the issue, so I can't say why 2.38.2 behaved there.
In a fleet this means one restart of every node on the run after the k3s image is built or rebuilt (a new host, or a k3s or Tailscale version bump), not on every run, but that is still every etcd member restarting at once.
Since both the fault and its fix are in Compose, the role now checks
docker compose versionas its first task and fails before touching the host on 2.37.1 up to but not including 2.39.0, saying why and to upgrade. I first wrote this as a plain minimum of 2.39.0, but a check-mode run against the fleet showed one host on Docker Desktop's Compose 2.29.7, which predates the fault and would have been refused for nothing, so the check refuses only the faulty range. The role README documents it and the top-level README's prerequisites point there.For CI, the mesh workflow gains two matrix entries: the hosted scenario on 2.39.0, whose rerun check shows the first fixed release keeps every node running, and a
compose_faultyscenario on 2.38.2 that asserts the role refuses it with that message and writes nothing to the node directory. The existing scenarios stay on 5.5.1.The diagnostics and the version matrix above ran from a scratch branch that has since been deleted. Check-mode runs of the fleet's
site.ymlwith this branch's collection on the two reachable hosts reported no change to the k3s Compose task on either; the only change reported was the kubernetes client virtualenv task on one host, which reports the same with the currently released collection.