Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions .github/workflows/node-recovery.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
name: Node recovery

# Runs the ARC role's node recovery watcher in a three-node kind cluster (tests/node_recovery/run.sh): the role's rendered Deployment and RBAC, the watcher image built from this checkout, and a stopped worker container standing in for a node that goes down. Only on changes that can affect the watcher.
#
# GitHub-hosted ubuntu-latest has Docker, kind, kubectl and jq already.

on:
pull_request:
paths:
- roles/github_runner_arc/**
- scripts/node-recovery.sh
- node-recovery/**
- tests/node_recovery/**
- .github/workflows/node-recovery.yml
workflow_dispatch:

concurrency:
group: node-recovery-${{ github.ref }}
cancel-in-progress: true

permissions:
contents: read

jobs:
recovery:
name: Recovery in kind
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v7

- name: Install Ansible
run: |
set -euo pipefail
if ! python3 -m venv "${RUNNER_TEMP}/venv" 2>/dev/null; then
sudo apt-get update -qq
sudo apt-get install -y -qq python3-venv
python3 -m venv "${RUNNER_TEMP}/venv"
fi
"${RUNNER_TEMP}/venv/bin/pip" install --quiet --prefer-binary ansible-core
echo "${RUNNER_TEMP}/venv/bin" >> "$GITHUB_PATH"

- name: Run the node recovery test
run: tests/node_recovery/run.sh
5 changes: 4 additions & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name: Release

# Runs semantic-release against every push to main. Almost every conventional-commit type triggers at least a patch release (see .releaserc.json's commit-analyzer releaseRules), so in practice this fires on nearly every push, stamping the collection version into galaxy.yml and the arc role's heartbeat, autoscaler and pull-secret-renewer image defaults, building the collection tarball, and, if a release was actually published, building and pushing all four container images (runner, heartbeat, autoscaler, pull-secret-renewer) tagged with the exact new version alongside floating major and latest tags on the same manifest.
# Runs semantic-release against every push to main. Almost every conventional-commit type triggers at least a patch release (see .releaserc.json's commit-analyzer releaseRules), so in practice this fires on nearly every push, stamping the collection version into galaxy.yml and the arc role's heartbeat, autoscaler, pull-secret-renewer and node-recovery image defaults, building the collection tarball, and, if a release was actually published, building and pushing all five container images (runner, heartbeat, autoscaler, pull-secret-renewer, node-recovery) tagged with the exact new version alongside floating major and latest tags on the same manifest.
#
# The release commit and tag are pushed over SSH with a write-enabled deploy key (secret RELEASE_DEPLOY_KEY): actions/checkout loads it for every later git command, and .releaserc.json's repositoryUrl is the SSH remote, so @semantic-release/git pushes with it. A deploy key is scoped to this one repository and can be listed as a ruleset bypass actor. It can't call the GitHub API, so the semantic-release process uses the ordinary secrets.GITHUB_TOKEN for creating the Release and uploading the collection tarball, which need no bypass.
#
Expand Down Expand Up @@ -76,6 +76,9 @@ jobs:
- name: pull-secret-renewer
dockerfile: pull-secret-renewer/Dockerfile
image: ghcr.io/exadev/github-runner-pull-secret-renewer
- name: node-recovery
dockerfile: node-recovery/Dockerfile
image: ghcr.io/exadev/github-runner-node-recovery
steps:
- uses: actions/checkout@v7
with:
Expand Down
20 changes: 16 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,8 @@ Which hosts are k3s servers is worked out from the inventory, not set per host.

## Build, test, and smoke-test

- **Build the runner image locally (the only working path right now):** `docker buildx build --platform linux/arm64,linux/amd64 --push -t ghcr.io/exadev/github-runner:latest .` from a machine with buildx/QEMU cross-platform support (Docker Desktop has this built in). A plain `docker build` with no `--platform` only produces an image for the building machine's own architecture, which silently breaks scheduling on the fleet's other architecture. The fleet-health platform's own two images need the identical treatment, for the identical reason (they're genuinely unpinned pods now too - see [Heartbeat](#heartbeat)/[Autoscaler](#autoscaler)): `docker buildx build --platform linux/arm64,linux/amd64 --push -f heartbeat/Dockerfile -t ghcr.io/exadev/github-runner-heartbeat:latest .` and the same with `autoscaler/Dockerfile` → `ghcr.io/exadev/github-runner-autoscaler:latest`, and `pull-secret-renewer/Dockerfile` → `ghcr.io/exadev/github-runner-pull-secret-renewer:latest` (the image of the ARC role's App-sourced pull Secret renewal).
- **Build and push via CI:** `.github/workflows/release.yml`'s `build-images` job builds and pushes all four images (this one, `heartbeat`, `autoscaler` and `pull-secret-renewer`) multi-arch on GitHub-hosted `ubuntu-latest` whenever a release is published - see [Releases](#releases) below.
- **Build the runner image locally (the only working path right now):** `docker buildx build --platform linux/arm64,linux/amd64 --push -t ghcr.io/exadev/github-runner:latest .` from a machine with buildx/QEMU cross-platform support (Docker Desktop has this built in). A plain `docker build` with no `--platform` only produces an image for the building machine's own architecture, which silently breaks scheduling on the fleet's other architecture. The fleet-health platform's own two images need the identical treatment, for the identical reason (they're genuinely unpinned pods now too - see [Heartbeat](#heartbeat)/[Autoscaler](#autoscaler)): `docker buildx build --platform linux/arm64,linux/amd64 --push -f heartbeat/Dockerfile -t ghcr.io/exadev/github-runner-heartbeat:latest .` and the same with `autoscaler/Dockerfile` → `ghcr.io/exadev/github-runner-autoscaler:latest`, `pull-secret-renewer/Dockerfile` → `ghcr.io/exadev/github-runner-pull-secret-renewer:latest` (the image of the ARC role's App-sourced pull Secret renewal), and `node-recovery/Dockerfile` → `ghcr.io/exadev/github-runner-node-recovery:latest` (see [Node recovery](#node-recovery)).
- **Build and push via CI:** `.github/workflows/release.yml`'s `build-images` job builds and pushes all five images (this one, `heartbeat`, `autoscaler`, `pull-secret-renewer` and `node-recovery`) multi-arch on GitHub-hosted `ubuntu-latest` whenever a release is published - see [Releases](#releases) below.
- **Smoke-test the ARC scale set** (confirms a job gets a real ephemeral pod): `gh workflow run test-arc-runner.yml`, then watch `kubectl get pods -n arc-runners-<org> -w`. A pod must appear only once the job is queued, run to completion, and be deleted within seconds.
- **Smoke-test `ubuntu-latest` routing** (confirms `runner-fallback-action` still reaches GitHub-hosted runners when needed): `gh workflow run test-ubuntu-latest.yml`.
- **Generate load for the autoscaler** (see [Autoscaler](#autoscaler) below): `gh workflow run test-autoscaler.yml`.
Expand All @@ -65,10 +65,10 @@ Every push to `main` runs `.github/workflows/release.yml`, driven by [semantic-r

A release:

1. Stamps the new version into `galaxy.yml`'s `version:` field (which stays `0.0.0` in source between releases) and into the published-image tag defaults in `roles/github_runner_arc/defaults/main.yml` (`github_runner_arc_heartbeat_image`, `github_runner_arc_autoscaler_image`, `github_runner_arc_image_pull_secret_renewer_image`) - `scripts/release/stamp_version.py` does this by matching each variable's own name, not a line number, so it stays correct even while those files are edited concurrently by unrelated work.
1. Stamps the new version into `galaxy.yml`'s `version:` field (which stays `0.0.0` in source between releases) and into the published-image tag defaults in `roles/github_runner_arc/defaults/main.yml` (`github_runner_arc_heartbeat_image`, `github_runner_arc_autoscaler_image`, `github_runner_arc_image_pull_secret_renewer_image`, `github_runner_arc_node_recovery_image`) - `scripts/release/stamp_version.py` does this by matching each variable's own name, not a line number, so it stays correct even while those files are edited concurrently by unrelated work.
2. Builds the collection tarball (`ansible-galaxy collection build --force`) and attaches it to the GitHub Release.
3. Updates `CHANGELOG.md` and commits it, `galaxy.yml`, and the stamped defaults file back to `main` as `chore(release): <version> [skip ci]`.
4. Builds and pushes all four container images (`Dockerfile`, `heartbeat/Dockerfile`, `autoscaler/Dockerfile`, `pull-secret-renewer/Dockerfile`) multi-arch (`linux/arm64,linux/amd64`) to their `ghcr.io/exadev/github-runner*` repositories, each tagged with the exact version, the floating major version, and `latest` - all three tags on the same manifest, via `docker/metadata-action`'s `type=semver` patterns. To pin a host to a specific release rather than floating on `latest`, set `roles/github_runner_arc/defaults/main.yml`'s image variables (or a host's own override) to an exact version tag instead.
4. Builds and pushes all five container images (`Dockerfile`, `heartbeat/Dockerfile`, `autoscaler/Dockerfile`, `pull-secret-renewer/Dockerfile`, `node-recovery/Dockerfile`) multi-arch (`linux/arm64,linux/amd64`) to their `ghcr.io/exadev/github-runner*` repositories, each tagged with the exact version, the floating major version, and `latest` - all three tags on the same manifest, via `docker/metadata-action`'s `type=semver` patterns. To pin a host to a specific release rather than floating on `latest`, set `roles/github_runner_arc/defaults/main.yml`'s image variables (or a host's own override) to an exact version tag instead.
5. Optionally publishes the collection to Ansible Galaxy, only if a `GALAXY_API_KEY` secret is present - the `[skip ci]` release commit and the image build both happen unconditionally, in this same run, regardless of whether Galaxy publishing is configured.

The `[skip ci]` in the release commit message is deliberate: the image build already happens in this same run using the exact version semantic-release just decided, so there is nothing useful for a second, separately-triggered run of this workflow (or of `ci.yml`) to do against a commit that only changed a version stamp and a changelog - `ci.yml`'s own lint jobs already ran against every commit before it reached `main`.
Expand Down Expand Up @@ -134,6 +134,18 @@ Ships with `AUTOSCALER_DRY_RUN=true` by default: it computes and logs the target

Every Helm upgrade above reverts `maxRunners` to the values file's safe floor. The role's `install_org.yml` triggers one immediate autoscaler poll after every such upgrade (`kubectl exec deploy/autoscaler -- ...`, reaching the pod through the API server regardless of which node it's on), so the safe-floor window after a redeploy is seconds, not a full poll interval.

### Node recovery

When a node stops (its host is switched off or asleep, Docker quits, or the k3s container is stopped), the node controller marks it `Unknown` once its kubelet has missed heartbeats for the node-monitor grace period, and after the pods' five-minute unreachable toleration marks them for deletion. Only the node's kubelet can confirm a deletion, so they then stay `Terminating` indefinitely. A Deployment's pods are replaced regardless, which is why the ARC controller comes back on its own, but ARC creates each scale set's listener as a single pod and waits for the old one to go, so a listener on the stopped node is never replaced and every job for that scale set queues until someone force-deletes it.

`node-recovery` (an in-cluster Deployment in the platform namespace, see `roles/github_runner_arc/templates/node-recovery.yaml.j2` and `scripts/node-recovery.sh`) applies Kubernetes' [non-graceful node shutdown](https://kubernetes.io/docs/concepts/cluster-administration/node-shutdown/#non-graceful-node-shutdown): every `NODE_RECOVERY_POLL_SECONDS` (default 15s) it gives each node whose `Ready` condition has been `Unknown` for `github_runner_arc_node_recovery_after_seconds` (default 120s) the `node.kubernetes.io/out-of-service=nodeshutdown:NoExecute` taint. The control plane then evicts every pod on that node at once, unless it tolerates that taint, and the pod garbage collector force-deletes the ones already terminating there, so ARC starts the listener again on a healthy node. When the node reports `Ready` again the watcher removes the taint, and the node's kubelet stops the containers of the pods that were deleted while it was away. With the defaults a stopped node stops holding up job pickup within roughly four minutes: the grace period (50s on current Kubernetes), the threshold, one poll, the garbage collector's 20-second sweep, and the listener's own restart. `github_runner_arc_node_recovery_enabled: false` removes the watcher and any taint it applied.

The watcher only acts on a node whose kubelet has gone silent (`Unknown`), never on one reporting `NotReady` (`False`), whose kubelet is alive and finishes its pods' deletion itself. It never taints the node it runs on, removes only taints it applied (recognised by its `github-runner.exadev/out-of-service-applied-at` annotation), and sends each change as a JSON patch that first tests the node's `resourceVersion`, so a node that came back since it was read is left for the next poll. RBAC is `get`, `list` and `patch` on nodes, with no `delete`: the watcher never removes a Node object, and the taint does not touch k3s's etcd membership, which only deleting a server's Node object changes. Its own pod tolerates an unreachable or not-ready node for only 10 seconds, so if it was running on the node that stopped, its ReplicaSet starts a replacement elsewhere well before the threshold passes.

The threshold is also where a stopped node's runner pods are given up, earlier than the five-minute default eviction would. Their containers stopped with the node, so their jobs are lost either way unless the node returns within that time; lower the threshold for faster recovery, or raise it to wait longer for a node that is only asleep.

Its image, `github_runner_arc_node_recovery_image`, is pulled without a pull Secret, so the published package must be public.

## Conventions

Adding a second org is additive, never a change to anything existing:
Expand Down
1 change: 1 addition & 0 deletions galaxy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ build_ignore:
- autoscaler
- heartbeat
- pull-secret-renewer
- node-recovery
- scripts
- Dockerfile
- bootstrap.sh
Expand Down
13 changes: 13 additions & 0 deletions node-recovery/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Image for the node recovery Deployment (see roles/github_runner_arc/templates/node-recovery.yaml.j2 and scripts/node-recovery.sh): bash, jq and kubectl, all from Alpine's apk, so it is small and builds for either architecture. Build context is the repository root so it can COPY the script. The Deployment pulls it without a pull Secret, so the published package must be public.
FROM alpine:3.20

RUN apk add --no-cache bash jq kubectl

COPY node-recovery/loop.sh /app/loop.sh
COPY scripts/node-recovery.sh /app/scripts/node-recovery.sh
RUN chmod +x /app/loop.sh /app/scripts/node-recovery.sh

# A non-root user by number, so the Deployment's runAsNonRoot check can verify it without reading /etc/passwd.
USER 65532:65532
WORKDIR /app
ENTRYPOINT ["/app/loop.sh"]
10 changes: 10 additions & 0 deletions node-recovery/loop.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
#!/usr/bin/env bash
# Container entrypoint for the node recovery Deployment (see roles/github_runner_arc/templates/node-recovery.yaml.j2). Runs scripts/node-recovery.sh every NODE_RECOVERY_POLL_SECONDS forever; a failed pass is logged and retried on the next one. kubectl finds the in-cluster config from the pod's mounted ServiceAccount token.
set -euo pipefail

NODE_RECOVERY_POLL_SECONDS="${NODE_RECOVERY_POLL_SECONDS:-15}"

while true; do
/app/scripts/node-recovery.sh || true
sleep "$NODE_RECOVERY_POLL_SECONDS"
done
5 changes: 5 additions & 0 deletions roles/github_runner_arc/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ Before touching the cluster the role then checks the secrets it is about to writ
- `github_runner_arc_image_pull_registry`, `github_runner_arc_image_pull_secret_name`, `github_runner_arc_verify_image_pull`: the pull Secret's registry and name, and whether to prove the credential before writing it.
- `github_runner_arc_heartbeat_gist_id`: install the fleet-health platform from this host, refreshing this gist. `github_runner_arc_heartbeat_bootstrap_gist`, `github_runner_arc_heartbeat_gist_description` and `github_runner_arc_heartbeat_gist_consumer` control the bootstrap described above.
- `github_runner_arc_autoscaler_usable_budget_gi`, `github_runner_arc_autoscaler_max_ceiling`, `github_runner_arc_autoscaler_floor`: required when a profile sets `autoscale: true` and the platform is installed. The floor must equal the autoscaled profile's `maxRunners`, which every Helm upgrade reverts to. The other `github_runner_arc_autoscaler_*` and `github_runner_arc_heartbeat_*` settings have defaults; see `defaults/main.yml`.
- `github_runner_arc_node_recovery_enabled` (default `true`), `github_runner_arc_node_recovery_after_seconds` (default 120), `github_runner_arc_node_recovery_poll_seconds` (default 15), `github_runner_arc_node_recovery_image`: the node recovery watcher (see [Node recovery](#node-recovery)).
- `github_runner_arc_app_setup_*`, `github_runner_arc_app_manifest_code`, `github_runner_arc_app_private_key_path`: inputs to `playbooks/github_app_setup.yml` (see below), including `github_runner_arc_app_setup_write_secret`, `github_runner_arc_app_setup_secret_namespaces`, `github_runner_arc_app_setup_secret_name`, `github_runner_arc_app_setup_keep_private_key` and `github_runner_arc_app_setup_command`.

## Sizing
Expand Down Expand Up @@ -96,6 +97,10 @@ github_runner_arc_orgs:
- max_runners: 4
```

## Node recovery

Every host that installs anything also installs a small watcher, the `node-recovery` Deployment in `github_runner_arc_platform_namespace`, so that a node that stops does not leave its pods `Terminating` for ever. ARC waits for a listener's old pod to go before starting a new one, so a listener on a stopped node otherwise stalls its scale set until the pod is force-deleted by hand. Once a node's `Ready` condition has been `Unknown` for `github_runner_arc_node_recovery_after_seconds`, the watcher gives it the `node.kubernetes.io/out-of-service=nodeshutdown:NoExecute` taint of Kubernetes' non-graceful node shutdown, which makes the control plane evict the node's pods and force-delete the terminating ones; it removes the taint when the node reports `Ready` again. It never acts on a node reporting `NotReady`, never on its own node, and never removes a taint it did not apply. Its ClusterRole allows `get`, `list` and `patch` on nodes and nothing else. `github_runner_arc_node_recovery_enabled: false` removes the watcher, its RBAC and any taint it left behind. Its image is pulled without a pull Secret, so it must be public. The repository README's Node recovery section explains the default threshold.

## GitHub App setup

`playbooks/github_app_setup.yml` creates the App the scale sets authenticate as, through GitHub's manifest flow. The first run renders a local form that posts the manifest (organisation self-hosted runner write access, no webhook, private) to GitHub; an organisation owner submits it and copies the one-time code GitHub returns. The second run, with that code and a path for the key, exchanges the code, writes the private key there, waits while the App is installed on the organisation, and prints the `app_id` and `installation_id` to record.
Expand Down
Loading
Loading