Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,15 +120,15 @@ Confirm it refreshes: `gh gist view "$HEARTBEAT_GIST_ID"` must show a timestamp

### Autoscaler

A scale set's static `maxRunners` has to be safe for the pool's real worst case, every pod at its limit at once, so it leaves real idle capacity unused whenever the jobs in flight need less than that. `autoscaler` (an in-cluster Deployment, same shape/placement as heartbeat above) raises and lowers `maxRunners` on the one profile's `AutoscalingRunnerSet` with `autoscale: true`, based on real, current memory usage and pressure, never on job identity or GitHub Actions queue depth: every consuming repo's CI keeps using the scale set's one `runs-on` label unchanged, and no job is ever classified as light or heavy. RBAC is scoped to exactly what it reads/writes: cluster-wide `nodes.metrics.k8s.io` get/list, `get`/`patch` on that `AutoscalingRunnerSet` by name and `pods.metrics.k8s.io` get/list in its own namespace, and `get`/`patch` on its own status ConfigMap by name, with no wildcards.
A scale set's static `maxRunners` has to be safe for the pool's real worst case, every pod at its limit at once, so it leaves real idle capacity unused whenever the jobs in flight need less than that. `autoscaler` (an in-cluster Deployment, same shape/placement as heartbeat above) raises and lowers `maxRunners` across every profile with `autoscale: true` (one or more, possibly from different orgs, named by `AUTOSCALER_TARGETS`), based on real, current memory usage and pressure, never on job identity or GitHub Actions queue depth: every consuming repo's CI keeps using its scale set's own `runs-on` label unchanged, and no job is ever classified as light or heavy. When more than one profile is pooled, they share one combined budget rather than each getting a fixed slice: the algorithm below runs against the pool's combined `currentRunners`/`maxRunners`, and a change is handed to whichever pooled profile's own headroom (its `maxRunners` minus its `currentRunners`) makes it the right target: raises go to the profile with the *least* headroom (closest to genuinely needing more), lowers take from the profile with the *most* (its capacity is going unused), one runner at a time, and a lower never takes a profile below its own `currentRunners`. RBAC is scoped to exactly what it reads/writes: cluster-wide `nodes.metrics.k8s.io` get/list; per pooled profile, `get`/`patch` on that `AutoscalingRunnerSet` by name and `pods.metrics.k8s.io` get/list in its own namespace; and `get`/`patch` on its own status ConfigMap by name, with no wildcards.

Each poll (`AUTOSCALER_POLL_SECONDS`, default 45s):
- Reads the currently-running count (`status.currentRunners`) and the per-pod hard memory limit (`spec.template.spec.containers[0].resources.limits.memory`) straight from the live `AutoscalingRunnerSet`, never a static config value, so both always match what is actually deployed.
- Sums real, current memory usage of the running runner pods (`kubectl top pod` in the scale set's namespace) — this, not a static per-pod request or limit, is the entire point of "usage-driven".
- For every pooled profile, reads its currently-running count (`status.currentRunners`) and its per-pod hard memory limit (`spec.template.spec.containers[0].resources.limits.memory`) straight from its live `AutoscalingRunnerSet`, never a static config value, so both always match what is actually deployed; sums the running counts into one combined total, and takes the largest of the pod memory limits as the threshold a raise or lower must leave room for.
- Sums real, current memory usage of every pooled profile's running runner pods (`kubectl top pod` in each profile's own namespace) — this, not a static per-pod request or limit, is the entire point of "usage-driven".
- Reads cluster-wide memory-availability pressure via `kubectl top nodes`, taking the **worst-case (minimum)** available-memory percentage across all nodes, not an average — a pool average can look healthy while the specific node a new runner pod would actually land on is not, consistent with the algorithm's own bias toward lowering eagerly (lowering never disrupts in-flight jobs). This is a genuine improvement over the pod's own earlier host-pinned design, which could only ever see one fixed machine's pressure via `/proc/meminfo`, not the whole pool's. **No swap-pressure signal**: metrics-server's API has no swap field at all, and the only alternative (the kubelet's own Summary API) needs a materially broader RBAC grant for a signal most kubelet versions don't even surface unless an off-by-default feature gate is on — dropping swap detection is a deliberate, documented trade-off once the pod is genuinely unpinned to any node, not an oversight; memory-availability pressure remains the dominant signal (the incident the values file's own comments were tuned from).
- Raises `maxRunners` by exactly +1 per cycle (never straight to a computed target) once headroom has covered a full pod's hard limit for `AUTOSCALER_RAISE_CONFIRM_POLLS` (default 2) consecutive polls, capped at `AUTOSCALER_MAX_CEILING` (`github_runner_arc_autoscaler_max_ceiling`, or the profile's `sizing`): the bin-packing-safe ceiling across the pool at the profile's pod size, not an arbitrary cap.
- Lowers immediately, with no delay or averaging, the moment headroom drops below a pod's hard limit or real pressure is detected — lowering never disrupts in-flight jobs, since ARC only gates new claims. Never patches below the currently-running count.
- Fails safe on any measurement error (`kubectl` unreachable, `kubectl top nodes`/`kubectl top pod` failing): patches down to `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`, which must equal the profile's own static `maxRunners`) immediately rather than skipping the cycle silently.
- Raises the combined `maxRunners` by exactly +1 per cycle (never straight to a computed target) once headroom has covered a full pod's hard limit for `AUTOSCALER_RAISE_CONFIRM_POLLS` (default 2) consecutive polls, capped at `AUTOSCALER_MAX_CEILING` (`github_runner_arc_autoscaler_max_ceiling`, or, when exactly one pooled profile sets `sizing`, that profile's own derived ceiling): the bin-packing-safe ceiling across the pool at the profile's pod size, not an arbitrary cap.
- Lowers the combined `maxRunners` immediately, with no delay or averaging, the moment headroom drops below a pod's hard limit or real pressure is detected — lowering never disrupts in-flight jobs, since ARC only gates new claims. Never patches any pooled profile below its own currently-running count.
- Fails safe on any measurement error (`kubectl` unreachable, `kubectl top nodes`/`kubectl top pod` failing): lowers the combined total toward `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`) immediately rather than skipping the cycle silently. This combined floor need not equal any one profile's own values-file `maxRunners`, since each profile's Helm upgrade reverts to its own static floor independently of how the pool's combined floor is set.

Ships with `AUTOSCALER_DRY_RUN=true` by default: it computes and logs the target and writes its own status, but never patches. Generate real concurrent traffic to watch it react: `gh workflow run test-autoscaler.yml`, then watch `kubectl logs -n github-runner-platform deploy/autoscaler -f` and `kubectl get configmap autoscaler-status -n github-runner-platform -o jsonpath='{.data.status\.json}'`. Only set `AUTOSCALER_DRY_RUN=false` after watching real dry-run output across genuine CI traffic. Status and its raise-confirm counter live in a shared Kubernetes ConfigMap (`autoscaler-status`, pre-created empty so neither pod's own ServiceAccount ever needs `create` RBAC on it, only `get`/`patch`) rather than a bind-mounted directory, since the two pods can land on different nodes now — `scripts/heartbeat.sh` reads that same ConfigMap back out and republishes it as a second file in the heartbeat gist, reusing its existing gist-write credential rather than giving the autoscaler its own.

Expand Down
22 changes: 14 additions & 8 deletions plugins/filter/arc.py
Original file line number Diff line number Diff line change
Expand Up @@ -394,11 +394,15 @@ def _expand_profile(org_name: str, app_secret: str, index: int, profile: Any, er
max_runners = sizing_result["max_runners"]
else:
max_runners = None
namespace = _optional_name(profile, "namespace", where, errors) or f"arc-runners-{org_lower}{suffix}"
release = _optional_name(profile, "release_name", where, errors) or f"{org_lower}-runners{suffix}"
expanded = {
"org_name": org_name,
"suffix": suffix,
"namespace": _optional_name(profile, "namespace", where, errors) or f"arc-runners-{org_lower}{suffix}",
"release": _optional_name(profile, "release_name", where, errors) or f"{org_lower}-runners{suffix}",
"namespace": namespace,
"release": release,
# namespace/release joined for the autoscaler's own AUTOSCALER_TARGETS list (see scripts/autoscaler.sh): both are validated DNS labels below, so neither can itself contain '/', keeping each token unambiguous.
"target": f"{namespace}/{release}",
"app_secret": app_secret,
"scale_set_labels": _scale_set_labels(profile, f"{org_lower}-runners", where, errors),
"values_file": _text(profile, "values_file"),
Expand All @@ -420,19 +424,19 @@ def _expand_profile(org_name: str, app_secret: str, index: int, profile: Any, er
def arc_profiles(orgs: Sequence[Any], require_app_id: bool = True) -> dict[str, Any]:
"""Expand github_runner_arc_orgs entries into one record per scale-set profile, and check them.

Pass every org configured anywhere in the inventory, not one host's, so that the cross-host checks (an org configured twice, two profiles sharing a namespace, more than one autoscaled profile) see the whole fleet.
Pass every org configured anywhere in the inventory, not one host's, so that the cross-host checks (an org configured twice, two profiles sharing a namespace, more than one autoscaled profile setting its own sizing) see the whole fleet.

Args:
require_app_id: whether each org must set app_id. The role needs it only when it writes the App Secret; otherwise the App's id is in the existing App Secret, and an app_id that is set is only checked against it.
orgs: github_runner_arc_orgs entries, each with name, app_id (see require_app_id), image and a non-empty scale_set_profiles list, at most one of private_key and private_key_op_reference, and optionally app_secret_name. A profile may override its namespace and release_name, and set scale_set_labels in place of runs_on_label.

Returns:
A dict with ``errors`` (messages, empty when valid), ``profiles`` (one dict per profile with org_name, suffix, namespace, release, app_secret, scale_set_labels, values_file, node_selector, autoscale, sizing, max_runners (the static maxRunners the role sets, or None), label and settings, the profile's own keys) and ``autoscaled`` (the one profile flagged autoscale, or None).
A dict with ``errors`` (messages, empty when valid), ``profiles`` (one dict per profile with org_name, suffix, namespace, release, target (namespace/release, joined), app_secret, scale_set_labels, values_file, node_selector, autoscale, sizing, max_runners (the static maxRunners the role sets, or None), label and settings, the profile's own keys) and ``autoscaled`` (every profile flagged autoscale, sharing one autoscaler-managed memory pool across their combined scale sets; empty when none are).
"""
errors: list[str] = []
profiles: list[dict[str, Any]] = []
if not isinstance(orgs, Sequence) or isinstance(orgs, (str, bytes)):
return {"errors": ["github_runner_arc_orgs must be a list"], "profiles": [], "autoscaled": None}
return {"errors": ["github_runner_arc_orgs must be a list"], "profiles": [], "autoscaled": []}
seen_orgs: dict[str, str] = {}
for index, org in enumerate(orgs):
where = f"github_runner_arc_orgs[{index}]"
Expand Down Expand Up @@ -474,10 +478,12 @@ def arc_profiles(orgs: Sequence[Any], require_app_id: bool = True) -> dict[str,
if release_key in releases:
errors.append(f"{profile['label']} and {releases[release_key]} both use release name '{profile['release']}', which is the scale set's name on GitHub; give one of them a different suffix or release_name")
releases.setdefault(release_key, profile["label"])
# Several autoscaled profiles are allowed - the autoscaler pools their combined scale sets against one shared memory budget - but each one's own sizing derives that budget's ceiling, so more than one sizing among them would leave the ceiling ambiguous.
autoscaled = [profile for profile in profiles if profile["autoscale"]]
if len(autoscaled) > 1:
errors.append(f"At most one scale-set profile may set autoscale: true, because the autoscaler budgets one pool's memory for one scale set; found {', '.join(profile['label'] for profile in autoscaled)}")
return {"errors": errors, "profiles": profiles, "autoscaled": autoscaled[0] if len(autoscaled) == 1 else None}
sized_autoscaled = [profile for profile in autoscaled if profile["sizing"] is not None]
if len(sized_autoscaled) > 1:
errors.append(f"At most one autoscaled scale-set profile may set sizing, because it derives the shared pool's ceiling; found {', '.join(profile['label'] for profile in sized_autoscaled)}")
return {"errors": errors, "profiles": profiles, "autoscaled": autoscaled}


def arc_image_ref(image: str) -> dict[str, str]:
Expand Down
2 changes: 1 addition & 1 deletion roles/github_runner_arc/tasks/install_org.yml
Original file line number Diff line number Diff line change
Expand Up @@ -25,4 +25,4 @@
environment: "{{ {'KUBECONFIG': github_runner_arc_kubeconfig_path} if github_runner_arc_kubeconfig_path | length > 0 else {} }}"
changed_when: false
failed_when: false
when: github_runner_arc_fleet.autoscaled is not none and github_runner_arc_fleet.autoscaled.org_name == org.name
when: github_runner_arc_fleet.autoscaled | selectattr('org_name', 'equalto', org.name) | list | length > 0
Loading
Loading