diff --git a/README.md b/README.md index 359152f..e1b7bef 100644 --- a/README.md +++ b/README.md @@ -120,15 +120,15 @@ Confirm it refreshes: `gh gist view "$HEARTBEAT_GIST_ID"` must show a timestamp ### Autoscaler -A scale set's static `maxRunners` has to be safe for the pool's real worst case, every pod at its limit at once, so it leaves real idle capacity unused whenever the jobs in flight need less than that. `autoscaler` (an in-cluster Deployment, same shape/placement as heartbeat above) raises and lowers `maxRunners` on the one profile's `AutoscalingRunnerSet` with `autoscale: true`, based on real, current memory usage and pressure, never on job identity or GitHub Actions queue depth: every consuming repo's CI keeps using the scale set's one `runs-on` label unchanged, and no job is ever classified as light or heavy. RBAC is scoped to exactly what it reads/writes: cluster-wide `nodes.metrics.k8s.io` get/list, `get`/`patch` on that `AutoscalingRunnerSet` by name and `pods.metrics.k8s.io` get/list in its own namespace, and `get`/`patch` on its own status ConfigMap by name, with no wildcards. +A scale set's static `maxRunners` has to be safe for the pool's real worst case, every pod at its limit at once, so it leaves real idle capacity unused whenever the jobs in flight need less than that. `autoscaler` (an in-cluster Deployment, same shape/placement as heartbeat above) raises and lowers `maxRunners` across every profile with `autoscale: true` (one or more, possibly from different orgs, named by `AUTOSCALER_TARGETS`), based on real, current memory usage and pressure, never on job identity or GitHub Actions queue depth: every consuming repo's CI keeps using its scale set's own `runs-on` label unchanged, and no job is ever classified as light or heavy. When more than one profile is pooled, they share one combined budget rather than each getting a fixed slice: the algorithm below runs against the pool's combined `currentRunners`/`maxRunners`, and a change is handed to whichever pooled profile's own headroom (its `maxRunners` minus its `currentRunners`) makes it the right target: raises go to the profile with the *least* headroom (closest to genuinely needing more), lowers take from the profile with the *most* (its capacity is going unused), one runner at a time, and a lower never takes a profile below its own `currentRunners`. RBAC is scoped to exactly what it reads/writes: cluster-wide `nodes.metrics.k8s.io` get/list; per pooled profile, `get`/`patch` on that `AutoscalingRunnerSet` by name and `pods.metrics.k8s.io` get/list in its own namespace; and `get`/`patch` on its own status ConfigMap by name, with no wildcards. Each poll (`AUTOSCALER_POLL_SECONDS`, default 45s): -- Reads the currently-running count (`status.currentRunners`) and the per-pod hard memory limit (`spec.template.spec.containers[0].resources.limits.memory`) straight from the live `AutoscalingRunnerSet`, never a static config value, so both always match what is actually deployed. -- Sums real, current memory usage of the running runner pods (`kubectl top pod` in the scale set's namespace) — this, not a static per-pod request or limit, is the entire point of "usage-driven". +- For every pooled profile, reads its currently-running count (`status.currentRunners`) and its per-pod hard memory limit (`spec.template.spec.containers[0].resources.limits.memory`) straight from its live `AutoscalingRunnerSet`, never a static config value, so both always match what is actually deployed; sums the running counts into one combined total, and takes the largest of the pod memory limits as the threshold a raise or lower must leave room for. +- Sums real, current memory usage of every pooled profile's running runner pods (`kubectl top pod` in each profile's own namespace) — this, not a static per-pod request or limit, is the entire point of "usage-driven". - Reads cluster-wide memory-availability pressure via `kubectl top nodes`, taking the **worst-case (minimum)** available-memory percentage across all nodes, not an average — a pool average can look healthy while the specific node a new runner pod would actually land on is not, consistent with the algorithm's own bias toward lowering eagerly (lowering never disrupts in-flight jobs). This is a genuine improvement over the pod's own earlier host-pinned design, which could only ever see one fixed machine's pressure via `/proc/meminfo`, not the whole pool's. **No swap-pressure signal**: metrics-server's API has no swap field at all, and the only alternative (the kubelet's own Summary API) needs a materially broader RBAC grant for a signal most kubelet versions don't even surface unless an off-by-default feature gate is on — dropping swap detection is a deliberate, documented trade-off once the pod is genuinely unpinned to any node, not an oversight; memory-availability pressure remains the dominant signal (the incident the values file's own comments were tuned from). -- Raises `maxRunners` by exactly +1 per cycle (never straight to a computed target) once headroom has covered a full pod's hard limit for `AUTOSCALER_RAISE_CONFIRM_POLLS` (default 2) consecutive polls, capped at `AUTOSCALER_MAX_CEILING` (`github_runner_arc_autoscaler_max_ceiling`, or the profile's `sizing`): the bin-packing-safe ceiling across the pool at the profile's pod size, not an arbitrary cap. -- Lowers immediately, with no delay or averaging, the moment headroom drops below a pod's hard limit or real pressure is detected — lowering never disrupts in-flight jobs, since ARC only gates new claims. Never patches below the currently-running count. -- Fails safe on any measurement error (`kubectl` unreachable, `kubectl top nodes`/`kubectl top pod` failing): patches down to `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`, which must equal the profile's own static `maxRunners`) immediately rather than skipping the cycle silently. +- Raises the combined `maxRunners` by exactly +1 per cycle (never straight to a computed target) once headroom has covered a full pod's hard limit for `AUTOSCALER_RAISE_CONFIRM_POLLS` (default 2) consecutive polls, capped at `AUTOSCALER_MAX_CEILING` (`github_runner_arc_autoscaler_max_ceiling`, or, when exactly one pooled profile sets `sizing`, that profile's own derived ceiling): the bin-packing-safe ceiling across the pool at the profile's pod size, not an arbitrary cap. +- Lowers the combined `maxRunners` immediately, with no delay or averaging, the moment headroom drops below a pod's hard limit or real pressure is detected — lowering never disrupts in-flight jobs, since ARC only gates new claims. Never patches any pooled profile below its own currently-running count. +- Fails safe on any measurement error (`kubectl` unreachable, `kubectl top nodes`/`kubectl top pod` failing): lowers the combined total toward `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`) immediately rather than skipping the cycle silently. This combined floor need not equal any one profile's own values-file `maxRunners`, since each profile's Helm upgrade reverts to its own static floor independently of how the pool's combined floor is set. Ships with `AUTOSCALER_DRY_RUN=true` by default: it computes and logs the target and writes its own status, but never patches. Generate real concurrent traffic to watch it react: `gh workflow run test-autoscaler.yml`, then watch `kubectl logs -n github-runner-platform deploy/autoscaler -f` and `kubectl get configmap autoscaler-status -n github-runner-platform -o jsonpath='{.data.status\.json}'`. Only set `AUTOSCALER_DRY_RUN=false` after watching real dry-run output across genuine CI traffic. Status and its raise-confirm counter live in a shared Kubernetes ConfigMap (`autoscaler-status`, pre-created empty so neither pod's own ServiceAccount ever needs `create` RBAC on it, only `get`/`patch`) rather than a bind-mounted directory, since the two pods can land on different nodes now — `scripts/heartbeat.sh` reads that same ConfigMap back out and republishes it as a second file in the heartbeat gist, reusing its existing gist-write credential rather than giving the autoscaler its own. diff --git a/plugins/filter/arc.py b/plugins/filter/arc.py index 74582fa..bdceaf6 100644 --- a/plugins/filter/arc.py +++ b/plugins/filter/arc.py @@ -394,11 +394,15 @@ def _expand_profile(org_name: str, app_secret: str, index: int, profile: Any, er max_runners = sizing_result["max_runners"] else: max_runners = None + namespace = _optional_name(profile, "namespace", where, errors) or f"arc-runners-{org_lower}{suffix}" + release = _optional_name(profile, "release_name", where, errors) or f"{org_lower}-runners{suffix}" expanded = { "org_name": org_name, "suffix": suffix, - "namespace": _optional_name(profile, "namespace", where, errors) or f"arc-runners-{org_lower}{suffix}", - "release": _optional_name(profile, "release_name", where, errors) or f"{org_lower}-runners{suffix}", + "namespace": namespace, + "release": release, + # namespace/release joined for the autoscaler's own AUTOSCALER_TARGETS list (see scripts/autoscaler.sh): both are validated DNS labels below, so neither can itself contain '/', keeping each token unambiguous. + "target": f"{namespace}/{release}", "app_secret": app_secret, "scale_set_labels": _scale_set_labels(profile, f"{org_lower}-runners", where, errors), "values_file": _text(profile, "values_file"), @@ -420,19 +424,19 @@ def _expand_profile(org_name: str, app_secret: str, index: int, profile: Any, er def arc_profiles(orgs: Sequence[Any], require_app_id: bool = True) -> dict[str, Any]: """Expand github_runner_arc_orgs entries into one record per scale-set profile, and check them. - Pass every org configured anywhere in the inventory, not one host's, so that the cross-host checks (an org configured twice, two profiles sharing a namespace, more than one autoscaled profile) see the whole fleet. + Pass every org configured anywhere in the inventory, not one host's, so that the cross-host checks (an org configured twice, two profiles sharing a namespace, more than one autoscaled profile setting its own sizing) see the whole fleet. Args: require_app_id: whether each org must set app_id. The role needs it only when it writes the App Secret; otherwise the App's id is in the existing App Secret, and an app_id that is set is only checked against it. orgs: github_runner_arc_orgs entries, each with name, app_id (see require_app_id), image and a non-empty scale_set_profiles list, at most one of private_key and private_key_op_reference, and optionally app_secret_name. A profile may override its namespace and release_name, and set scale_set_labels in place of runs_on_label. Returns: - A dict with ``errors`` (messages, empty when valid), ``profiles`` (one dict per profile with org_name, suffix, namespace, release, app_secret, scale_set_labels, values_file, node_selector, autoscale, sizing, max_runners (the static maxRunners the role sets, or None), label and settings, the profile's own keys) and ``autoscaled`` (the one profile flagged autoscale, or None). + A dict with ``errors`` (messages, empty when valid), ``profiles`` (one dict per profile with org_name, suffix, namespace, release, target (namespace/release, joined), app_secret, scale_set_labels, values_file, node_selector, autoscale, sizing, max_runners (the static maxRunners the role sets, or None), label and settings, the profile's own keys) and ``autoscaled`` (every profile flagged autoscale, sharing one autoscaler-managed memory pool across their combined scale sets; empty when none are). """ errors: list[str] = [] profiles: list[dict[str, Any]] = [] if not isinstance(orgs, Sequence) or isinstance(orgs, (str, bytes)): - return {"errors": ["github_runner_arc_orgs must be a list"], "profiles": [], "autoscaled": None} + return {"errors": ["github_runner_arc_orgs must be a list"], "profiles": [], "autoscaled": []} seen_orgs: dict[str, str] = {} for index, org in enumerate(orgs): where = f"github_runner_arc_orgs[{index}]" @@ -474,10 +478,12 @@ def arc_profiles(orgs: Sequence[Any], require_app_id: bool = True) -> dict[str, if release_key in releases: errors.append(f"{profile['label']} and {releases[release_key]} both use release name '{profile['release']}', which is the scale set's name on GitHub; give one of them a different suffix or release_name") releases.setdefault(release_key, profile["label"]) + # Several autoscaled profiles are allowed - the autoscaler pools their combined scale sets against one shared memory budget - but each one's own sizing derives that budget's ceiling, so more than one sizing among them would leave the ceiling ambiguous. autoscaled = [profile for profile in profiles if profile["autoscale"]] - if len(autoscaled) > 1: - errors.append(f"At most one scale-set profile may set autoscale: true, because the autoscaler budgets one pool's memory for one scale set; found {', '.join(profile['label'] for profile in autoscaled)}") - return {"errors": errors, "profiles": profiles, "autoscaled": autoscaled[0] if len(autoscaled) == 1 else None} + sized_autoscaled = [profile for profile in autoscaled if profile["sizing"] is not None] + if len(sized_autoscaled) > 1: + errors.append(f"At most one autoscaled scale-set profile may set sizing, because it derives the shared pool's ceiling; found {', '.join(profile['label'] for profile in sized_autoscaled)}") + return {"errors": errors, "profiles": profiles, "autoscaled": autoscaled} def arc_image_ref(image: str) -> dict[str, str]: diff --git a/roles/github_runner_arc/tasks/install_org.yml b/roles/github_runner_arc/tasks/install_org.yml index 3c6b679..db9c975 100644 --- a/roles/github_runner_arc/tasks/install_org.yml +++ b/roles/github_runner_arc/tasks/install_org.yml @@ -25,4 +25,4 @@ environment: "{{ {'KUBECONFIG': github_runner_arc_kubeconfig_path} if github_runner_arc_kubeconfig_path | length > 0 else {} }}" changed_when: false failed_when: false - when: github_runner_arc_fleet.autoscaled is not none and github_runner_arc_fleet.autoscaled.org_name == org.name + when: github_runner_arc_fleet.autoscaled | selectattr('org_name', 'equalto', org.name) | list | length > 0 diff --git a/roles/github_runner_arc/tasks/install_platform.yml b/roles/github_runner_arc/tasks/install_platform.yml index bced258..69c0893 100644 --- a/roles/github_runner_arc/tasks/install_platform.yml +++ b/roles/github_runner_arc/tasks/install_platform.yml @@ -9,7 +9,7 @@ kind: Namespace metadata: name: "{{ item }}" - loop: "{{ [github_runner_arc_platform_namespace, github_runner_arc_controller_namespace] + ([github_runner_arc_fleet.autoscaled.namespace] if github_runner_arc_fleet.autoscaled else []) }}" + loop: "{{ ([github_runner_arc_platform_namespace, github_runner_arc_controller_namespace] + (github_runner_arc_fleet.autoscaled | map(attribute='namespace') | list)) | unique }}" - name: "Platform: create heartbeat/autoscaler service accounts" kubernetes.core.k8s: @@ -114,7 +114,7 @@ apiGroup: rbac.authorization.k8s.io - name: "Platform: autoscaler RBAC" - when: github_runner_arc_fleet.autoscaled is not none + when: github_runner_arc_fleet.autoscaled | length > 0 block: # autoscaler: read per-node memory usage via metrics-server (cluster-wide, so a ClusterRole). Needs both rules, not metrics.k8s.io alone - confirmed live, kubectl top nodes fails outright with "nodes is forbidden ... at the cluster scope" without the plain core grant too, the same gap as kubectl top pod needing core pods alongside pods.metrics.k8s.io below. - name: "Platform: autoscaler RBAC - read node metrics" @@ -148,19 +148,19 @@ name: github-runner-autoscaler-node-metrics-reader apiGroup: rbac.authorization.k8s.io - # autoscaler: read and patch the autoscaled profile's AutoscalingRunnerSet (scoped to that one object by resourceNames), and read pod metrics in its namespace. - - name: "Platform: autoscaler RBAC - manage the autoscaled AutoscalingRunnerSet and read pod metrics" + # autoscaler: read and patch each pooled profile's own AutoscalingRunnerSet (scoped to that one object by resourceNames), and read pod metrics in its namespace. One Role/RoleBinding pair per pooled profile, since each lives in its own namespace. + - name: "Platform: autoscaler RBAC - manage each pooled AutoscalingRunnerSet and read pod metrics" kubernetes.core.k8s: definition: apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: github-runner-autoscaler-runnerset-manager - namespace: "{{ github_runner_arc_fleet.autoscaled.namespace }}" + namespace: "{{ item.namespace }}" rules: - apiGroups: ["actions.github.com"] resources: ["autoscalingrunnersets"] - resourceNames: ["{{ github_runner_arc_fleet.autoscaled.release }}"] + resourceNames: ["{{ item.release }}"] verbs: ["get", "patch"] - apiGroups: ["metrics.k8s.io"] resources: ["pods"] @@ -169,6 +169,9 @@ - apiGroups: [""] resources: ["pods"] verbs: ["get", "list"] + loop: "{{ github_runner_arc_fleet.autoscaled }}" + loop_control: + label: "{{ item.label }}" - name: "Platform: autoscaler RBAC - bind runnerset-manager" kubernetes.core.k8s: @@ -177,7 +180,7 @@ kind: RoleBinding metadata: name: github-runner-autoscaler-runnerset-manager - namespace: "{{ github_runner_arc_fleet.autoscaled.namespace }}" + namespace: "{{ item.namespace }}" subjects: - kind: ServiceAccount name: autoscaler @@ -186,6 +189,9 @@ kind: Role name: github-runner-autoscaler-runnerset-manager apiGroup: rbac.authorization.k8s.io + loop: "{{ github_runner_arc_fleet.autoscaled }}" + loop_control: + label: "{{ item.label }}" # autoscaler: read/patch the shared status ConfigMap, scoped by resourceNames. - name: "Platform: autoscaler RBAC - manage the status ConfigMap" @@ -328,24 +334,32 @@ if github_runner_arc_controller_namespace != 'actions-runner-controller' else [] }} -- name: "Platform: measure the autoscaled profile's ceiling from the eligible nodes" +- name: "Platform: measure the autoscaled pool's ceiling from the eligible nodes" + # arc_profiles already rejects more than one autoscaled profile setting sizing, so at most one iteration of this loop ever runs its task. ansible.builtin.include_tasks: measure_sizing.yml - vars: - github_runner_arc_measure: "{{ github_runner_arc_fleet.autoscaled }}" - when: github_runner_arc_fleet.autoscaled is not none and github_runner_arc_fleet.autoscaled.sizing is not none and github_runner_arc_fleet.autoscaled.sizing.measured + loop: "{{ github_runner_arc_fleet.autoscaled }}" + loop_control: + loop_var: github_runner_arc_measure + label: "{{ github_runner_arc_measure.label }}" + when: github_runner_arc_measure.sizing is not none and github_runner_arc_measure.sizing.measured - name: "Platform: fail when the autoscaler's floor is above its measured ceiling" ansible.builtin.fail: msg: >- github_runner_arc_autoscaler_floor ({{ github_runner_arc_autoscaler_floor }}) is above the ceiling measured for - {{ github_runner_arc_fleet.autoscaled.label }} ({{ github_runner_arc_measured_sizing[github_runner_arc_fleet.autoscaled.label].max_runners }}). + {{ item.label }} ({{ github_runner_arc_measured_sizing[item.label].max_runners }}). + loop: "{{ github_runner_arc_fleet.autoscaled }}" + loop_control: + label: "{{ item.label }}" when: - - github_runner_arc_fleet.autoscaled is not none - - github_runner_arc_fleet.autoscaled.label in (github_runner_arc_measured_sizing | default({})) - - (github_runner_arc_autoscaler_floor | int) > (github_runner_arc_measured_sizing[github_runner_arc_fleet.autoscaled.label].max_runners | int) + - item.label in (github_runner_arc_measured_sizing | default({})) + - (github_runner_arc_autoscaler_floor | int) > (github_runner_arc_measured_sizing[item.label].max_runners | int) - name: "Platform: install/upgrade the autoscaler Deployment" - when: github_runner_arc_fleet.autoscaled is not none + when: github_runner_arc_fleet.autoscaled | length > 0 + vars: + # arc_profiles already rejects more than one autoscaled profile setting sizing, so this is at most one profile. + github_runner_arc_sized: "{{ github_runner_arc_fleet.autoscaled | selectattr('sizing', 'defined') | selectattr('sizing', 'ne', none) | list }}" kubernetes.core.k8s: definition: apiVersion: apps/v1 @@ -378,9 +392,9 @@ - name: AUTOSCALER_MAX_CEILING value: >- {{ - (github_runner_arc_measured_sizing[github_runner_arc_fleet.autoscaled.label].max_runners - if github_runner_arc_fleet.autoscaled.label in (github_runner_arc_measured_sizing | default({})) - else github_runner_arc_fleet.autoscaled.sizing.max_runners if github_runner_arc_fleet.autoscaled.sizing + (github_runner_arc_measured_sizing[(github_runner_arc_sized | first).label].max_runners + if github_runner_arc_sized | length > 0 and (github_runner_arc_sized | first).label in (github_runner_arc_measured_sizing | default({})) + else (github_runner_arc_sized | first).sizing.max_runners if github_runner_arc_sized | length > 0 else github_runner_arc_autoscaler_max_ceiling) | string }} - name: AUTOSCALER_FLOOR @@ -391,11 +405,9 @@ value: "{{ github_runner_arc_autoscaler_mem_available_pressure_pct | string }}" - name: AUTOSCALER_STATE_NAMESPACE value: "{{ github_runner_arc_platform_namespace }}" - # The scale set it manages, and where its status lives, named explicitly so the script's own defaults never apply. - - name: AUTOSCALER_NAMESPACE - value: "{{ github_runner_arc_fleet.autoscaled.namespace }}" - - name: AUTOSCALER_RELEASE_NAME - value: "{{ github_runner_arc_fleet.autoscaled.release }}" + # The pooled scale sets it manages, one namespace/release token per profile (see plugins/filter/arc.py's own target field), space-separated: the script splits on whitespace. + - name: AUTOSCALER_TARGETS + value: "{{ github_runner_arc_fleet.autoscaled | map(attribute='target') | join(' ') }}" - name: AUTOSCALER_STATUS_CONFIGMAP value: autoscaler-status resources: diff --git a/roles/github_runner_arc/tasks/validate.yml b/roles/github_runner_arc/tasks/validate.yml index e35e4e1..46ad6b4 100644 --- a/roles/github_runner_arc/tasks/validate.yml +++ b/roles/github_runner_arc/tasks/validate.yml @@ -81,15 +81,17 @@ - name: Fail when the autoscaler's pool settings are missing or inconsistent ansible.builtin.fail: msg: >- - The autoscaled profile {{ github_runner_arc_fleet.autoscaled.label }} needs - github_runner_arc_autoscaler_usable_budget_gi, github_runner_arc_autoscaler_floor and, unless the profile sets sizing, - github_runner_arc_autoscaler_max_ceiling, all whole numbers, with the floor no higher than the ceiling. + The autoscaled pool ({{ github_runner_arc_fleet.autoscaled | map(attribute='label') | join(', ') }}) needs + github_runner_arc_autoscaler_usable_budget_gi, github_runner_arc_autoscaler_floor and, unless one of its profiles sets + sizing, github_runner_arc_autoscaler_max_ceiling, all whole numbers, with the floor no higher than the ceiling. vars: - github_runner_arc_ceiling: "{{ github_runner_arc_fleet.autoscaled.sizing.max_runners if github_runner_arc_fleet.autoscaled.sizing else (github_runner_arc_autoscaler_max_ceiling | default('')) }}" + # arc_profiles already rejects more than one autoscaled profile setting sizing, so this is at most one profile. + github_runner_arc_sized: "{{ github_runner_arc_fleet.autoscaled | selectattr('sizing', 'defined') | selectattr('sizing', 'ne', none) | list }}" + github_runner_arc_ceiling: "{{ (github_runner_arc_sized | first).sizing.max_runners if github_runner_arc_sized | length > 0 else (github_runner_arc_autoscaler_max_ceiling | default('')) }}" # A measured ceiling is only known once the cluster is read, so install_platform.yml checks the floor against it then. - github_runner_arc_ceiling_measured: "{{ github_runner_arc_fleet.autoscaled.sizing.measured | default(false) if github_runner_arc_fleet.autoscaled.sizing else false }}" + github_runner_arc_ceiling_measured: "{{ (github_runner_arc_sized | first).sizing.measured | default(false) if github_runner_arc_sized | length > 0 else false }}" when: - - github_runner_arc_fleet.autoscaled is not none + - github_runner_arc_fleet.autoscaled | length > 0 - github_runner_arc_heartbeat_gist_id | length > 0 - >- github_runner_arc_autoscaler_usable_budget_gi | default('') | string is not match('^[0-9]+$') diff --git a/scripts/autoscaler.sh b/scripts/autoscaler.sh index 77dc034..1a59593 100755 --- a/scripts/autoscaler.sh +++ b/scripts/autoscaler.sh @@ -1,21 +1,20 @@ #!/usr/bin/env bash -# Usage-driven maxRunners autoscaler. Runs inside the `autoscaler` in-cluster Deployment (see autoscaler/loop.sh for the poll loop). Never inspects job identity or the GitHub Actions queue: it only watches real, current memory usage and pressure, and adjusts spec.maxRunners on the single AutoscalingRunnerSet named by AUTOSCALER_NAMESPACE and AUTOSCALER_RELEASE_NAME accordingly. See the project README's Autoscaler section for the full algorithm and rationale. +# Usage-driven maxRunners autoscaler. Runs inside the `autoscaler` in-cluster Deployment (see autoscaler/loop.sh for the poll loop). Never inspects job identity or the GitHub Actions queue: it only watches real, current memory usage and pressure, and adjusts spec.maxRunners across every AutoscalingRunnerSet named in AUTOSCALER_TARGETS, pooling them against one combined memory budget rather than budgeting each separately. See the project README's Autoscaler section for the full algorithm and rationale. # # Environment (provided by the Deployment - see roles/github_runner_arc/tasks/install_platform.yml): # - kubectl needs no KUBECONFIG here: running as a pod with a mounted ServiceAccount token, client-go auto-detects in-cluster config. # - AUTOSCALER_DRY_RUN: "true" computes and logs only, never patches -# - AUTOSCALER_USABLE_BUDGET_GI: proven-safe memory budget for all runner pods combined, across every node the scale set can use -# - AUTOSCALER_MAX_CEILING: hard operator cap on maxRunners, independent of the live arithmetic -# - AUTOSCALER_FLOOR: the static floor every Helm upgrade reverts to +# - AUTOSCALER_USABLE_BUDGET_GI: proven-safe memory budget for every pooled target's runner pods combined, across every node any of them can use +# - AUTOSCALER_MAX_CEILING: hard operator cap on the pooled maxRunners total, independent of the live arithmetic +# - AUTOSCALER_FLOOR: the combined floor every poll reverts the pooled total to under pressure; each Helm upgrade separately reverts its own target's maxRunners to that target's own values-file floor, which need not equal this combined figure divided evenly # - AUTOSCALER_RAISE_CONFIRM_POLLS: consecutive polls of confirmed headroom before raising # - AUTOSCALER_MEM_AVAILABLE_PRESSURE_PCT: MemAvailable-of-MemTotal percentage treated as real pressure, evaluated as the worst case across all nodes (see the kubectl top nodes block below) - no swap-pressure equivalent, see that block's own comment for why -# - AUTOSCALER_NAMESPACE, AUTOSCALER_RELEASE_NAME: the AutoscalingRunnerSet to manage (the scale-set profile with autoscale: true) +# - AUTOSCALER_TARGETS: the AutoscalingRunnerSets to pool, as whitespace-separated namespace/release tokens (one per scale-set profile with autoscale: true; see plugins/filter/arc.py's own target field, which is exactly this namespace/release form) # - AUTOSCALER_STATE_NAMESPACE, AUTOSCALER_STATUS_CONFIGMAP: the ConfigMap this pod's own status/raise-confirm-count state lives in (shared with scripts/heartbeat.sh, which reads status.json back out to republish in the heartbeat gist) set -euo pipefail -# The target and the pool-specific figures have no defaults: they depend on the cluster, and the role always sets them. -NAMESPACE="${AUTOSCALER_NAMESPACE:?AUTOSCALER_NAMESPACE must be set}" -RELEASE_NAME="${AUTOSCALER_RELEASE_NAME:?AUTOSCALER_RELEASE_NAME must be set}" +# The targets and the pool-specific figures have no defaults: they depend on the cluster, and the role always sets them. +TARGETS_RAW="${AUTOSCALER_TARGETS:?AUTOSCALER_TARGETS must be set}" DRY_RUN="${AUTOSCALER_DRY_RUN:-true}" USABLE_BUDGET_GI="${AUTOSCALER_USABLE_BUDGET_GI:?AUTOSCALER_USABLE_BUDGET_GI must be set}" MAX_CEILING="${AUTOSCALER_MAX_CEILING:?AUTOSCALER_MAX_CEILING must be set}" @@ -25,6 +24,16 @@ MEM_AVAILABLE_PRESSURE_PCT="${AUTOSCALER_MEM_AVAILABLE_PRESSURE_PCT:-15}" STATE_NAMESPACE="${AUTOSCALER_STATE_NAMESPACE:-github-runner-platform}" CONFIGMAP="${AUTOSCALER_STATUS_CONFIGMAP:-autoscaler-status}" +# Splits on whitespace: one token or several, however the Deployment env's own spacing renders. +read -ra TARGETS <<< "$TARGETS_RAW" +[ "${#TARGETS[@]}" -gt 0 ] || { echo "AUTOSCALER_TARGETS named no targets" >&2; exit 1; } +NAMESPACES=() +RELEASES=() +for target in "${TARGETS[@]}"; do + NAMESPACES+=("${target%%/*}") + RELEASES+=("${target#*/}") +done + # Kubernetes memory quantities as reported by `kubectl get -o jsonpath` and `kubectl top pod` use binary suffixes (Ki/Mi/Gi). Convert to whole MiB so every comparison below is plain integer arithmetic — bash has no floats, and shelling out to `bc`/`python3` for this is unnecessary weight in a tiny Alpine image. to_mib() { local qty="$1" @@ -46,64 +55,140 @@ configmap_set() { -p "$(jq -n --arg k "$1" --arg v "$2" '{data: {($k): $v}}')" } +# targets_json: a compact JSON array, one object per target, built from the parallel NAMESPACES/RELEASES/MAX arrays - so the status ConfigMap (and the heartbeat gist it feeds) shows each pooled target's own current maxRunners, not just the combined total. +targets_json() { + local i out="[]" + for ((i = 0; i < ${#TARGETS[@]}; i++)); do + out="$(jq -c --argjson acc "$out" --arg namespace "${NAMESPACES[$i]}" --arg release "${RELEASES[$i]}" --argjson max_runners "${MAX[$i]}" --argjson current_runners "${R[$i]}" \ + '$acc + [{namespace: $namespace, release: $release, max_runners: $max_runners, current_runners: $current_runners}]' <<< null)" + done + echo "$out" +} + write_status() { - local current_max="$1" mode="$2" headroom_mib="${3:-null}" headroom_gi="null" + local current_max_total="$1" mode="$2" headroom_mib="${3:-null}" headroom_gi="null" if [ "$headroom_mib" != "null" ]; then headroom_gi="$(awk -v m="$headroom_mib" 'BEGIN{printf "%.1f", m/1024}')" fi local status_json status_json="$(jq -n \ --arg last_run "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \ - --argjson current_max_runners "$current_max" \ + --argjson current_max_runners "$current_max_total" \ --arg headroom_gi "$headroom_gi" \ --arg mode "$mode" \ - '{last_run: $last_run, current_max_runners: $current_max_runners, headroom_gi: ($headroom_gi | tonumber? // null), mode: $mode}')" + --argjson targets "$(targets_json)" \ + '{last_run: $last_run, current_max_runners: $current_max_runners, headroom_gi: ($headroom_gi | tonumber? // null), mode: $mode, targets: $targets}')" configmap_set "status.json" "$status_json" } -patch_max_runners() { - local target="$1" reason="$2" headroom_mib="${3:-null}" - if [ "$DRY_RUN" = "true" ]; then - echo "DRY RUN: would patch maxRunners to ${target} (${reason})" - else - kubectl patch autoscalingrunnerset "$RELEASE_NAME" -n "$NAMESPACE" \ - --type merge -p "{\"spec\":{\"maxRunners\": ${target}}}" - echo "Patched maxRunners to ${target} (${reason})" - fi - write_status "$target" "$reason" "$headroom_mib" +# Patches every target whose NEW_MAX differs from its live MAX (set by the caller before calling this), skipping the rest: an unchanged target needs no API call. Updates MAX in place afterwards so targets_json (via write_status) reports what was actually applied. +apply_new_max() { + local reason="$1" headroom_mib="${2:-null}" i total=0 + for ((i = 0; i < ${#TARGETS[@]}; i++)); do + total=$(( total + NEW_MAX[i] )) + [ "${NEW_MAX[$i]}" -eq "${MAX[$i]}" ] && continue + if [ "$DRY_RUN" = "true" ]; then + echo "DRY RUN: would patch ${TARGETS[$i]}'s maxRunners to ${NEW_MAX[$i]} (${reason})" + else + kubectl patch autoscalingrunnerset "${RELEASES[$i]}" -n "${NAMESPACES[$i]}" \ + --type merge -p "{\"spec\":{\"maxRunners\": ${NEW_MAX[$i]}}}" + echo "Patched ${TARGETS[$i]}'s maxRunners to ${NEW_MAX[$i]} (${reason})" + fi + MAX[i]=${NEW_MAX[$i]} + done + write_status "$total" "$reason" "$headroom_mib" +} + +# Lowers the combined total toward $1 by taking slack (a target's own MAX minus its own R) from whichever pooled target currently has the most of it, one runner at a time, so the target closest to actually using its allotment keeps the most of what it has. Never takes a target below its own R: a currently-running job must never be orphaned by its own scale set shrinking under it, so the combined result can stay above $1 when demand alone already exceeds it. +lower_to() { + local combined_target=$1 i combined=0 + for ((i = 0; i < ${#TARGETS[@]}; i++)); do + NEW_MAX[i]=${MAX[$i]} + combined=$(( combined + NEW_MAX[i] )) + done + while [ "$combined" -gt "$combined_target" ]; do + local best=-1 best_slack=0 + for ((i = 0; i < ${#TARGETS[@]}; i++)); do + local slack=$(( NEW_MAX[i] - R[i] )) + if [ "$slack" -gt "$best_slack" ]; then + best_slack=$slack + best=$i + fi + done + [ "$best" -ge 0 ] || break # every target is already down at its own R; can't lower further without orphaning a running job + NEW_MAX[best]=$(( NEW_MAX[best] - 1 )) + combined=$(( combined - 1 )) + done +} + +# Raises the combined total by exactly one runner, handed to whichever pooled target currently has the least slack (its own MAX minus its own R): the one closest to fully using what it already has, so the extra capacity goes where it is actually needed rather than to whichever target happens to be first. +raise_by_one() { + local i best=0 best_slack=$(( MAX[0] - R[0] )) + for ((i = 1; i < ${#TARGETS[@]}; i++)); do + local slack=$(( MAX[i] - R[i] )) + if [ "$slack" -lt "$best_slack" ]; then + best_slack=$slack + best=$i + fi + done + for ((i = 0; i < ${#TARGETS[@]}; i++)); do NEW_MAX[i]=${MAX[$i]}; done + NEW_MAX[best]=$(( NEW_MAX[best] + 1 )) } fail_safe() { - echo "FAIL-SAFE: $1 — patching maxRunners down to the static floor (${FLOOR})" >&2 + echo "FAIL-SAFE: $1 — lowering the pooled total toward the combined floor (${FLOOR})" >&2 configmap_set "raise-confirm-count" "0" - patch_max_runners "$FLOOR" "fail-safe: $1" + lower_to "$FLOOR" + apply_new_max "fail-safe: $1" exit 0 } # ---- Gather signals --------------------------------------------------------- -# R: the authoritative currently-running count, read from the AutoscalingRunnerSet CRD's own status (not a pod-label guess). The role gives every scale-set profile its own namespace, so the target namespace holds exactly this one scale set and no scale-set-specific label selector is needed for anything below. -R="$(kubectl get autoscalingrunnerset "$RELEASE_NAME" -n "$NAMESPACE" -o jsonpath='{.status.currentRunners}' 2>/dev/null)" \ - || fail_safe "could not read ${RELEASE_NAME}'s status.currentRunners" -case "$R" in ''|*[!0-9]*) fail_safe "status.currentRunners was not a plain integer ('$R')" ;; esac - -# Wh: the pod's own hard memory limit, read from the live deployed spec, not re-parsed from values/*.yaml, so this always matches what is actually running even after a manual --set override. -WH_RAW="$(kubectl get autoscalingrunnerset "$RELEASE_NAME" -n "$NAMESPACE" \ - -o jsonpath='{.spec.template.spec.containers[0].resources.limits.memory}' 2>/dev/null)" \ - || fail_safe "could not read the runner pod's own memory limit" -[ -n "$WH_RAW" ] || fail_safe "the runner pod's memory limit was empty" -WH_MIB="$(to_mib "$WH_RAW")" || fail_safe "could not parse the runner pod's memory limit ('$WH_RAW')" - -# Actual current memory usage of running runner pods, summed — this, not a static per-pod request/limit, is the entire point of "usage-driven". Requires metrics-server; k3s bundles it by default unless explicitly disabled with --disable metrics-server (this repo's docker-compose.yml only disables traefik). -TOP_OUTPUT="$(kubectl top pod -n "$NAMESPACE" --no-headers 2>/dev/null)" \ - || fail_safe "kubectl top pod failed (metrics-server unreachable?)" +declare -a R MAX USAGE_MIB=0 -if [ -n "$TOP_OUTPUT" ]; then - while read -r _name _cpu mem _rest; do - pod_mib="$(to_mib "$mem")" || fail_safe "could not parse a runner pod's memory usage ('$mem')" - USAGE_MIB=$(( USAGE_MIB + pod_mib )) - done <<< "$TOP_OUTPUT" -fi +WH_MIB=0 # the largest pooled target's own pod memory limit, used as the headroom threshold: raising or lowering must leave room for whichever pooled target's next pod would be the largest. +for i in "${!TARGETS[@]}"; do + namespace=${NAMESPACES[$i]} + release=${RELEASES[$i]} + target=${TARGETS[$i]} + + # R: the authoritative currently-running count, read from the AutoscalingRunnerSet CRD's own status (not a pod-label guess). The role gives every scale-set profile its own namespace, so each target namespace holds exactly this one scale set and no scale-set-specific label selector is needed for anything below. + r="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" -o jsonpath='{.status.currentRunners}' 2>/dev/null)" \ + || fail_safe "could not read ${target}'s status.currentRunners" + case "$r" in ''|*[!0-9]*) fail_safe "${target}'s status.currentRunners was not a plain integer ('$r')" ;; esac + R[i]=$r + + max="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" -o jsonpath='{.spec.maxRunners}' 2>/dev/null)" \ + || fail_safe "could not read ${target}'s current spec.maxRunners" + case "$max" in ''|*[!0-9]*) fail_safe "${target}'s spec.maxRunners was not a plain integer ('$max')" ;; esac + MAX[i]=$max + + # Wh: the pod's own hard memory limit, read from the live deployed spec, not re-parsed from values/*.yaml, so this always matches what is actually running even after a manual --set override. + wh_raw="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" \ + -o jsonpath='{.spec.template.spec.containers[0].resources.limits.memory}' 2>/dev/null)" \ + || fail_safe "could not read ${target}'s runner pod memory limit" + [ -n "$wh_raw" ] || fail_safe "${target}'s runner pod memory limit was empty" + wh_mib="$(to_mib "$wh_raw")" || fail_safe "could not parse ${target}'s runner pod memory limit ('$wh_raw')" + [ "$wh_mib" -gt "$WH_MIB" ] && WH_MIB=$wh_mib + + # Actual current memory usage of running runner pods, summed — this, not a static per-pod request/limit, is the entire point of "usage-driven". Requires metrics-server; k3s bundles it by default unless explicitly disabled with --disable metrics-server (this repo's docker-compose.yml only disables traefik). + top_output="$(kubectl top pod -n "$namespace" --no-headers 2>/dev/null)" \ + || fail_safe "kubectl top pod failed for ${namespace} (metrics-server unreachable?)" + if [ -n "$top_output" ]; then + while read -r _name _cpu mem _rest; do + pod_mib="$(to_mib "$mem")" || fail_safe "could not parse a runner pod's memory usage in ${namespace} ('$mem')" + USAGE_MIB=$(( USAGE_MIB + pod_mib )) + done <<< "$top_output" + fi +done + +R_TOTAL=0 +MAX_TOTAL=0 +for i in "${!TARGETS[@]}"; do + R_TOTAL=$(( R_TOTAL + R[i] )) + MAX_TOTAL=$(( MAX_TOTAL + MAX[i] )) +done # Cluster-wide memory-availability pressure via `kubectl top nodes` (metrics-server), not a single node's own /proc/meminfo - the autoscaler pod is genuinely unpinned now (see docker-compose.yml's own removal of host pinning), so a single-node reading would only ever reflect wherever the scheduler happened to place THIS pod, not the pool as a whole. Takes the WORST-CASE (minimum) available-memory percentage across all nodes, not an average: a pool average can look healthy while the specific node a new runner pod would actually land on is not - consistent with the algorithm's own existing bias toward lowering eagerly, since lowering never disrupts in-flight jobs. No swap-pressure signal: metrics-server's API has no swap field at all (confirmed against its type, not just "probably not"), and the only alternative (the kubelet's own Summary API) needs a materially broader RBAC grant - node/proxy, effectively "reach anything a kubelet exposes" - for a signal most kubelet versions don't even surface unless the off-by-default NodeSwap feature gate is on. Real cluster-wide swap visibility needs Prometheus/node-exporter, which this fleet doesn't run and which would be genuine scope creep to add here - dropping swap detection is a deliberate, documented trade-off, not an oversight. TOP_NODES_OUTPUT="$(kubectl top nodes --no-headers 2>/dev/null)" \ @@ -127,38 +212,35 @@ fi USABLE_BUDGET_MIB=$(( USABLE_BUDGET_GI * 1024 )) HEADROOM_MIB=$(( USABLE_BUDGET_MIB - USAGE_MIB )) -CURRENT_MAX="$(kubectl get autoscalingrunnerset "$RELEASE_NAME" -n "$NAMESPACE" -o jsonpath='{.spec.maxRunners}' 2>/dev/null)" \ - || fail_safe "could not read ${RELEASE_NAME}'s current spec.maxRunners" -case "$CURRENT_MAX" in ''|*[!0-9]*) fail_safe "spec.maxRunners was not a plain integer ('$CURRENT_MAX')" ;; esac - -echo "R=${R} maxRunners=${CURRENT_MAX} Wh=${WH_MIB}MiB usage=${USAGE_MIB}MiB headroom=${HEADROOM_MIB}MiB mem_available=${mem_available_pct}% pressure=${pressure}" +echo "R_total=${R_TOTAL} maxRunners_total=${MAX_TOTAL} Wh=${WH_MIB}MiB usage=${USAGE_MIB}MiB headroom=${HEADROOM_MIB}MiB mem_available=${mem_available_pct}% pressure=${pressure} targets=${TARGETS[*]}" if [ "$pressure" = "true" ] || [ "$HEADROOM_MIB" -lt "$WH_MIB" ]; then - # Lower immediately, no delay or averaging — lowering never disrupts in-flight jobs (ARC only gates new claims), so there is no cost to being trigger-happy in this direction. Can drop below FLOOR (even to 0) if pressure is severe enough, but never below R: a currently-running job must never be orphaned by its own scale set shrinking under it. + # Lower immediately, no delay or averaging — lowering never disrupts in-flight jobs (ARC only gates new claims), so there is no cost to being trigger-happy in this direction. Can drop below FLOOR (even to 0, spread across targets by lower_to) if pressure is severe enough. configmap_set "raise-confirm-count" "0" - target=$FLOOR - [ "$pressure" = "true" ] && target=1 - [ "$target" -lt "$R" ] && target=$R - if [ "$target" -lt "$CURRENT_MAX" ]; then + target_total=$FLOOR + [ "$pressure" = "true" ] && target_total=1 + [ "$target_total" -lt "$R_TOTAL" ] && target_total=$R_TOTAL + if [ "$target_total" -lt "$MAX_TOTAL" ]; then reason="lower: headroom=${HEADROOM_MIB}MiB < Wh=${WH_MIB}MiB" [ "$pressure" = "true" ] && reason="lower: host pressure (mem_available=${mem_available_pct}%)" - patch_max_runners "$target" "$reason" "$HEADROOM_MIB" + lower_to "$target_total" + apply_new_max "$reason" "$HEADROOM_MIB" else - write_status "$CURRENT_MAX" "steady: already at or below the safe target" "$HEADROOM_MIB" + write_status "$MAX_TOTAL" "steady: already at or below the safe target" "$HEADROOM_MIB" fi -elif [ "$HEADROOM_MIB" -ge "$WH_MIB" ] && [ "$CURRENT_MAX" -lt "$MAX_CEILING" ]; then +elif [ "$HEADROOM_MIB" -ge "$WH_MIB" ] && [ "$MAX_TOTAL" -lt "$MAX_CEILING" ]; then confirm_count="$(configmap_get 'raise-confirm-count')" [ -n "$confirm_count" ] || confirm_count=0 confirm_count=$(( confirm_count + 1 )) configmap_set "raise-confirm-count" "$confirm_count" if [ "$confirm_count" -ge "$RAISE_CONFIRM_POLLS" ]; then configmap_set "raise-confirm-count" "0" - target=$(( CURRENT_MAX + 1 )) - patch_max_runners "$target" "raise: headroom=${HEADROOM_MIB}MiB confirmed for ${confirm_count} polls" "$HEADROOM_MIB" + raise_by_one + apply_new_max "raise: headroom=${HEADROOM_MIB}MiB confirmed for ${confirm_count} polls" "$HEADROOM_MIB" else - write_status "$CURRENT_MAX" "watching: headroom confirmed for ${confirm_count}/${RAISE_CONFIRM_POLLS} polls" "$HEADROOM_MIB" + write_status "$MAX_TOTAL" "watching: headroom confirmed for ${confirm_count}/${RAISE_CONFIRM_POLLS} polls" "$HEADROOM_MIB" fi else configmap_set "raise-confirm-count" "0" - write_status "$CURRENT_MAX" "steady" "$HEADROOM_MIB" + write_status "$MAX_TOTAL" "steady" "$HEADROOM_MIB" fi diff --git a/tests/unit/test_arc_filters.py b/tests/unit/test_arc_filters.py index 8f4a32d..ff57068 100644 --- a/tests/unit/test_arc_filters.py +++ b/tests/unit/test_arc_filters.py @@ -33,11 +33,11 @@ def test_derives_names_from_the_org_and_suffix(self) -> None: result = arc.arc_profiles([org(profiles=[{"suffix": "", "max_runners": 1}, {"suffix": "-builder", "values_file": "b.yaml", "runs_on_label": "image-builder", "node_selector": {"kubernetes.io/hostname": "n1"}}])]) self.assertEqual(result["errors"], []) default, builder = result["profiles"] - self.assertEqual((default["namespace"], default["release"], default["app_secret"], default["scale_set_labels"]), ("arc-runners-example", "example-runners", "example-github-app", ["example-runners"])) - self.assertEqual((builder["namespace"], builder["release"], builder["scale_set_labels"]), ("arc-runners-example-builder", "example-runners-builder", ["image-builder"])) + self.assertEqual((default["namespace"], default["release"], default["target"], default["app_secret"], default["scale_set_labels"]), ("arc-runners-example", "example-runners", "arc-runners-example/example-runners", "example-github-app", ["example-runners"])) + self.assertEqual((builder["namespace"], builder["release"], builder["target"], builder["scale_set_labels"]), ("arc-runners-example-builder", "example-runners-builder", "arc-runners-example-builder/example-runners-builder", ["image-builder"])) self.assertEqual(builder["node_selector"], {"kubernetes.io/hostname": "n1"}) self.assertEqual(default["node_selector"], {}) - self.assertIsNone(result["autoscaled"]) + self.assertEqual(result["autoscaled"], []) def test_a_missing_suffix_means_the_unsuffixed_profile(self) -> None: result = arc.arc_profiles([org(profiles=[{"values_file": "v.yaml"}])]) @@ -93,12 +93,20 @@ def test_two_profiles_of_an_org_sharing_a_release_name_is_an_error(self) -> None def test_one_autoscaled_profile_is_returned(self) -> None: result = arc.arc_profiles([org(profiles=[{"suffix": "", "autoscale": True, "max_runners": 1}, {"suffix": "-b", "max_runners": 1}])]) self.assertEqual(result["errors"], []) - self.assertEqual(result["autoscaled"]["release"], "example-runners") + self.assertEqual([profile["release"] for profile in result["autoscaled"]], ["example-runners"]) - def test_more_than_one_autoscaled_profile_is_an_error_across_orgs(self) -> None: - result = arc.arc_profiles([org("A", [{"autoscale": True}]), org("B", [{"autoscale": "yes"}])]) - self.assertTrue(any("At most one" in message for message in result["errors"])) - self.assertIsNone(result["autoscaled"]) + def test_more_than_one_autoscaled_profile_across_orgs_is_pooled_not_an_error(self) -> None: + result = arc.arc_profiles([org("A", [{"autoscale": True, "max_runners": 1}]), org("B", [{"autoscale": "yes", "max_runners": 1}])]) + self.assertEqual(result["errors"], []) + self.assertEqual({profile["label"] for profile in result["autoscaled"]}, {"A", "B"}) + + def test_more_than_one_autoscaled_profile_setting_sizing_is_an_error(self) -> None: + result = arc.arc_profiles([org("A", [{"autoscale": True, "sizing": SIZING, "values_file": "v.yaml"}]), org("B", [{"autoscale": True, "sizing": SIZING, "values_file": "v.yaml"}])]) + self.assertTrue(any("At most one autoscaled scale-set profile may set sizing" in message for message in result["errors"]), result["errors"]) + + def test_one_autoscaled_profile_may_set_sizing_alongside_a_pooled_static_one(self) -> None: + result = arc.arc_profiles([org("A", [{"autoscale": True, "sizing": SIZING, "values_file": "v.yaml"}]), org("B", [{"autoscale": True, "max_runners": 1}])]) + self.assertEqual(result["errors"], []) def test_invalid_autoscale_value_is_an_error(self) -> None: result = arc.arc_profiles([org(profiles=[{"autoscale": "sometimes"}])])