feat(arc): pool multiple autoscaled profiles into one shared budget - #23
Merged
Merged
Conversation
arc_profiles previously rejected more than one scale-set profile setting autoscale: true, since the autoscaler budgeted one pool's memory for exactly one scale set. Lets several profiles (from different orgs) share that budget instead: at most one may still set its own sizing, since that derives the shared ceiling. autoscaler.sh now takes a whitespace-separated AUTOSCALER_TARGETS list instead of one AUTOSCALER_NAMESPACE/AUTOSCALER_RELEASE_NAME pair. Each poll gathers every pooled target's own currentRunners/maxRunners and memory usage, applies the existing usage-driven algorithm to the combined totals, then hands a raise to whichever target has the least headroom and takes a lower from whichever has the most, one runner at a time, never below a target's own currentRunners. install_platform.yml loops its per-profile RBAC and namespace creation over every pooled profile, and install_org.yml's post-upgrade reconcile trigger fires for any org with a pooled profile rather than a single fixed one. Verified: the full unit suite (47 tests, including new coverage for pooling and for the one-sizing-profile constraint), a real --check dry run against the live fleet inventory with a second org's profile added (the validation layer accepted it with zero errors), and a scripted end-to-end run of autoscaler.sh itself against a fake two-target kubectl, confirming both the lower and raise distribution logic and the per-target status reporting.
shellcheck SC2004: an array subscript on the left of an assignment is already an arithmetic context, so $i/$best there is redundant.
Mearman
marked this pull request as ready for review
September 28, 2026 18:54
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
🎉 This PR is included in version 1.5.0 🎉 The release is available on:
Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
plugins/filter/arc.py'sarc_profilespreviously rejected more than one scale-set profile settingautoscale: truefleet-wide. Now allows several (from different orgs), pooling their combined scale sets against one shared memory budget instead of a fixed per-org slice. At most one pooled profile may still set its ownsizing, since that's what derives the shared ceiling.scripts/autoscaler.shtakes a newAUTOSCALER_TARGETSenv var (whitespace-separatednamespace/releasetokens) instead of a singleAUTOSCALER_NAMESPACE/AUTOSCALER_RELEASE_NAMEpair. Each poll gathers every pooled target'scurrentRunners/maxRunners/memory usage, runs the existing usage-driven algorithm against the combined totals, then distributes any change: a raise goes to whichever target has the least headroom (closest to actually needing it), a lower comes from whichever has the most, one runner at a time, never below a target's owncurrentRunners.install_platform.ymlloops its per-profile RBAC (Role/RoleBinding, one pair per pooled namespace) and namespace creation over every pooled profile instead of assuming one.install_org.yml's post-upgrade autoscaler reconcile trigger fires for any org that owns a pooled profile, not a single fixed one.Motivation
A second org (ExaCapLtd) wants to share this fleet's runner capacity with ExaDev's existing pool, on the same physical nodes, without carving out a fixed, wasteful static slice of the budget for it.
Test plan
python3 -m unittest tests.unit.test_arc_filters, 47 tests, including new coverage for pooling (test_more_than_one_autoscaled_profile_across_orgs_is_pooled_not_an_error) and the one-sizing-profile constraint (test_more_than_one_autoscaled_profile_setting_sizing_is_an_error,test_one_autoscaled_profile_may_set_sizing_alongside_a_pooled_static_one).ansible-playbook playbooks/arc.yml --syntax-checkandplaybooks/site.yml --syntax-check.ansible-playbook arc.yml --check --diff --limit miniagainst the live fleet inventory with a second org's profile added: the validation layer (arc_profilesexpansion plus every "Fail on invalid" check) accepted it with zero errors, correctly expandingExaCapLtdalongsideExaDev/ExaDev-builder.autoscaler.shagainst a fake two-targetkubectl, in both dry-run "steady" and "raise" scenarios: confirmed the distribution logic (tie-break to first, least-headroom-gets-the-raise, most-headroom-gives-up-the-lower) and the per-target status JSON.Not yet run against the live cluster past the validation layer: reading the fleet's 1Password secrets needs an interactive biometric approval I couldn't complete in this environment, so the actual Helm/RBAC apply hasn't been dry-run end to end. Recommend a
--checkrun (or a real, targeted apply) from a session that can complete that approval before merging.