Skip to content

feat(arc): pool multiple autoscaled profiles into one shared budget - #23

Merged
Mearman merged 2 commits into
mainfrom
feat/pool-multiple-autoscaled-profiles
Sep 28, 2026
Merged

Mearman merged 2 commits into
mainfrom
feat/pool-multiple-autoscaled-profiles

Conversation

@Mearman

@Mearman Mearman commented Sep 28, 2026

Copy link
Copy Markdown
Member

Summary

  • plugins/filter/arc.py's arc_profiles previously rejected more than one scale-set profile setting autoscale: true fleet-wide. Now allows several (from different orgs), pooling their combined scale sets against one shared memory budget instead of a fixed per-org slice. At most one pooled profile may still set its own sizing, since that's what derives the shared ceiling.
  • scripts/autoscaler.sh takes a new AUTOSCALER_TARGETS env var (whitespace-separated namespace/release tokens) instead of a single AUTOSCALER_NAMESPACE/AUTOSCALER_RELEASE_NAME pair. Each poll gathers every pooled target's currentRunners/maxRunners/memory usage, runs the existing usage-driven algorithm against the combined totals, then distributes any change: a raise goes to whichever target has the least headroom (closest to actually needing it), a lower comes from whichever has the most, one runner at a time, never below a target's own currentRunners.
  • install_platform.yml loops its per-profile RBAC (Role/RoleBinding, one pair per pooled namespace) and namespace creation over every pooled profile instead of assuming one.
  • install_org.yml's post-upgrade autoscaler reconcile trigger fires for any org that owns a pooled profile, not a single fixed one.
  • README's Autoscaler section rewritten for pooling.

Motivation

A second org (ExaCapLtd) wants to share this fleet's runner capacity with ExaDev's existing pool, on the same physical nodes, without carving out a fixed, wasteful static slice of the budget for it.

Test plan

  • Full unit suite: python3 -m unittest tests.unit.test_arc_filters, 47 tests, including new coverage for pooling (test_more_than_one_autoscaled_profile_across_orgs_is_pooled_not_an_error) and the one-sizing-profile constraint (test_more_than_one_autoscaled_profile_setting_sizing_is_an_error, test_one_autoscaled_profile_may_set_sizing_alongside_a_pooled_static_one).
  • ansible-playbook playbooks/arc.yml --syntax-check and playbooks/site.yml --syntax-check.
  • Real ansible-playbook arc.yml --check --diff --limit mini against the live fleet inventory with a second org's profile added: the validation layer (arc_profiles expansion plus every "Fail on invalid" check) accepted it with zero errors, correctly expanding ExaCapLtd alongside ExaDev/ExaDev-builder.
  • Scripted end-to-end run of autoscaler.sh against a fake two-target kubectl, in both dry-run "steady" and "raise" scenarios: confirmed the distribution logic (tie-break to first, least-headroom-gets-the-raise, most-headroom-gives-up-the-lower) and the per-target status JSON.

Not yet run against the live cluster past the validation layer: reading the fleet's 1Password secrets needs an interactive biometric approval I couldn't complete in this environment, so the actual Helm/RBAC apply hasn't been dry-run end to end. Recommend a --check run (or a real, targeted apply) from a session that can complete that approval before merging.

arc_profiles previously rejected more than one scale-set profile
setting autoscale: true, since the autoscaler budgeted one pool's
memory for exactly one scale set. Lets several profiles (from
different orgs) share that budget instead: at most one may still set
its own sizing, since that derives the shared ceiling.

autoscaler.sh now takes a whitespace-separated AUTOSCALER_TARGETS list
instead of one AUTOSCALER_NAMESPACE/AUTOSCALER_RELEASE_NAME pair. Each
poll gathers every pooled target's own currentRunners/maxRunners and
memory usage, applies the existing usage-driven algorithm to the
combined totals, then hands a raise to whichever target has the least
headroom and takes a lower from whichever has the most, one runner at
a time, never below a target's own currentRunners.

install_platform.yml loops its per-profile RBAC and namespace creation
over every pooled profile, and install_org.yml's post-upgrade
reconcile trigger fires for any org with a pooled profile rather than
a single fixed one.

Verified: the full unit suite (47 tests, including new coverage for
pooling and for the one-sizing-profile constraint), a real --check
dry run against the live fleet inventory with a second org's profile
added (the validation layer accepted it with zero errors), and a
scripted end-to-end run of autoscaler.sh itself against a fake
two-target kubectl, confirming both the lower and raise distribution
logic and the per-target status reporting.
@Mearman Mearman changed the title Pool multiple autoscaled profiles into one shared memory budget feat(arc): pool multiple autoscaled profiles into one shared budget Sep 28, 2026
shellcheck SC2004: an array subscript on the left of an assignment is
already an arithmetic context, so $i/$best there is redundant.
@Mearman
Mearman marked this pull request as ready for review September 28, 2026 18:54
@Mearman
Mearman merged commit 78be6df into main Sep 28, 2026
20 checks passed
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-28T19:04:14.775901Z e21d93a Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@Mearman
Mearman deleted the feat/pool-multiple-autoscaled-profiles branch September 28, 2026 18:54
@github-actions

Copy link
Copy Markdown

🎉 This PR is included in version 1.5.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant