maintenance: move agents between management servers without failing or hanging jobs - #16
Merged
calvix merged 3 commits intoOct 5, 2026
Conversation
…ance countPendingNonPseudoJobs() filtered on instance_type != 'Thread'. VmWork jobs (VmWorkStart, VmWorkStop, VmWorkAttachVolume, ...) carry a NULL instance_type, and NULL != 'Thread' is not true in SQL, so they were never counted. The maintenance drain then reported 0 pending jobs and moved the agents away while VM work was still running on the management server, which failed that work. Count NULL instance types as well, and use the same ownership rule as the jobs a restarting management server fails (getResetJobs): executing on this server, or queued on it and not picked up yet.
…ng commands Preparing a management server for maintenance migrates its indirect agents with MigrateAgentConnectionCommand. The agent reconnected 3 s later whatever was still running, and the old management server treated the closed link as a failure: commands in flight lost their answers, the host went Alert with an agent investigation and HA work, and commands sent meanwhile failed with AgentUnavailableException. Under load every agent move failed operations, and once a live migration completed in libvirt while CloudStack recorded it as failed. Make the move a planned handoff: - Before the move command, put the host into Rebalancing. A disconnect while Rebalancing stays Rebalancing (no Alert, no investigation), and the agent's ping does not ask for a new startup that would pull it back. - Wait until this management server has no command queued or awaiting an answer for the host (indirect.agent.migration.idle.wait, default 600 s); otherwise leave the agent where it is. - Send the move command with send(): easySend() refuses hosts that are not Up/Connecting. - Hold new commands for a host that is Rebalancing or Connecting until it is Up on its new owner (agent.handoff.wait, default 180 s), instead of failing them. The move command itself is never held. - On the agent, finish running and queued commands before reconnecting, so their answers still go back over the current link. - Move one host at a time (indirect.agent.migration.parallelism, default 1), because the planner skips Rebalancing hosts. - If the agent does not come Up elsewhere, reconnect it instead of leaving the host Rebalancing.
When a management server restarts quickly, ClusterManagerImpl only notices the
new runid ("left and rejoined quickly") and raises no left/joined event, so the
other server keeps its cached channel to the old process. The first write to a
channel the peer has closed still succeeds locally and its data is lost; only
the next write fails with "Broken pipe" and triggers a reconnect. The lost write
is typically an agent answer routed back to the restarted server, so the job
waiting for it hangs until its timeout.
Peer channels are only ever written to. Before reusing a cached one, do a
non-blocking read: -1 means the peer closed it, so open a new connection first.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Putting a management server into maintenance (
prepareForMaintenance) is meant to allow rolling restartswithout disruption. Under load it was not: every agent move failed operations, and after the drained server was
restarted, jobs could hang until their timeout. Three separate defects, one commit each.
1. The drain did not wait for VM work (
jobs:)countPendingNonPseudoJobs()filtered oninstance_type != 'Thread'. VmWork jobs (VmWorkStart,VmWorkStop,VmWorkAttachVolume, ...) have a NULLinstance_type, andNULL != 'Thread'is not true in SQL, so they werenever counted. The drain reported 0 pending jobs and moved the agents while VM work was still running there. The
count now includes NULL instance types and uses the ownership rule of
getResetJobs(): executing here, or queuedhere and not picked up yet.
2. Moving an agent lost the commands in flight (
agent:)MigrateAgentConnectionCommandmade the agent reconnect 3 s later whatever was still running, and the oldmanagement server treated the closed link as a failure:
Alertwith an agent investigation and HA work;AgentUnavailableException/ "not in the right state: Alert";The move is now a planned handoff:
Rebalancingbefore the move command. A disconnect whileRebalancingstaysRebalancing(theFSM already has
Rebalancing + AgentDisconnected -> Rebalancing), so there is no Alert, investigation or HAwork, and the agent's ping no longer asks for a startup that would pull it back;
(
indirect.agent.migration.idle.wait, default 600 s); if that does not happen, the agent stays;send(), becauseeasySend()refuses hosts that are notUp/Connecting;RebalancingorConnectingwait until it isUpon its new owner(
agent.handoff.wait, default 180 s) instead of failing. Only the move command is never held;indirect.agent.migration.parallelism, default 1), because the planner skipsRebalancinghosts;Upelsewhere, it is reconnected instead of being leftRebalancing.3. The first answer after a peer restart was lost (
cluster:)After a quick restart
ClusterManagerImplonly notes the newrunid("left and rejoined quickly") and raises noleft/joined event, so the other server keeps its cached channel to the old process. The first write to a channel
the peer has closed still succeeds locally and is lost; only the next write fails with "Broken pipe" and
reconnects. The lost write is typically an agent answer routed back to the restarted server, so that job hangs
until its timeout. Peer channels are only ever written to, so a non-blocking read before reuse detects a closed
one (
-1) and a new connection is opened first.New settings (all
Advanced, dynamic):Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
Bug Severity
Screenshots (if appropriate):
n/a
How Has This Been Tested?
On a 2-MS KVM cluster (4.23.0.0, Ubuntu 26.04 hosts, Ceph RBD primary storage, 6 hosts + SSVM + CPVM agents),
with a continuous load running in parallel: VM stop/start, reboot, volume attach/detach, live migration, root
volume snapshot create/delete and VR firewall rule create/delete. Full procedure each time: LB maint ->
prepareForMaintenance-> wait forMaintenance, 0 agents, 0 jobs -> restart the MS ->cancelMaintenance rebalance=false-> LB ready.The last run moved all 8 agents one by one in 2 min 45 s; the slowest operation took 34 s (held while its host
moved). The remaining server logged one "has closed our connection (restarted?)" reconnect, no "Broken pipe" and
no command timeout. A separate end-to-end test suite ran on the same cluster at the same time: 96 jobs, all
successful, 30 of them during an agent move. Afterwards every VM's host and every volume's attachment in the DB
matched libvirt.
Unit tests:
ClusteredAgentManagerImplTest(6, including the newisPeerChannelClosedDetectsPeerThatClosedTheConnectionon a real local socket pair),AgentManagerImplTest(8),ManagementServerMaintenanceManagerImplTest(33), all passing.How did you try to break this feature and the system with this change?
easySend(), whichrefuses a
Rebalancinghost): the host must not stayRebalancing. It is now reconnected afteragent.handoff.wait, andreconnect()acceptsRebalancing.Connectingon the new owner:getAttache()only switches to a forwardingattache once the host is
Up, so held commands keep waiting throughConnecting.the agent waiting.
(the case that lost an answer before the third commit).
rebalance=false,cancelMaintenancemoves the agents back at once through the agents' preferred-hostcheck; that path does not use this handoff. Unchanged here, documented in our runbook.