Skip to content

maintenance: move agents between management servers without failing or hanging jobs - #16

Merged
calvix merged 3 commits into
integration/all-fixes-4.23.0.0from
fix/ms-maintenance-safe-agent-handoff-4.23
Oct 5, 2026
Merged

calvix merged 3 commits into
integration/all-fixes-4.23.0.0from
fix/ms-maintenance-safe-agent-handoff-4.23

Conversation

@calvix

@calvix calvix commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Description

Putting a management server into maintenance (prepareForMaintenance) is meant to allow rolling restarts
without disruption. Under load it was not: every agent move failed operations, and after the drained server was
restarted, jobs could hang until their timeout. Three separate defects, one commit each.

1. The drain did not wait for VM work (jobs:)

countPendingNonPseudoJobs() filtered on instance_type != 'Thread'. VmWork jobs (VmWorkStart, VmWorkStop,
VmWorkAttachVolume, ...) have a NULL instance_type, and NULL != 'Thread' is not true in SQL, so they were
never counted. The drain reported 0 pending jobs and moved the agents while VM work was still running there. The
count now includes NULL instance types and uses the ownership rule of getResetJobs(): executing here, or queued
here and not picked up yet.

2. Moving an agent lost the commands in flight (agent:)

MigrateAgentConnectionCommand made the agent reconnect 3 s later whatever was still running, and the old
management server treated the closed link as a failure:

  • answers to commands in flight were lost ("Unable to send response" on the agent);
  • the host went Alert with an agent investigation and HA work;
  • commands sent during the move failed with AgentUnavailableException / "not in the right state: Alert";
  • once a live migration completed in libvirt while CloudStack recorded it as failed, leaving the VM's host wrong.

The move is now a planned handoff:

  • the host goes Rebalancing before the move command. A disconnect while Rebalancing stays Rebalancing (the
    FSM already has Rebalancing + AgentDisconnected -> Rebalancing), so there is no Alert, investigation or HA
    work, and the agent's ping no longer asks for a startup that would pull it back;
  • the server waits until it has no command queued or awaiting an answer for that host before moving it
    (indirect.agent.migration.idle.wait, default 600 s); if that does not happen, the agent stays;
  • the move command is sent with send(), because easySend() refuses hosts that are not Up/Connecting;
  • new commands for a host that is Rebalancing or Connecting wait until it is Up on its new owner
    (agent.handoff.wait, default 180 s) instead of failing. Only the move command is never held;
  • the agent finishes running and queued commands before reconnecting, so their answers still use the old link;
  • one host at a time (indirect.agent.migration.parallelism, default 1), because the planner skips
    Rebalancing hosts;
  • if the agent does not come Up elsewhere, it is reconnected instead of being left Rebalancing.

3. The first answer after a peer restart was lost (cluster:)

After a quick restart ClusterManagerImpl only notes the new runid ("left and rejoined quickly") and raises no
left/joined event, so the other server keeps its cached channel to the old process. The first write to a channel
the peer has closed still succeeds locally and is lost; only the next write fails with "Broken pipe" and
reconnects. The lost write is typically an agent answer routed back to the restarted server, so that job hangs
until its timeout. Peer channels are only ever written to, so a non-blocking read before reuse detects a closed
one (-1) and a new connection is opened first.

New settings (all Advanced, dynamic):

setting                                 default  meaning
--------------------------------------  -------  ------------------------------------------------------------
agent.handoff.wait                      180      seconds a command waits for a host whose agent is moving
indirect.agent.migration.idle.wait      600      seconds to wait for a host's commands to finish before moving it
indirect.agent.migration.parallelism    1        agents moved at the same time during maintenance

Types of changes

  • Breaking change (fix or feature that would cause existing functionality to change)
  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement (improves an existing feature and functionality)
  • Cleanup (Code refactoring and cleanup, that may add test cases)
  • Build/CI
  • Test (unit or integration test code)

Feature/Enhancement Scale or Bug Severity

Feature/Enhancement Scale

  • Major
  • Minor

Bug Severity

  • BLOCKER
  • Critical
  • Major
  • Minor
  • Trivial

Screenshots (if appropriate):

n/a

How Has This Been Tested?

On a 2-MS KVM cluster (4.23.0.0, Ubuntu 26.04 hosts, Ceph RBD primary storage, 6 hosts + SSVM + CPVM agents),
with a continuous load running in parallel: VM stop/start, reboot, volume attach/detach, live migration, root
volume snapshot create/delete and VR firewall rule create/delete. Full procedure each time: LB maint ->
prepareForMaintenance -> wait for Maintenance, 0 agents, 0 jobs -> restart the MS ->
cancelMaintenance rebalance=false -> LB ready.

build                    ops   failed  hung  notes
-----------------------  ----  ------  ----  ---------------------------------------------------------------
unpatched, run 1         ~120  21      yes   failures in the 30 s move; a migration recorded wrongly
unpatched, run 2         ~140  16      yes   same; jobs hung up to 30 min until the drained MS was restarted
this branch, commits 1-2 344   0       1     one job hung right after the MS restart (defect 3)
this branch, all three   512   0       0     124 ops during the move (34 of them VR), 256 after the restart

The last run moved all 8 agents one by one in 2 min 45 s; the slowest operation took 34 s (held while its host
moved). The remaining server logged one "has closed our connection (restarted?)" reconnect, no "Broken pipe" and
no command timeout. A separate end-to-end test suite ran on the same cluster at the same time: 96 jobs, all
successful, 30 of them during an agent move. Afterwards every VM's host and every volume's attachment in the DB
matched libvirt.

Unit tests: ClusteredAgentManagerImplTest (6, including the new
isPeerChannelClosedDetectsPeerThatClosedTheConnection on a real local socket pair), AgentManagerImplTest (8),
ManagementServerMaintenanceManagerImplTest (33), all passing.

How did you try to break this feature and the system with this change?

  • A move command that never reaches the agent (the first version of this change sent it with easySend(), which
    refuses a Rebalancing host): the host must not stay Rebalancing. It is now reconnected after
    agent.handoff.wait, and reconnect() accepts Rebalancing.
  • Commands released while the host is Connecting on the new owner: getAttache() only switches to a forwarding
    attache once the host is Up, so held commands keep waiting through Connecting.
  • System VM agents (SSVM/CPVM) still run the old agent code: their move works, the server side does not depend on
    the agent waiting.
  • Restarting the drained MS immediately after the move and putting it straight back into the load balancer
    (the case that lost an answer before the third commit).
  • Without rebalance=false, cancelMaintenance moves the agents back at once through the agents' preferred-host
    check; that path does not use this handoff. Unchanged here, documented in our runbook.

calvix added 3 commits October 5, 2026 08:55
…ance

countPendingNonPseudoJobs() filtered on instance_type != 'Thread'. VmWork jobs
(VmWorkStart, VmWorkStop, VmWorkAttachVolume, ...) carry a NULL instance_type,
and NULL != 'Thread' is not true in SQL, so they were never counted. The
maintenance drain then reported 0 pending jobs and moved the agents away while
VM work was still running on the management server, which failed that work.

Count NULL instance types as well, and use the same ownership rule as the jobs a
restarting management server fails (getResetJobs): executing on this server, or
queued on it and not picked up yet.
…ng commands

Preparing a management server for maintenance migrates its indirect agents with
MigrateAgentConnectionCommand. The agent reconnected 3 s later whatever was still
running, and the old management server treated the closed link as a failure:
commands in flight lost their answers, the host went Alert with an agent
investigation and HA work, and commands sent meanwhile failed with
AgentUnavailableException. Under load every agent move failed operations, and
once a live migration completed in libvirt while CloudStack recorded it as
failed.

Make the move a planned handoff:
- Before the move command, put the host into Rebalancing. A disconnect while
  Rebalancing stays Rebalancing (no Alert, no investigation), and the agent's
  ping does not ask for a new startup that would pull it back.
- Wait until this management server has no command queued or awaiting an answer
  for the host (indirect.agent.migration.idle.wait, default 600 s); otherwise
  leave the agent where it is.
- Send the move command with send(): easySend() refuses hosts that are not
  Up/Connecting.
- Hold new commands for a host that is Rebalancing or Connecting until it is Up
  on its new owner (agent.handoff.wait, default 180 s), instead of failing them.
  The move command itself is never held.
- On the agent, finish running and queued commands before reconnecting, so their
  answers still go back over the current link.
- Move one host at a time (indirect.agent.migration.parallelism, default 1),
  because the planner skips Rebalancing hosts.
- If the agent does not come Up elsewhere, reconnect it instead of leaving the
  host Rebalancing.
When a management server restarts quickly, ClusterManagerImpl only notices the
new runid ("left and rejoined quickly") and raises no left/joined event, so the
other server keeps its cached channel to the old process. The first write to a
channel the peer has closed still succeeds locally and its data is lost; only
the next write fails with "Broken pipe" and triggers a reconnect. The lost write
is typically an agent answer routed back to the restarted server, so the job
waiting for it hangs until its timeout.

Peer channels are only ever written to. Before reusing a cached one, do a
non-blocking read: -1 means the peer closed it, so open a new connection first.
@calvix
calvix merged commit 55e13dc into integration/all-fixes-4.23.0.0 Oct 5, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant