problem
With Database High Availability enabled (db.ha.enabled=true), after the MySQL source becomes unreachable the Management Server UI becomes extremely slow / barely usable.
Failover to the configured replica seems to happen via the MySQL connector, but each (or almost each) request appears to retry the dead source first, causing severe UI latency.
Workaround that restores normal performance:
- disable db.ha.enabled
- manually promote the replica (STOP REPLICA / RESET REPLICA ALL / read_only=OFF)
- point db.cloud.host / db.usage.host to the promoted DB in db.properties
- restart cloudstack-management
This matches the install guide "Failover" procedure better than the client-side db.ha mechanism.
Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way replication.
Ref: https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability
versions
CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
OS: Ubuntu 24.04
MySQL: 8.0.46 (async replication source→replica, GTID)
Hypervisor: VMware
Topology: 2 Management Servers (multi-site), reverse proxy in front of UI
The steps to reproduce the bug
- Deploy 2 MS pointing to MySQL source; configure two-way replica (GTID)
- Set in db.properties:
db.ha.enabled=true
db.cloud.replicas=replica-ip
db.usage.replicas=replica-ip
- Restart cloudstack-management on both nodes
- Stop / isolate the MySQL source (simulate site/DB failure)
- Use UI/API through the remaining Management Server
What to do about it?
Expected: MS fail over to replica promptly; UI remains usable.
Actual: UI becomes very slow after source outage (connector retries against unreachable source; defaults secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 3600/5000/3600).
Suggestions:
- fix connector failover so replica is used without per-request source retries hanging the UI
- and/or document clearly that db.ha.enabled is not recommended for production on MySQL 8.x and that manual promotion (install guide Failover) is the supported path
problem
With Database High Availability enabled (db.ha.enabled=true), after the MySQL source becomes unreachable the Management Server UI becomes extremely slow / barely usable.
Failover to the configured replica seems to happen via the MySQL connector, but each (or almost each) request appears to retry the dead source first, causing severe UI latency.
Workaround that restores normal performance:
This matches the install guide "Failover" procedure better than the client-side db.ha mechanism.
Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way replication.
Ref: https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability
versions
CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
OS: Ubuntu 24.04
MySQL: 8.0.46 (async replication source→replica, GTID)
Hypervisor: VMware
Topology: 2 Management Servers (multi-site), reverse proxy in front of UI
The steps to reproduce the bug
db.ha.enabled=true
db.cloud.replicas=replica-ip
db.usage.replicas=replica-ip
What to do about it?
Expected: MS fail over to replica promptly; UI remains usable.
Actual: UI becomes very slow after source outage (connector retries against unreachable source; defaults secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 3600/5000/3600).
Suggestions: