Repository navigation
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## 4.22 #14344 +/- ##
============================================
+ Coverage 18.02% 18.06% +0.04%
- Complexity 16250 16272 +22
============================================
Files 5936 5936
Lines 535823 535866 +43
Branches 65612 65617 +5
============================================
+ Hits 96582 96817 +235
+ Misses 428242 428037 -205
- Partials 10999 11012 +13
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
0a1c6da to
70d43cd
Compare
| MigrateAnswer migrateAnswer = (MigrateAnswer) _agentManager.send(srcHost.getId(), migrateCommand); | ||
| boolean success = migrateAnswer != null && migrateAnswer.getResult(); | ||
|
|
||
| postMigrationHandled = true; |
There was a problem hiding this comment.
what happens if the migrate call times out while the vm is still moving? this flag is only set after it returns, so the cleanup could delete the new volume the vm ends up running on
There was a problem hiding this comment.
Thanks, now the new volumes are only cleaned up if the migration itself failed.
|
|
||
| String devPath = LinstorUtil.createResource( | ||
| destVolumeInfo, destStoragePool, _storagePoolDao, exactSize); | ||
| LinstorUtil.createResource(destVolumeInfo, destStoragePool, _storagePoolDao, exactSize); |
There was a problem hiding this comment.
if creating the linstor resource fails here, does the new volume row still get left behind? it only goes into the cleanup list a few lines later
There was a problem hiding this comment.
Also thanks, it is added to the cleanup list now after the DB object was created
…migration Live migration with storage into a Linstor pool spawned the destination resource through the resource group's auto-placement, which knows nothing about the target host. PrepareForMigrationCommand still describes the source disks, so the target agent never connected the new volume either. If the placement left the target host without a copy (place count 1, no quorum tiebreaker, 4+ nodes with 2 copies, node filters), the block device was missing there and libvirt aborted with "cannot precreate storage for disk type 'block'". Make the resource available on the target host right after creating it and take the device path from there. The source is not a Linstor resource of this pool, so no dual-primary is needed. A failure before the migration command (creating the resource, make-available, prepare for migration) now also removes the already created destination volumes; they used to be left behind. Once the migration command is sent, the destination volumes are only removed if the source host reports a failure. If the command times out, the VM may still end up on the new volumes: like StorageSystemDataMotionStrategy, check whether it runs on the target host and finish the migration in that case, otherwise keep the volumes. The error message names the destination pools instead of the source host. Unit tests cover make-available before the migration with the device path from the target host, the cleanup when creating the resource, make-available, prepare for migration or the migration itself fails, and a migration timeout with and without the VM running on the target host.
…rage LinstorDataMotionStrategy set the MigrateCommand CPU shares to the raw cpus * speed value, which is only valid on cgroup v1. On cgroup v2 hosts the source agent's updateVmSharesIfNeeded() sees that it differs from its own scaled value and writes it into the migration XML, so libvirt on the target rejects VMs above 10000 with "shares '<n>' must be in range [1, 10000]", and smaller VMs end up with an unscaled CPU weight. Use the value the target host calculated and returned in the PrepareForMigrationAnswer, as StorageSystemDataMotionStrategy and VirtualMachineManagerImpl do. A unit test checks that the target host's value is sent, not cpus * speed.
70d43cd to
97a1170
Compare
|
thanks @rp-, fix looks right. LGTM. |
Description
This PR fixes live migration with storage (
migrateVirtualMachineWithVolume) from another primarystorage (e.g. NFS) into a LINSTOR primary storage, which fails whenever the LINSTOR auto-placement
does not put a copy of the new resource on the migration target host:
LinstorDataMotionStrategy.copyAsyncspawns the destination resource from the resource group, soLINSTOR picks the nodes without knowing about the target host. The
PrepareForMigrationCommandsent to the target still describes the source disks, so the target agent never connects the new
volume either. If the target host has no copy,
/dev/drbd/by-res/cs-<uuid>/0does not exist there,and libvirt can only precreate missing file disks, not block devices.
On 3-node clusters with 2 replicas this rarely shows up, because LINSTOR adds a diskless quorum
tiebreaker on the third node. It fails with place count 1, with the tiebreaker disabled, with 4+
nodes and 2 replicas, or when the resource group's filters exclude the target host. Offline
migration (stop the VM, migrate the volume) is not affected, because starting the VM goes through
connectPhysicalDisk.Changes:
(
resourceMakeAvailableOnNode, a diskless resource for DRBD) and take the device path fromthere. No dual-primary is needed, the source disk is not this LINSTOR resource.
the already created destination volumes are now destroyed and expunged. Before, they were left
behind in LINSTOR and in the
volumestable.The strategy was added in #12532, so 4.20.3.0 and 4.22.1.0 are affected.
Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
Bug Severity
Screenshots (if appropriate):
How Has This Been Tested?
3 KVM hosts (Ubuntu 24.04), LINSTOR 1.35.2 / DRBD 9.3, a 4.22.1.1-based CloudStack build with this
patch on the management server (
LinstorDataMotionStrategyis identical to the 4.22 branch). To make the failure deterministic on 3 nodes, the target LINSTOR pool uses aresource group with place count 1 restricted to a storage pool that only exists on host 2. A
running VM with its root disk on an NFS primary storage is migrated from host 1 to host 3 with
migrateVirtualMachineWithVolume, root volume to the LINSTOR pool.cannot precreate storage for disk type 'block'(reproduced)./dev/drbd/by-res/cs-<uuid>/0(diskless, InUse; diskful copy on host 2 only), a 64 MiB randomfile written in the guest before the migration has the same sha256 afterwards, the NFS volume
is expunged.
How did you try to break this feature and the system with this change?
LINSTOR error, the VM stays running on the source host on NFS, the destination volume is
expunged and its LINSTOR resource definition is deleted.
target host. The migration fails in
blockdev-add, the VM stays on the source host, thedestination resource definition, including the new diskless resource on the target, is deleted.
Not tested: shared (thick LVM) LINSTOR storage pools, VMs with multiple volumes.