# Proxmox snapshot backup of containers fails sometimes

**URL:** <https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134>\
**Category:** Proxmox VE\
**Tags:** drbd, linstor, proxmox\
**Created:** [January 13, 2026, 9:24am UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134 "2026-01-13T09:24:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![espenu](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@espenu](https://forums.linbit.com/u/espenu)\
**Post date:** [January 13, 2026, 9:24am UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/1 "2026-01-13T09:24:40Z")

</div>

My setup has two storage nodes, each with 2x nvme drives mirrored with mdadm, an lvm on that, and then drbd. In addition I have a VM running on TrueNAS as a diskless witness.  
The two storage nodes are connected with a dedicated 25Gb link.

However, sometimes container backups will fail. It seems to happen when both hypervisors start a backup at the same time. It’s usually not the same container that fails. When failed I have to manually remove the snapshot and snapshot related resources.

Any ideas?

Linstor controller version: 1.33.1  
Kernel: 6.17.2-2-pve.

Error from Proxmox:

> INFO: starting new backup job: vzdump 106 --mode snapshot --storage nas-backup --notes-template ‘{{guestname}}’ --notification-mode notification-system --node hypervisor1 --compress zstd --remove 1  
> INFO: Starting Backup of VM 106 (lxc)  
> INFO: Backup started at 2026-01-07 20:39:10  
> INFO: status = running  
> INFO: CT Name: nextcloud  
> INFO: including mount point rootfs (‘/’) in backup  
> INFO: excluding bind mount point mp0 (‘/mnt/data’) from backup (not a volume)  
> INFO: excluding bind mount point mp1 (‘/mnt/uploadtemp’) from backup (not a volume)  
> INFO: found old vzdump snapshot (force removal)  
> INFO: backup mode: snapshot  
> INFO: ionice priority: 7  
> INFO: create storage snapshot ‘vzdump’  
> mount: /mnt/vzsnap0: fsconfig() failed: /dev/drbd1019: Can’t open blockdev.  
> dmesg(1) may have more information after failed mount system call.  
> umount: /mnt/vzsnap0/: not mounted.  
> command ‘umount -l -d /mnt/vzsnap0/’ failed: exit code 32  
> ERROR: Backup of VM 106 failed - command ‘mount -o ro,noload /dev/drbd1019 /mnt/vzsnap0//’ failed: exit code 32  
> INFO: Failed at 2026-01-07 20:39:27  
> INFO: Backup job finished with errors

> **Error log from hypervisor where the backup failed (reverse direction):**
>
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: Began resync as SyncSource (will sync 14680064 KB [3670016 bits set]).  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: pdsk( Outdated → Inconsistent ) repl( WFBitMapS → SyncSource ) replication( yes → no ) [receive-b\>  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: helper command: /sbin/drbdadm before-resync-source exit code 0  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: helper command: /sbin/drbdadm before-resync-source  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: receive bitmap stats [Bytes(packets)]: plain 0(0), RLE 24(1), total 24; compression: 100.0%  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: send bitmap stats [Bytes(packets)]: plain 0(0), RLE 24(1), total 24; compression: 100.0%  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: pdsk( Consistent → Outdated ) [peer-state]  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: pdsk( DUnknown → Consistent ) repl( Off → WFBitMapS ) [connected]  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: quorum( no → yes ) [connected]  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( Connecting → Connected ) peer( Unknown → Secondary ) [connected]  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Committing cluster-wide state change 1782363112 (23ms)  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: State change 1782363112: primary\_nodes=0, weak\_nodes=0  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: uuid\_compare()=source-if-both-failed by rule=both-off  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: peer C1B6BD4A543399DC:0000000000000000:0000000000000000:0000000000000000 bits:1750016 flags:1020  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: self C1B6BD4A543399DC:0000000000000000:0000000000000000:0000000000000000 bits:3670016 flags:22  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 hypervisor2: drbd\_sync\_handshake:  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Preparing cluster-wide state change 1782363112: 0-\>1 role( Secondary ) conn( Connected )  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Peer authenticated using 20 bytes HMAC  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Feature flags enabled on protocol level: 0x1ff TRIM THIN\_RESYNC WRITE\_SAME WRITE\_ZEROES RESYNC\_DAGTAG  
> Jan 07 20:39:37 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Handshake to peer 1 successful: Agreed network protocol version 123  
> Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: INFO: Backup job finished with errors  
> Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: ERROR: Backup of VM 106 failed - command ‘mount -o ro,noload /dev/drbd1019 /mnt/vzsnap0//’ failed: exit code 32  
> Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: command ‘umount -l -d /mnt/vzsnap0/’ failed: exit code 32  
> Jan 07 20:39:27 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:27.204 [DeviceManager] INFO LINSTOR/Satellite/3c8fd2 SYSTEM - Begin DeviceManager cycle 106  
> Jan 07 20:39:27 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:27.204 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 105  
> Jan 07 20:39:27 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:27.143 [DeviceManager] INFO LINSTOR/Satellite/952b0b SYSTEM - Resource ‘snap\_pm-b5d0916f\_vzdump’ [DRBD] adjusted.  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( Unconnected → Connecting ) [connecting]  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Restarting receiver thread  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( NetworkFailure → Unconnected ) [disconnected]  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: helper command: /sbin/drbdadm disconnected exit code 0  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: helper command: /sbin/drbdadm disconnected  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Connection closed  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Starting sender thread (peer-node-id 1)  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Terminating sender thread  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Failure to connect: Interrupted state change (-21); retrying  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Aborting cluster-wide state change 884033480 (19ms) rv = -21  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: meta connection shut down by peer.  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( Connecting → NetworkFailure )  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: sock was shut down by peer  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: meta connection shut down by peer.  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Preparing cluster-wide state change 884033480: 0-\>1 role( Secondary ) conn( Connected )  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Peer authenticated using 20 bytes HMAC  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Feature flags enabled on protocol level: 0x1ff TRIM THIN\_RESYNC WRITE\_SAME WRITE\_ZEROES RESYNC\_DAGTAG  
> Jan 07 20:39:25 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Handshake to peer 1 successful: Agreed network protocol version 123  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.540 [DeviceManager] INFO LINSTOR/Satellite/952b0b SYSTEM - Begin DeviceManager cycle 105  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.540 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 104  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.484 [MainWorkerPool-12] INFO LINSTOR/Satellite/6d995a SYSTEM - Storage pool ‘pve-storage’ for node ‘hypervisor1’ updated.  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.484 [DeviceManager] INFO LINSTOR/Satellite/f42f14 SYSTEM - Begin DeviceManager cycle 104  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.484 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 103  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.483 [MainWorkerPool-11] INFO LINSTOR/Satellite/6fac8f SYSTEM - Storage pool ‘pve-storage’ for node ‘hypervisor2’ updated.  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.482 [MainWorkerPool-8] INFO LINSTOR/Satellite/001f93 SYSTEM - SpaceInfo: pve-storage → 458506245/742039552  
> Jan 07 20:39:16 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:16.419 [MainWorkerPool-8] INFO LINSTOR/Satellite/001f93 SYSTEM - SpaceInfo: DfltDisklessStorPool → 9223372036854775807/922\>  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Committing remote state change 260165721 (primary\_nodes=0)  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Preparing remote state change 260165721: 1-\>2 role( Secondary ) conn( Connected )  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 qdevice: pdsk( DUnknown → Diskless ) repl( Off → Established ) [connected]  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: conn( Connecting → Connected ) peer( Unknown → Secondary ) [connected]  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Committing cluster-wide state change 296415547 (21ms)  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: State change 296415547: primary\_nodes=0, weak\_nodes=0  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 qdevice: peer’s exposed UUID: 0000000000000000  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019 qdevice: self C1B6BD4A543399DC:0000000000000000:0000000000000000:0000000000000000 bits:0 flags:0  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Preparing cluster-wide state change 296415547: 0-\>2 role( Secondary ) conn( Connected )  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Peer authenticated using 20 bytes HMAC  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Feature flags enabled on protocol level: 0x1ff TRIM THIN\_RESYNC WRITE\_SAME WRITE\_ZEROES RESYNC\_DAGTAG  
> Jan 07 20:39:12 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Handshake to peer 2 successful: Agreed network protocol version 123  
> Jan 07 20:39:12 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:12.057 [DeviceManager] INFO LINSTOR/Satellite/b6d228 SYSTEM - Begin DeviceManager cycle 103  
> Jan 07 20:39:12 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:12.057 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 102  
> Jan 07 20:39:12 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:12.052 [DeviceManager] INFO LINSTOR/Satellite/86027d SYSTEM - Resource ‘snap\_pm-b5d0916f\_vzdump’ [DRBD] adjusted.  
> Jan 07 20:39:12 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:12.039 [DeviceManager] INFO LINSTOR/Satellite/86027d SYSTEM - Begin DeviceManager cycle 102  
> Jan 07 20:39:12 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:12.039 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 101  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: conn( Unconnected → Connecting ) [connecting]  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Starting receiver thread (peer-node-id 2)  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( Unconnected → Connecting ) [connecting]  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.913 [DeviceManager] INFO LINSTOR/Satellite/f73358 SYSTEM - Resource ‘snap\_pm-b5d0916f\_vzdump’ [DRBD] adjusted.  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: conn( StandAlone → Unconnected ) [connect]  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Starting receiver thread (peer-node-id 1)  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: conn( StandAlone → Unconnected ) [connect]  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: Setting exposed data uuid: C1B6BD4A543399DC  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: attached to current UUID: C1B6BD4A543399DC  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: disk( Attaching → UpToDate ) [attach]  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: size = 14 GB (14680984 KB)  
> Jan 07 20:39:11 hypervisor1 kernel: drbd1019: detected capacity change from 0 to 29361968  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: resync bitmap: bits=3670246 bits\_4k=3670246 words=401436 pages=785  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: drbd\_bm\_resize called with capacity == 29361968  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Method to ensure write ordering: flush  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: Maximum number of peer devices = 7  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: disk( Diskless → Attaching ) [attach]  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump/0 drbd1019: meta-data IO uses: blk-bio  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump qdevice: Starting sender thread (peer-node-id 2)  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump hypervisor2: Starting sender thread (peer-node-id 1)  
> Jan 07 20:39:11 hypervisor1 kernel: drbd snap\_pm-b5d0916f\_vzdump: Starting worker thread (node-id 0)  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.874 [DeviceManager] INFO LINSTOR/Satellite/f73358 SYSTEM - DRBD regenerated resource file: /var/lib/linstor.d/snap\_pm-b5\>  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.871 [DeviceManager] INFO LINSTOR/Satellite/f73358 SYSTEM - Volume number 0 of resource ‘snap\_pm-b5d0916f\_vzdump’ [LVM-Th\>  
> Jan 07 20:39:11 hypervisor1 dmeventd[979]: Monitoring thin pool linstor\_vg-thinpool-tpool.  
> Jan 07 20:39:11 hypervisor1 dmeventd[979]: No longer monitoring thin pool linstor\_vg-thinpool-tpool.  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.658 [DeviceManager] INFO LINSTOR/Satellite/f73358 SYSTEM - Begin DeviceManager cycle 101  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.658 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 100  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.653 [DeviceManager] INFO LINSTOR/Satellite/f09f31 SYSTEM - Resource ‘pm-b5d0916f’ [DRBD] adjusted.  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.642 [MainWorkerPool-1] INFO LINSTOR/Satellite/318524 SYSTEM - Snapshot ‘snap\_pm-b5d0916f\_vzdump’ of resource 'pm-b5d0916\>  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.571 [DeviceManager] INFO LINSTOR/Satellite/f09f31 SYSTEM - Begin DeviceManager cycle 100  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.571 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 99  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.566 [DeviceManager] INFO LINSTOR/Satellite/8d9e26 SYSTEM - Resource ‘pm-b5d0916f’ [DRBD] adjusted.  
> Jan 07 20:39:11 hypervisor1 kernel: drbd pm-b5d0916f: susp-io( user → no ) [resume-io]  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.548 [MainWorkerPool-6] INFO LINSTOR/Satellite/fcbfd7 SYSTEM - Snapshot ‘snap\_pm-b5d0916f\_vzdump’ of resource 'pm-b5d0916\>  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.544 [DeviceManager] INFO LINSTOR/Satellite/8d9e26 SYSTEM - Begin DeviceManager cycle 99  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.544 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 98  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.443 [DeviceManager] INFO LINSTOR/Satellite/537707 SYSTEM - Snapshot [LVM-Thin] with name ‘snap\_pm-b5d0916f\_vzdump’ of re\>  
> Jan 07 20:39:11 hypervisor1 dmeventd[979]: Monitoring thin pool linstor\_vg-thinpool-tpool.  
> Jan 07 20:39:11 hypervisor1 dmeventd[979]: No longer monitoring thin pool linstor\_vg-thinpool-tpool.  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.314 [DeviceManager] INFO LINSTOR/Satellite/537707 SYSTEM - Resource ‘pm-b5d0916f’ [DRBD] adjusted.  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.303 [MainWorkerPool-3] INFO LINSTOR/Satellite/de03aa SYSTEM - Snapshot ‘snap\_pm-b5d0916f\_vzdump’ of resource 'pm-b5d0916\>  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.301 [DeviceManager] INFO LINSTOR/Satellite/537707 SYSTEM - Begin DeviceManager cycle 98  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.300 [DeviceManager] INFO LINSTOR/Satellite/ SYSTEM - End DeviceManager cycle 97  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.243 [DeviceManager] INFO LINSTOR/Satellite/f80f19 SYSTEM - Resource ‘pm-b5d0916f’ [DRBD] adjusted.  
> Jan 07 20:39:11 hypervisor1 kernel: drbd pm-b5d0916f: susp-io( no → user ) [suspend-io]  
> Jan 07 20:39:11 hypervisor1 Satellite[3554249]: 2026-01-07 20:39:11.123 [MainWorkerPool-16] INFO LINSTOR/Satellite/05a4c2 SYSTEM - Snapshot ‘snap\_pm-b5d0916f\_vzdump’ of resource 'pm-b5d091\>  
> Jan 07 20:39:10 hypervisor1 pvedaemon[2641647]: INFO: Starting Backup of VM 106 (lxc)

---

<div class="post-metadata">

**Author:** ![ghernadi](https://avatars.discourse-cdn.com/v4/letter/g/9e8a1a/32.png) [@ghernadi](https://forums.linbit.com/u/ghernadi)\
**Post date:** [January 27, 2026, 12:16pm UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/2 "2026-01-27T12:16:57Z")

</div>

Hello, although I do apologize for the late response, unfortunately I am not quite sure how to help you other than reading the logs back to you:

You do have some `NetworkFailure`:

```auto
Jan 07 20:39:25 hypervisor1 kernel: drbd snap_pm-b5d0916f_vzdump hypervisor2: conn( Connecting → NetworkFailure )

```

But I don’t think that this is the root cause. First I thought that the DRBD device cannot be mounted, before I realized that I misread the error message:

```auto
# chronological order, not reversed
Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: command ‘umount -l -d /mnt/vzsnap0/’ failed: exit code 32
Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: ERROR: Backup of VM 106 failed - command ‘mount -o ro,noload /dev/drbd1019 /mnt/vzsnap0//’ failed: exit code 32
Jan 07 20:39:27 hypervisor1 pvedaemon[2641647]: INFO: Backup job finished with errors

```

The first command that fails is an `umount`, not a `mount`. I assume the following `mount` fails because the target directory is still used by a different `mount`?

I am not a Proxmox expert here so I cannot really tell what Proxmox tries to do any why or what would still have the target-directory mounted. From my perspective this has nothing to do with LINSTOR. Even the `NetworkFailure` shown by DRBD should not have an effect here, since `hypervisor1` and `qdevice` are connected and should have quorum anyways.

---

<div class="post-metadata">

**Author:** ![espenu](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@espenu](https://forums.linbit.com/u/espenu)\
**Post date:** [January 30, 2026, 8:01am UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/3 "2026-01-30T08:01:31Z")

</div>

Could it be related to the way I’ve set up my paths? I have a 25Gb direct link primary path for the two storage nodes. Then a backup path over a 10Gb switched link. The qdevice has only one path to each storage node, which is over the 10Gb network.  
If it’s relevant, here are my DRBD options:

`DrbdOptions/Disk/al-extents: 6433`  
`DrbdOptions/Disk/on-io-error: detach`  
`DrbdOptions/Net/max-buffers: 36000`  
`DrbdOptions/Net/max-epoch-size: 20000`  
`DrbdOptions/Net/protocol: C`  
`DrbdOptions/Net/rcvbuf-size: 10485760`  
`DrbdOptions/Net/sndbuf-size: 10485760`  
`DrbdOptions/PeerDevice/c-fill-target: 49152`  
`DrbdOptions/PeerDevice/c-max-rate: 2252800`

I also tried to set the buffer and epoch values to the DfltRscGrp as I noticed that the snapshots were created under that group (which is odd since I never assigned a storage pool to the default group, but I guess that’s just how it works?)

---

<div class="post-metadata">

**Author:** ![ghernadi](https://avatars.discourse-cdn.com/v4/letter/g/9e8a1a/32.png) [@ghernadi](https://forums.linbit.com/u/ghernadi)\
**Post date:** [February 2, 2026, 6:02am UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/4 "2026-02-02T06:02:00Z")

</div>

I cannot really tell you more about the `NetworkFailure` than what you shared here. I do not think that only because the backup network is slower than the replication network should be an issue here.

My guess is that you only shared some DRBD specific logs (which was good), but if you want to find out the issue of the network failure you might want to check the general logs and see if there are other journal entries for example around that time.

---

<div class="post-metadata">

**Author:** ![espenu](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@espenu](https://forums.linbit.com/u/espenu)\
**Post date:** [February 3, 2026, 6:39am UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/5 "2026-02-03T06:39:49Z")

</div>

Those are actually a full journalctl dump, and not specifically DRBD logs.  
I have an issue with one of my servers (it occasionally kernel panics) and as a result most VMs and containers have been moved to a single server until that’s sorted. This has shown that the backup issue happens even if it’s one server doing sequential backups.

I’ll be re-doing my network to use 25Gb LAG on each server later this year, and hopefully that will resolve it. In the meantime I’ll just make sure container backups are staggered by a few minutes, as it seems the first backup always goes through.

---

<div class="post-metadata">

**Author:** ![espenu](https://avatars.discourse-cdn.com/v4/letter/e/9d8465/32.png) [@espenu](https://forums.linbit.com/u/espenu)\
**Post date:** [August 17, 2026, 5:32pm UTC](https://forums.linbit.com/t/proxmox-snapshot-backup-of-containers-fails-sometimes/1134/6 "2026-08-17T17:32:54Z")

</div>

This issue seems to have been resolved with a sw update within the last 2-3 months 😄
